Glossary
Document capture
Also called: capture, intake
Document capture is the process of getting documents into a system and making them usable: scanning or importing, recognising text, classifying, extracting fields and applying metadata. Scanning is only the first step, and a project that stops there produces storage rather than records.
Document capture explained
Scanning is not capture
Scanning produces an image. Capture produces a record: recognised text, a document type, mandatory metadata, a document number and a retention rule. The distinction is the reason so many digitisation projects deliver a large volume of unsearchable image files and no change in how work is done.
The intake route decides everything
Documents arrive by email, by post, from a scanner, from a mobile phone, from another system, and occasionally by hand. Whichever route is easiest is the one people use, so the practical design question is not which route is best but which is easiest — and then making that one produce a proper record.
An intake route that requires leaving the application people are already in does not get used.
Where human review belongs
A fully automatic pipeline handles the ordinary case and mishandles the exceptions, which is where the difficult documents concentrate. An intake queue with a small amount of confirmation — accepting or correcting a proposed classification and extracted fields — usually outperforms both full automation and full manual filing.
Quality at the point of capture
Recognition quality is decided by the scan, not by the software. Resolution, contrast and skew settings belong in a documented scanning standard rather than being left to whoever is operating the device, because inconsistent capture is the main cause of unreliable recognition.
FAQ
Document capture: common questions
Should capture be centralised or distributed?
Centralised for high-volume, uniform material where a scanning standard and dedicated hardware pay off. Distributed for documents produced where the work happens — site paperwork, delivery notes, clinical consents — because central capture adds a delay that causes people to keep local copies.
Is mobile capture good enough?
For printed documents in reasonable light, usually. For anything that will be relied on legally, a proper scan is worth the extra step — and either way, indexing by a reference captured at the same time makes the document findable regardless of recognition quality.
Related terms
- Duplicate detectionDuplicate detection identifies documents whose content is already present in the repository, normally by comparing a cryptographic hash of the file.
- Email-to-folder importEmail-to-folder import monitors a mailbox and files incoming messages and their attachments into a specified folder automatically.
- Full-text searchFull-text search queries the words inside documents rather than only their names and metadata.
- Intelligent document processingIntelligent document processing combines optical character recognition with machine learning to classify a document, extract named fields from it and route it onward.
- MetadataMetadata is structured information about a document rather than inside it: its type, owner, date, status, retention class and any fields specific to its kind.
- Optical character recognitionOptical character recognition converts an image of text into machine-readable characters.
آخر مراجعة: 28 أغسطس 2026. Browse the full glossary.