Glossary
Optical character recognition
Also called: OCR, text recognition
Optical character recognition converts an image of text into machine-readable characters. In document management it makes scanned paper and image-only PDFs searchable by their contents, which turns an archive you can merely store into one you can actually answer questions from.
Optical character recognition explained
Why it matters more than it sounds
Scanning a filing cabinet produces image files. Without recognition, those files are findable only by whatever someone typed into the filename, which in practice means they are not findable at all. OCR is the step that converts a digitisation project from a storage exercise into a retrieval capability, and skipping it is the most common reason paperless programmes deliver nothing.
What it reads reliably, and what it does not
Printed text on a clean scan at 300 dpi recognises very well. Recognition degrades with poor contrast, skew, unusual fonts, tables with ruled lines, and photographs taken at an angle in bad light — which describes most documents captured on a phone.
Handwriting does not recognise reliably, and any vendor implying otherwise is overstating. Handwritten annotations on a delivery note or a clinical record generally stay unsearchable, so the practical approach is to index those documents by structured metadata captured at the point of scanning, rather than relying on full-text search.
Searchable PDF versus extracted text
Two outputs are commonly confused. A searchable PDF embeds a text layer behind the page image, so the file itself can be searched in a reader. Extracted text is written into a search index, which is what makes a document findable from outside the file. A document management system needs the second; the first is a convenience.
Practical implications
Run recognition on arrival rather than on demand, so a document is searchable by the time the person who uploaded it has finished filing. Expect to re-run it across an existing archive as a background task after a migration, and expect a proportion of that archive — usually the oldest and worst-scanned part — never to become reliably searchable.
FAQ
Optical character recognition: common questions
Does OCR work on handwriting?
Not reliably. Printed text recognises well; handwriting does not. For handwritten records, index by structured metadata captured at scanning time so they are findable even though their contents are not searchable.
Should we OCR our whole archive?
Only the part still being referenced. In most archives a small fraction is ever opened again, and recognising the rest costs money to produce documents nobody reads — while creating discoverable records you then have to govern.
Related terms
- Document captureDocument capture is the process of getting documents into a system and making them usable: scanning or importing, recognising text, classifying, extracting fields and applying metadata.
- Duplicate detectionDuplicate detection identifies documents whose content is already present in the repository, normally by comparing a cryptographic hash of the file.
- Email-to-folder importEmail-to-folder import monitors a mailbox and files incoming messages and their attachments into a specified folder automatically.
- Full-text searchFull-text search queries the words inside documents rather than only their names and metadata.
- Intelligent document processingIntelligent document processing combines optical character recognition with machine learning to classify a document, extract named fields from it and route it onward.
- MetadataMetadata is structured information about a document rather than inside it: its type, owner, date, status, retention class and any fields specific to its kind.
آخر مراجعة: 28 أغسطس 2026. Browse the full glossary.