OCR that makes scanned paper findable
Optical character recognition runs on uploaded images and on PDFs that contain only a scanned page image rather than a text layer. The recognised text is written into the search index and attached to the document, so a signed delivery note photographed by a driver becomes findable by the consignment number printed on it.
This is the difference between a digital archive and a searchable one. Scanning a filing cabinet into image PDFs produces something you can store but not use; running OCR over it produces something you can answer questions from. Recognition accuracy depends on the source — a clean 300 dpi scan reads far better than a photograph taken at an angle in poor light — so the honest expectation is that OCR makes most of your archive searchable, not all of it.
