Skip to main content
DocumentMS

Search & metadata

OCR document search across scans, PDFs and metadata

OCR document search makes scanned paper findable by its contents. DocumentMS runs optical character recognition over uploaded images and image-only PDFs, writes the recognised text into the search index, and queries it alongside custom metadata fields so one search covers both structured and unstructured content.

Who this is for

Search is the module people notice first, because failing to find a document is the failure they experience most often.
  • Teams whose archive is largely scanned paper and therefore currently unsearchable
  • Anyone answering subject access, disclosure or audit requests against years of correspondence
  • Finance and procurement teams who need to find an invoice by supplier and amount rather than by filename
  • Organisations where the same document is described differently by different departments

Capabilities

OCR that makes scanned paper findable

Optical character recognition runs on uploaded images and on PDFs that contain only a scanned page image rather than a text layer. The recognised text is written into the search index and attached to the document, so a signed delivery note photographed by a driver becomes findable by the consignment number printed on it.

This is the difference between a digital archive and a searchable one. Scanning a filing cabinet into image PDFs produces something you can store but not use; running OCR over it produces something you can answer questions from. Recognition accuracy depends on the source — a clean 300 dpi scan reads far better than a photograph taken at an angle in poor light — so the honest expectation is that OCR makes most of your archive searchable, not all of it.

One search across contents and metadata

Global search queries the extracted text and the structured metadata together. That matters because the useful question is almost never purely one or the other: "the supplier agreement with a renewal date this quarter that mentions indexation" needs the date from a metadata field and the word from the document body, and running two searches and intersecting them by hand is how people give up.

Results are permission-filtered before they are returned. A user never sees a result they are not entitled to open, and the search itself does not leak the existence of documents through result counts.

Custom metadata that reflects your record types

Metadata fields are defined per document type, so a contract carries counterparty, value, term and renewal date while a policy carries owner, review date and approval reference. Mandatory fields can be enforced at upload, which is the only reliable way to keep a metadata model from decaying into blanks.

Where AI extraction is enabled, those fields are populated from the document content and presented for confirmation rather than silently accepted, so the metadata stays trustworthy enough to report from.

  • Field types for text, number, date, currency, single-select and multi-select
  • Mandatory fields enforced at the point of upload
  • Metadata filters that combine with free-text search
  • Links between related files and folders, so an amendment sits with its parent agreement

Filters that narrow rather than restart

Filters apply on top of a query instead of replacing it: category, folder, document type, owner, date range, approval status, retention class and any custom metadata field. The behaviour people expect from an online shop is the behaviour that works here too — refine, look, refine again.

A filtered search can be saved and shared, which turns an ad-hoc query into a working view for a team. It is also how disclosure exercises stay reproducible: the same saved search run a month later returns the same set plus anything new that matches.

Semantic search for when you cannot remember the words

Keyword search fails in a specific and frustrating way: you remember what a document was about but not what it said. Semantic search matches meaning, so a query for "the lease that renews in the third quarter" can return the right agreement even though it uses the phrase "term expires 30 September".

Semantic and keyword results are presented together rather than as separate modes, because forcing a user to choose a search algorithm is asking them to solve our problem.

In the product

What this looks like in use

DocumentMS search results showing metadata filters applied, matched text from a scanned PDF, and the document version and status for each hit
DocumentMS search results showing metadata filters applied, matched text from a scanned PDF, and the document version and status for each hit
client verification — interface wireframe. Replace with a capture of real OCR search results.

How it works

How OCR search is built

Indexing happens on arrival, so a document is searchable by the time the person who uploaded it has finished filing it.
  1. Step 1: Detect

    On upload, DocumentMS checks whether the file already contains a text layer. Native documents skip OCR entirely; images and image-only PDFs are queued for recognition.

  2. Step 2: Recognise

    OCR extracts the text, preserving page boundaries so a search result can point at the page it matched rather than the whole document.

  3. Step 3: Enrich

    Custom metadata is applied, and where AI extraction is enabled, named fields such as supplier or renewal date are proposed from the recognised text for confirmation.

  4. Step 4: Index

    Text and metadata are written to the search index together, with the document’s permissions attached so results can be filtered per user at query time.

Specifications

Technical specifications

The numbers a technical evaluation asks for, stated rather than described. Where a limit is configurable, the default and the ceiling are both given.
Search and OCR specifications
PropertyValue
OCR inputsUploaded images and image-only PDFs
Text layer detectionAutomatic; native documents bypass OCR
Page granularityMatches resolve to the page within a document
Search scopeExtracted text, custom metadata, filename, document number
FiltersCategory, folder, type, owner, date range, status, retention class, custom fields
Semantic searchMeaning-based matching presented alongside keyword results
Permission filteringApplied at query time; results never include documents the user cannot open
Saved searchesShareable within a team, reproducible for disclosure exercises
OCR languagesOver 100 recognition languages including Arabic, Chinese, Japanese and Cyrillic scripts, with mixed-language pages handled in a single pass
Indexing latencyTypically under 30 seconds from upload to searchable for an office document, and under 2 minutes for a scanned multi-page PDF requiring OCR

Security

Security notes

Search is the module most likely to leak information by accident, because a result list is itself a disclosure.

Read the trust centre

  • Results are filtered by the requesting user’s permissions before they are returned, not hidden in the interface afterwards
  • Result counts do not reveal the existence of documents a user cannot access
  • Search queries are recorded in the audit trail, which is how you detect someone systematically looking for records outside their remit
  • OCR text is stored with the same encryption and permissions as the document it came from

FAQ

OCR & full-text search: common questions

Answers to what procurement, IT and compliance teams ask us about this module.
Does OCR work on handwriting?

Printed text recognises reliably; handwriting does not, and we would rather say so than let you discover it after a scanning project. For handwritten records the practical approach is to index them by structured metadata — date, author, reference, case number — captured at the point of scanning, so they are findable even though their contents are not searchable.

Will OCR run over documents we have already uploaded?

Yes. Recognition can be run retrospectively across an existing archive, which is normally how a migration finishes: bring the files in first so people can start working, then process the backlog in the background.

How is this different from search in SharePoint or Google Drive?

Both index document contents competently. The difference is what search returns alongside the hit: DocumentMS shows the version, approval status, retention class and owner, and filters on enforced metadata fields rather than optional ones. That turns a search into an answer about whether the document may be relied on.

Can we search inside documents we do not have permission to open?

No. Permissions are applied at query time, so an unauthorised document does not appear in results, does not affect result counts and cannot be inferred from the interface.

What happens when OCR gets a word wrong?

The document remains findable by its metadata, filename and document number, and the extracted text can be corrected. Recognition errors are a reason to enforce a small number of mandatory metadata fields rather than relying on full-text search alone.

A 30-minute session with a solutions engineer, using a folder structure and approval chain that resemble yours — not a generic demonstration tenant.