Operations
Making Document Search Actually Work
Document search disappoints mainly because the attributes people search by were never recorded anywhere. Full-text indexing finds words inside documents, and it cannot answer questions about document type, status, owner or effective date unless each of those exists as a field.
Making Document Search Actually Work
Search is the feature people evaluate least carefully and complain about most. The complaint is almost always the same: I know the document exists and I cannot find it.
The cause is almost never the search engine.
What full-text search can and cannot do
Full-text indexing finds words that appear inside documents. That is genuinely useful and it is narrower than people assume.
It cannot find all contracts expiring in the next ninety days, because expiry is not a word in the document. It cannot find the current approved version of a procedure, because approval status is not text. It cannot find everything owned by a department, or everything of a given document type, or everything subject to a particular retention class.
Every one of those is a real query, and none of them is a text query. They are attribute queries, and they only work if the attribute exists as a field.
Half the archive may have no text at all
A scanned document with no recognition applied is an image. Its content is invisible to any index, and the failure is silent — the search returns results, just not that one.
Two checks are worth running on any existing repository. What proportion of documents have a text layer, and what the recognition quality is on the ones that do. Poor-quality recognition is worse than none, because it produces confident matches on the wrong documents and near-misses on the right ones.
A structural fix matters more than a retrospective one here: recognise at the point of capture, with a documented scanning standard, so the problem stops growing while you deal with the backlog.
Give people a reference to search by
Recognition is imperfect and always will be for handwriting, poor originals and unusual layouts.
The reliable answer is not better recognition. It is a reference captured at intake — an invoice number, a contract number, a case reference, a purchase order — recorded as a field. A document with an indexed reference is findable regardless of how legible its content is, and references are usually what people actually have in front of them when they search.
Four changes, in order of return
Record document type on everything. It is the field nearly every real query filters on, and it simultaneously enables retention and permission rules. Documents with no type are the ones that disappear.
Recognise text at capture. With a scanning standard, so quality is consistent rather than dependent on whoever operated the device.
Capture one reference per document type. The number the person searching will already know.
Expose filters, not just a box. Most retrieval is a narrowing task rather than a lookup: this type, this status, this owner, this date range. A single text box asks people to guess vocabulary; filters let them describe what they know.
Search results have to show enough to choose
A results list of filenames forces people to open documents to identify them, which is slow enough that they stop trusting search and go back to asking colleagues.
Show the attributes that distinguish candidates: type, status, version, owner, date, and where it sits. Most of the time, choosing correctly from the list is the whole task.
Saved views beat repeated searches
A large share of searching is the same query run repeatedly — contracts expiring soon, documents past their review date, items awaiting my approval, records with no retention class.
Those should be saved views rather than remembered searches, because a view is always current, shareable, and reviewable. They also double as governance reports: the view listing documents with no retention class is the most useful completeness measure a records programme has.
The diagnostic worth running
Ask five people to find five documents each, and watch. You will learn which attributes they searched by, which they expected to exist, and where they gave up.
That produces a short, specific list of fields to add — which is a more reliable basis for improving retrieval than any evaluation of search technology.
It is also worth recording what people searched for and found nothing. A log of zero-result queries is the cheapest ongoing source of improvement available: it names the vocabulary people use, which is frequently not the vocabulary the taxonomy uses, and it identifies the document classes that exist in the organisation but not in the system. Both are fixable, and neither is visible from inside the system's own structure.
FAQ
Questions this raises
Why does search miss documents that definitely exist?
Usually because the document is a scanned image with no recognised text layer, or because the search term is an attribute rather than a word in the content. Both are capture problems rather than search problems.
Is better search software the answer?
Rarely on its own. If the attributes people search by were never recorded and half the archive has no text layer, a better index searches the same missing information faster.
What is the single most useful field to add?
Document type, because it is the field almost every real query filters on and it makes retention and permission rules possible at the same time.
About the author
Written and reviewed by the DocumentMS product and compliance team.
Dernière revue: 27 août 2026