Glossary
Full-text search
Also called: content search, text search
Full-text search queries the words inside documents rather than only their names and metadata. In a document system it depends on text extraction — including OCR for scanned material — and it must be filtered by the requesting user's permissions at query time.
Full-text search explained
What it depends on
An index of extracted text. For native documents that extraction is straightforward; for scanned images and image-only PDFs it depends on recognition, which is why OCR quality determines search quality for any archive that began on paper.
Why it is insufficient alone
Full-text search finds documents that mention a word. It cannot answer "every contract expiring next quarter", because the renewal date exists as a sentence rather than as a sortable value. It also cannot tell you whether the document it found is the approved version.
Records need metadata for the questions that matter and full text for the questions you did not anticipate. Systems that offer only one of the two are frustrating in predictable ways.
Permission filtering
Results must be filtered by the requesting user's permissions before they are returned — not hidden in the interface afterwards. Result counts must not reveal the existence of documents the user cannot open, because a count is a disclosure.
This is easy to get wrong and worth testing explicitly during an evaluation: search for a term you know appears only in a restricted document, as a user without access.
Stemming and phrase matching
A useful index matches "retained" against a search for "retention", and treats a quoted phrase as exact. Both behaviours are expected and their absence is noticed immediately by anyone who has used a search engine.
FAQ
Full-text search: common questions
Should search cover document contents or metadata?
Both, in one query. Forcing a user to choose which index to search is asking them to solve the system's problem, and the useful question is almost never purely one or the other.
How current should the index be?
A document should be searchable by the time the person who uploaded it has finished filing. Indexing latency measured in hours produces users who do not trust search.
Related terms
- Document captureDocument capture is the process of getting documents into a system and making them usable: scanning or importing, recognising text, classifying, extracting fields and applying metadata.
- Duplicate detectionDuplicate detection identifies documents whose content is already present in the repository, normally by comparing a cryptographic hash of the file.
- Email-to-folder importEmail-to-folder import monitors a mailbox and files incoming messages and their attachments into a specified folder automatically.
- Intelligent document processingIntelligent document processing combines optical character recognition with machine learning to classify a document, extract named fields from it and route it onward.
- MetadataMetadata is structured information about a document rather than inside it: its type, owner, date, status, retention class and any fields specific to its kind.
- Optical character recognitionOptical character recognition converts an image of text into machine-readable characters.
آخر مراجعة: 28 أغسطس 2026. Browse the full glossary.