Glossary
Semantic search
Also called: vector search, meaning-based search
Semantic search matches meaning rather than exact strings, using vector representations of text. It answers queries where the searcher remembers what a document was about but not the words it used, which is the most common way keyword search fails.
Semantic search explained
The failure it addresses
Keyword search fails in a specific and frustrating way: you remember the substance of a document and not its wording. A query for "the lease that renews in the third quarter" finds nothing when the agreement says "term expires 30 September", even though it is exactly the right document.
Semantic search handles that case because the two phrasings are close in meaning even though they share no keywords.
Why it should not replace exact matching
When a user searches for an invoice number, a document reference or a person's surname, they want exact matches and nothing else. Semantic matching on a precise identifier returns plausible near-misses, which is worse than returning nothing.
The right implementation returns both and presents them together, rather than making the user choose a retrieval algorithm.
Practical caveats
Semantic results are harder to explain — a user cannot see why a document matched, which undermines trust when the match is poor. And the same permission filtering applies: a vector index must respect access rights at query time, or it becomes a route to content a user cannot open.
Where it is most valuable
On large archives of narrative documents — correspondence, reports, meeting records — where the vocabulary varies between authors and no naming convention was ever enforced. On structured, well-classified content it adds less, because metadata already answers the questions people ask.
FAQ
Semantic search: common questions
Does semantic search replace metadata?
No. It improves retrieval of unstructured text; it does not turn a renewal date into a sortable field. Reporting still needs metadata.
Can it search scanned documents?
Only as well as recognition allows — semantic search operates on extracted text, so a poorly recognised scan is poorly searchable by any method.
Related terms
- Document captureDocument capture is the process of getting documents into a system and making them usable: scanning or importing, recognising text, classifying, extracting fields and applying metadata.
- Duplicate detectionDuplicate detection identifies documents whose content is already present in the repository, normally by comparing a cryptographic hash of the file.
- Email-to-folder importEmail-to-folder import monitors a mailbox and files incoming messages and their attachments into a specified folder automatically.
- Full-text searchFull-text search queries the words inside documents rather than only their names and metadata.
- Intelligent document processingIntelligent document processing combines optical character recognition with machine learning to classify a document, extract named fields from it and route it onward.
- MetadataMetadata is structured information about a document rather than inside it: its type, owner, date, status, retention class and any fields specific to its kind.
Zuletzt geprüft: 28. August 2026. Browse the full glossary.