Glossary
Duplicate detection
Also called: deduplication
Duplicate detection identifies documents whose content is already present in the repository, normally by comparing a cryptographic hash of the file. Filename comparison misses the common case entirely, because the same document has usually been renamed somewhere between the two copies.
Duplicate detection explained
Why hashing rather than filenames
The same document arrives twice as "invoice.pdf" and "INV-4471 scan.pdf". A filename comparison finds nothing; a hash comparison finds an exact match regardless of what it has been called. Hashing also catches the case where a file is re-uploaded years later by someone who did not know it was already there.
What hashing cannot catch is the same document arriving as a scan and as a native PDF — different bytes, same content. That case needs metadata matching, typically supplier plus document number.
The real reason it matters
Storage cost is the obvious argument and the less important one. The document control problem is that a duplicate is a second copy with an independent life: it does not receive the new version, it is not covered by the retention rule applied to the original, and it circulates as current after the original has been superseded.
A duplicate report is therefore a document control report, not a housekeeping one.
Flag rather than block
There are legitimate reasons to hold the same bytes twice — a contract filed under both parties, a certificate relevant to two assets. Blocking the upload forces a workaround; flagging it and recording the decision preserves the information.
FAQ
Duplicate detection: common questions
Does duplicate detection slow uploads?
Hashing a file is fast relative to transferring and indexing it, so the added cost is negligible. Comparison is an index lookup.
What is the better answer than a duplicate?
A link between the two folders, so both routes lead to one record with one version history. Two copies will diverge; one linked record cannot.
Related terms
- Document captureDocument capture is the process of getting documents into a system and making them usable: scanning or importing, recognising text, classifying, extracting fields and applying metadata.
- Email-to-folder importEmail-to-folder import monitors a mailbox and files incoming messages and their attachments into a specified folder automatically.
- Full-text searchFull-text search queries the words inside documents rather than only their names and metadata.
- Intelligent document processingIntelligent document processing combines optical character recognition with machine learning to classify a document, extract named fields from it and route it onward.
- MetadataMetadata is structured information about a document rather than inside it: its type, owner, date, status, retention class and any fields specific to its kind.
- Optical character recognitionOptical character recognition converts an image of text into machine-readable characters.
Last reviewed: 28 August 2026. Browse the full glossary.