Glossary
Intelligent document processing
Also called: IDP, document AI
Intelligent document processing combines optical character recognition with machine learning to classify a document, extract named fields from it and route it onward. Where OCR turns an image into text, IDP turns that text into structured data a system can act on.
Intelligent document processing explained
The distinction from OCR
Recognition produces characters. It cannot tell you that the document is an invoice, that the supplier is a particular company, or that the total is a specific amount — those require understanding the layout and the semantics, which is what the "intelligent" part adds.
The practical consequence is that OCR makes a document findable and IDP makes it actionable. An invoice that has only been recognised still needs someone to read it and type the amount into a finance system; one that has been processed arrives with the amount as a field.
What it does reliably
Classification into a known set of document types works well when the types are visually distinct and there are enough examples. Extraction of clearly labelled fields — invoice number, date, total, supplier — works well on printed documents in familiar layouts.
It degrades on unusual layouts, on handwritten annotations, and on documents where the field you want is implied rather than stated. A renewal date expressed as "this agreement shall continue for a period of thirty-six months from the commencement date" is a calculation, not an extraction.
Why confirmation matters
An extracted value that nobody verified is worse than one nobody extracted, because it will be relied on. The defensible pattern is to propose the value, retain the source passage it came from, and require confirmation before it becomes metadata of record — so a reviewer can check a figure in seconds without opening the document.
FAQ
Intelligent document processing: common questions
How accurate is field extraction?
Good enough to propose, not good enough to accept silently. Accuracy varies enormously by document type and layout quality, which is why a general percentage is not a useful claim and per-type measurement is.
Does IDP replace data entry staff?
It removes the retyping, not the judgement. Someone still confirms the extracted values and handles exceptions, and the exceptions are where the difficult cases concentrate.
Related terms
- Document captureDocument capture is the process of getting documents into a system and making them usable: scanning or importing, recognising text, classifying, extracting fields and applying metadata.
- Duplicate detectionDuplicate detection identifies documents whose content is already present in the repository, normally by comparing a cryptographic hash of the file.
- Email-to-folder importEmail-to-folder import monitors a mailbox and files incoming messages and their attachments into a specified folder automatically.
- Full-text searchFull-text search queries the words inside documents rather than only their names and metadata.
- MetadataMetadata is structured information about a document rather than inside it: its type, owner, date, status, retention class and any fields specific to its kind.
- Optical character recognitionOptical character recognition converts an image of text into machine-readable characters.
最終レビュー: 2026年8月28日. Browse the full glossary.