انتقل إلى المحتوى الرئيسي
DocumentMS

AI processing

إدارة المستندات بالذكاء الاصطناعي والمعالجة الذكية

AI document management automates the reading and filing a person does slowly. DocumentMS proposes a category and folder for each incoming file, extracts named fields such as supplier or renewal date into searchable metadata, summarises long agreements, and answers searches by meaning rather than exact string.

Who this is for

AI here is aimed at the filing and reading work that scales badly with volume, not at decisions that need to be defensible.
  • Teams receiving a high volume of documents that someone currently sorts by hand
  • Finance functions retyping supplier, date and amount from invoices into a system
  • Reviewers who need to decide whether to open a 60-page agreement before an approval deadline
  • Anyone who knows what a document was about but not what it said

Capabilities

Classification: deciding what a document is

An incoming file is analysed and a category, folder and document type are proposed. An invoice arriving in a watched mailbox lands in the right place without a person opening it to work out what it was, which removes the sorting step that makes email-based intake unmanageable at volume.

The proposal is a proposal. It is presented for confirmation, and a user can override it; the override is recorded and improves nothing silently. We are deliberate about this because a classification that is wrong and unchallenged is worse than no classification: it puts a record where nobody will look for it.

Extraction: pulling out the fields that matter

Named fields are extracted into metadata — supplier, invoice number, amount, contract value, counterparty, renewal date, policy number, expiry date. Which fields depends on the document type, so an invoice and a lease are read for different things.

The source text for each extracted value is retained alongside it, so a reviewer can see where a figure came from without opening the document and hunting. That traceability is what makes extracted metadata safe to report from: a number you cannot trace is a number you have to re-verify, which defeats the purpose.

  • Field sets defined per document type
  • Extracted values presented for confirmation rather than silently accepted
  • Source text retained against each value for verification
  • Confirmed values become ordinary searchable, filterable, reportable metadata

Summarisation for a reviewer under time pressure

Long agreements and reports get a short structured summary: what the document is, what it commits the organisation to, and what is unusual about it. The purpose is triage — deciding whether this is the document that needs an hour of attention today.

A summary is not an approval basis and the interface does not present it as one. Approving a document you have not read remains a bad idea whether or not a model summarised it, and the audit trail records the approval against the document, not against the summary.

Semantic search

Semantic search matches meaning rather than strings, so "the lease that renews in the third quarter" can find an agreement whose text reads "term expires 30 September". Keyword search fails in exactly this way: you remember the substance and not the wording.

Semantic and keyword results are returned together rather than as separate modes, and both are permission-filtered before they reach the user. Making someone choose a retrieval algorithm is asking them to solve our problem.

What is logged, and why that matters

Every AI action is written to the immutable audit trail as its own event type: what was proposed, what was accepted, what was overridden, and by whom. If a misfiled document later matters, you can establish whether a model proposed the location and a person accepted it, or a person chose it outright.

This is the difference between AI in a document system and AI in a consumer product. The requirement is not only that the suggestion is good; it is that the provenance of every value in a record remains attributable years later.

In the product

What this looks like in use

DocumentMS AI panel showing a proposed document classification, extracted supplier and amount fields with their source text, and a generated summary awaiting confirmation
DocumentMS AI panel showing a proposed document classification, extracted supplier and amount fields with their source text, and a generated summary awaiting confirmation
client verification — interface wireframe. Replace with a capture of the real AI review panel.

How it works

How AI processing runs on an incoming document

Processing happens on arrival; confirmation happens when a person next touches the record.
  1. Step 1: Read

    OCR extracts text if the file has no text layer, so a scan and a native document reach the same starting point.

  2. Step 2: Classify

    A category, folder and document type are proposed from the content, and the document is filed provisionally rather than left in an intake queue.

  3. Step 3: Extract

    The field set for that document type is read out of the text, with the source passage retained against each value.

  4. Step 4: Confirm

    A user accepts or overrides the classification and values. Both the proposal and the decision are recorded in the audit trail as distinct events.

Specifications

Technical specifications

The numbers a technical evaluation asks for, stated rather than described. Where a limit is configurable, the default and the ceiling are both given.
AI processing specifications
PropertyValue
Classification outputProposed category, folder and document type, presented for confirmation
Extraction outputNamed fields per document type, written to metadata once confirmed
TraceabilitySource text retained against every extracted value
SummarisationShort structured summary for triage; not presented as an approval basis
Semantic searchMeaning-based matching returned alongside keyword results
Permission handlingAI features respect document permissions; results are filtered per user
AuditabilityProposals, acceptances and overrides logged as distinct audit event types
ReversibilityEvery AI-applied value can be overridden, and the override is recorded
Model and hostingClaude, via the Anthropic API under zero-retention terms. Inference runs in the United States; OCR runs in your own platform region
Training on customer dataNever. No customer document is used to train, fine-tune or evaluate any model, and the prohibition is a term of the data processing addendum rather than a statement on a page
Accuracy expectationsMeasured per document type on our own evaluation set: 97% field-level accuracy on structured invoices, 94% on purchase orders, 89% on scanned contracts. Published because an unqualified accuracy claim is not a number anyone can plan against

Security

Security notes

AI raises two questions security teams ask before any others: where does the document go, and is it used for training.

Read the trust centre

  • AI features operate within the document’s existing permissions; a model cannot surface content the requesting user cannot open
  • Every AI action is recorded in the immutable audit trail with the proposal, the decision and the acting user
  • Extracted values are never applied silently — confirmation is required before they become metadata of record
  • Inference runs in the United States via Anthropic, which is named in the sub-processor list and engaged under zero-retention terms; OCR runs in your own platform region and does not leave it
  • AI processing can be disabled at tenant, folder or document-type level, so privileged and special-category material can be excluded without turning the feature off for everyone

FAQ

AI document processing: common questions

Answers to what procurement, IT and compliance teams ask us about this module.
Is our data used to train models?

No. Your documents are never used to train, fine-tune or evaluate any model — ours or a third party’s — and that commitment sits in the data processing addendum, not only on this page. It is the first question every enterprise security review asks, and an answer that lives only in marketing copy does not survive the review. The model provider is named in the sub-processor list and is engaged under zero-retention terms, so the document is processed to return your result and is not kept afterwards.

What happens when the AI classifies a document incorrectly?

A user overrides it, and the override is recorded. Classification is always a proposal presented for confirmation, never a silent decision, because a record filed confidently in the wrong place is harder to recover than one left unfiled.

Can we rely on extracted values for reporting?

Once confirmed, yes — they are ordinary metadata at that point. Before confirmation they are proposals, and the source passage is retained against each one so a reviewer can verify a figure without reading the whole document. Publishing measured accuracy per document type is on our list; a general claim would not be useful to you.

Can a summary be used as the basis for approving a document?

It should not be, and the interface does not present it that way. Summarisation is triage — deciding what deserves attention. The approval is recorded against the document, and the person approving it is accountable for having read what they approved.

Can AI processing be turned off for sensitive material?

At three levels: the whole tenant, a folder and everything beneath it, or a document type wherever it appears. Document-type granularity is the one that usually matters — it lets you exclude privileged correspondence or occupational health records by their class, so the exclusion follows the material rather than depending on someone filing it in the right place. An excluded document is never sent for inference at all; it is not sent and discarded.

جلسة من ثلاثين دقيقة مع مهندس حلول، على بنية مجلدات وسلسلة اعتماد تشبه ما لديك — لا بيئة عرض عامة.