Agents
Turn documents into data you can trust
The OCR Agent reads invoices, forms and scans, extracts the fields that matter, checks them against what you already know, and hands off only the pages it is unsure about.
What it does
Reads the document and returns structured fields
Scanned documents are unstructured until someone types them in. The OCR Agent does that reading for you. It locates the fields you care about — supplier, invoice number, line items, totals, tax — regardless of layout, returns them as structured data, and attaches a confidence score to each field so downstream systems know which values are safe to trust automatically.
Layout-independent
Finds fields by meaning, so a new supplier's format does not require a new template.
Line-item detail
Extracts tables of line items, not just header totals, so the full document is captured.
Per-field confidence
Scores each value so trusted fields flow through and doubtful ones are checked.
Inputs it accepts
What you feed it
Invoices and receipts
Supplier invoices, receipts and credit notes in PDF or image form.
Forms and contracts
Structured and semi-structured forms, applications and agreements.
Scans and photos
Photographs and scanned pages, including skewed or lower-quality captures.
Decisions it makes on its own
What it commits without asking
Where a field's confidence clears your threshold and validation passes, the agent commits the value. It normalises formats, such as dates and currencies, cross-checks the supplier and amounts against your records, and flags duplicates before they enter a workflow. Values below the threshold are never quietly guessed — they are held for review.
Normalises formats
Standardises dates, currencies and numbers to your conventions.
Validates against records
Confirms suppliers, purchase orders and amounts against master data.
Flags duplicates
Catches documents already seen before they create a second payment.
What escalates to a human
Where a person still reads the page
Low-confidence fields
A field the agent cannot read with confidence is shown against the source image for a quick human confirmation, not committed on a guess.
Failed validation
When an extracted value contradicts your records — an unknown supplier or a total that does not add up — the document is routed for review with the mismatch highlighted.
Systems it connects to
Where it reads and writes
Document sources
Mailboxes, shared drives and object storage where documents land.
ERPs and AP systems
Posts structured records into accounts payable and ERP workflows.
Data products
Writes extracted data to governed products for downstream use and audit.
A worked example
An invoice from a new supplier
A PDF invoice arrives from a supplier never seen before, in an unfamiliar layout. The agent extracts the header and all fourteen line items, normalises the date and currency, and cross-checks the totals — which add up. Because the supplier is unknown, it does not create a payment; it commits the extracted record, marks the supplier as new, and routes it to accounts payable to confirm the vendor before the invoice enters the approval flow.
Questions
Frequently asked
- Do we need a template per document type?
- No. The agent finds fields by meaning rather than fixed positions, so it handles new layouts without a template, though you can add hints for unusual documents.
- What accuracy should we expect?
- It depends on document quality, but the important control is the confidence threshold: you decide how sure the agent must be before a field is accepted automatically, and the rest is reviewed.
- Can it run on sensitive documents?
- Yes. It runs inside your deployment, including on-premise and air-gapped, so documents never have to leave your environment to be read.
Extract from your own documents
Send a batch of real invoices or forms and we will show you the structured output, the confidence scores and what the agent chooses to escalate.

