Document processing
AIMY Extract
Paper in, structured records out.
Invoices, delivery orders, forms and handwritten dockets read into typed, validated fields your systems can accept — with a confidence score on every field and a review queue for the ones that need a person.
- Reads scans, phone photos and handwriting, not only clean PDFs
- Every field carries a confidence score and the region it came from
- Low-confidence fields route to a reviewer instead of being guessed
What deployments look like
- Typed fields
- out, against a schema you define
- not a wall of extracted text
- Per field
- confidence score and source region on the page
- by design
- A queue
- for anything below your confidence threshold
- configurable per document type
The document arrived. Someone still has to type it in.
Invoices from two hundred suppliers, each in its own layout. Delivery orders photographed on a phone in a loading bay. Forms filled in by hand. The information is all there, and the only thing between it and your system is a person with a keyboard and a deadline.
- Template-based extraction breaks the first time a supplier redesigns their invoice.
- Generic OCR returns text, not fields — somebody still has to decide what is what.
- Typing errors surface downstream, in a payment or a stock count, long after the fact.
- The documents needing the most care are the ones least likely to be clean scans.
What AIMY Extract actually does.
Layout-independent extraction
Fields are identified by what they mean rather than where they sat on the last document, so a supplier changing their template is not an incident.
Handwriting and phone photos
Skewed, creased, shadowed and handwritten inputs are first-class, because that is what actually arrives.
A schema you define
You describe the fields and their types. What comes back is a record shaped the way your ERP expects, not a transcript to be parsed later.
Confidence, field by field
Every value returns with a score and the region of the page it was read from, so a reviewer can confirm it at a glance instead of re-reading the page.
A review queue, not a guess
Anything below your threshold goes to a person. The system is allowed to be unsure; it is not allowed to be confidently wrong.
Validation before delivery
Totals that do not add up, dates outside a plausible range and identifiers that fail a checksum are flagged before the record reaches you.
From the document that arrived to the record you needed.
-
01
Define the schema
One document type at a time: the fields you need, their types, and the confidence threshold below which a person looks.
-
02
Ingest
From a watched folder, a mailbox, a scanner or an API call. Multi-page and mixed-batch inputs are handled.
-
03
Extract
Text, tables, stamps and checkboxes are read and mapped to your fields, each with a score and a source region.
-
04
Review
Low-confidence and failed-validation records surface in a queue with the page alongside, so correcting one takes seconds.
-
05
Deliver
Records are written to your system as JSON, CSV or a direct API call — and the source document is retained or discarded per your policy.
Technical profile
- Deployment
- Managed cloud · Customer VPC · On-premise on GB10
- Accepts
- PDF, JPEG, PNG, HEIC, TIFF, multi-page scans
- Reads
- Printed text, handwriting, tables, stamps, checkboxes
- Returns
- JSON or CSV, or a direct write to your ERP or DMS
- Languages
- English, Bahasa Malaysia and Chinese
- Scoping
- Per document type, delivered on enquiry
Security & data handling
- Documents are processed inside the boundary you deploy into
- Configurable retention, including deleting the source after delivery
- Identity documents can be redacted once the fields have been extracted
- Full processing log — what was read, when, and with what confidence
- Runs air-gapped on GB10 where documents may not leave the site
- No customer documents used to train third-party models
The ones procurement always asks.
Something not covered here? Ask directly — you will get a straight answer, including when the answer is that we are not the right fit.
Ask usHow is this different from the OCR built into our scanner?
Scanner OCR gives you searchable text. This gives you a record: named, typed fields, validated, with a confidence score and a review path for what it is unsure about. The difference is whether a person still has to read the page.
What happens with a supplier we have never seen?
It is handled the same way as any other. Extraction is not template-matched, so there is no per-supplier setup and no first-time failure — though early documents from an unfamiliar layout are more likely to route for review.
What does it do when it is not sure?
It says so. The field comes back with a low score and lands in the review queue with the source region highlighted. Silence and a plausible-looking wrong number is the failure mode we design against.
Is this part of AIMY Expert?
No. It shares the brand and the deployment options, but it is a separate application with its own scope. You can run it without AIMY Expert, and most customers do.
Sectors where AIMY Extract is deployed
Also on the AIMY Platform
Agentic process automation
AIMY Agents
Multi-step work, executed inside boundaries you can defend.
Explore AIMY AgentsImage generation
AIMY Avatar
One photograph, a consistent set of portraits.
Explore AIMY AvatarEvent access
AIMY Smart Face
Register once, then walk in.
Explore AIMY Smart FaceSend us the documents you dread.
Scoping starts with a real sample — the crumpled ones, the handwritten ones, the supplier who redesigns their invoice every year. A clean PDF proves nothing.