Skip to content

Document processing

On Request

AIMY Extract

Paper in, structured records out.

Invoices, delivery orders, forms and handwritten dockets read into typed, validated fields your systems can accept — with a confidence score on every field and a review queue for the ones that need a person.

  • Reads scans, phone photos and handwriting, not only clean PDFs
  • Every field carries a confidence score and the region it came from
  • Low-confidence fields route to a reviewer instead of being guessed

What deployments look like

Typed fields
out, against a schema you define
not a wall of extracted text
Per field
confidence score and source region on the page
by design
A queue
for anything below your confidence threshold
configurable per document type
The problem

The document arrived. Someone still has to type it in.

Invoices from two hundred suppliers, each in its own layout. Delivery orders photographed on a phone in a loading bay. Forms filled in by hand. The information is all there, and the only thing between it and your system is a person with a keyboard and a deadline.

  • Template-based extraction breaks the first time a supplier redesigns their invoice.
  • Generic OCR returns text, not fields — somebody still has to decide what is what.
  • Typing errors surface downstream, in a payment or a stock count, long after the fact.
  • The documents needing the most care are the ones least likely to be clean scans.
Capabilities

What AIMY Extract actually does.

Layout-independent extraction

Fields are identified by what they mean rather than where they sat on the last document, so a supplier changing their template is not an incident.

Handwriting and phone photos

Skewed, creased, shadowed and handwritten inputs are first-class, because that is what actually arrives.

A schema you define

You describe the fields and their types. What comes back is a record shaped the way your ERP expects, not a transcript to be parsed later.

Confidence, field by field

Every value returns with a score and the region of the page it was read from, so a reviewer can confirm it at a glance instead of re-reading the page.

A review queue, not a guess

Anything below your threshold goes to a person. The system is allowed to be unsure; it is not allowed to be confidently wrong.

Validation before delivery

Totals that do not add up, dates outside a plausible range and identifiers that fail a checksum are flagged before the record reaches you.

How it works

From the document that arrived to the record you needed.

  1. 01

    Define the schema

    One document type at a time: the fields you need, their types, and the confidence threshold below which a person looks.

  2. 02

    Ingest

    From a watched folder, a mailbox, a scanner or an API call. Multi-page and mixed-batch inputs are handled.

  3. 03

    Extract

    Text, tables, stamps and checkboxes are read and mapped to your fields, each with a score and a source region.

  4. 04

    Review

    Low-confidence and failed-validation records surface in a queue with the page alongside, so correcting one takes seconds.

  5. 05

    Deliver

    Records are written to your system as JSON, CSV or a direct API call — and the source document is retained or discarded per your policy.

Technical profile

Deployment
Managed cloud · Customer VPC · On-premise on GB10
Accepts
PDF, JPEG, PNG, HEIC, TIFF, multi-page scans
Reads
Printed text, handwriting, tables, stamps, checkboxes
Returns
JSON or CSV, or a direct write to your ERP or DMS
Languages
English, Bahasa Malaysia and Chinese
Scoping
Per document type, delivered on enquiry

Security & data handling

  • Documents are processed inside the boundary you deploy into
  • Configurable retention, including deleting the source after delivery
  • Identity documents can be redacted once the fields have been extracted
  • Full processing log — what was read, when, and with what confidence
  • Runs air-gapped on GB10 where documents may not leave the site
  • No customer documents used to train third-party models
Questions

The ones procurement always asks.

Something not covered here? Ask directly — you will get a straight answer, including when the answer is that we are not the right fit.

Ask us
How is this different from the OCR built into our scanner?

Scanner OCR gives you searchable text. This gives you a record: named, typed fields, validated, with a confidence score and a review path for what it is unsure about. The difference is whether a person still has to read the page.

What happens with a supplier we have never seen?

It is handled the same way as any other. Extraction is not template-matched, so there is no per-supplier setup and no first-time failure — though early documents from an unfamiliar layout are more likely to route for review.

What does it do when it is not sure?

It says so. The field comes back with a low score and lands in the review queue with the source region highlighted. Silence and a plausible-looking wrong number is the failure mode we design against.

Is this part of AIMY Expert?

No. It shares the brand and the deployment options, but it is a separate application with its own scope. You can run it without AIMY Expert, and most customers do.

Next step

Send us the documents you dread.

Scoping starts with a real sample — the crumpled ones, the handwritten ones, the supplier who redesigns their invoice every year. A clean PDF proves nothing.