Skip to content

Multimodal retrieval

Enterprise
An advanced capability of AIMY Expert. Enabled on an existing deployment — there is no separate system to buy or run.

AIMY Visual

Show it the problem instead of describing it.

Upload a photograph, a screenshot or a scanned drawing straight into the chat. AIMY Visual works out what it is looking at and queries the same knowledge base from there — with the same citations, the same permissions and the same audit trail.

  • Ask with a photo when you cannot name the part
  • Reads nameplates, error screens, drawings and handwriting
  • Same index, same citations, same permission boundary

What deployments look like

One photo
replaces the part number nobody can find
the common case on a plant floor
100%
of answers still cite the source document
unchanged from AIMY Expert
Same index
no second corpus to build or maintain
by design
The problem

The person with the question often cannot describe it.

A technician standing in front of a stopped machine does not know the part is called a spindle encoder coupling. An associate holding a scanned form does not know which schedule it belongs to. Text search assumes you already have the vocabulary — which is precisely what the person asking is missing.

  • Faults are described by symptom and appearance, not in the terms the manual uses.
  • Much of the source material — drawings, nameplates, scanned forms — is not text at all.
  • Staff photograph the problem and message a colleague, because that is the only thing that works.
  • Error dialogs get retyped into a search box, badly, and the search fails.
Capabilities

What AIMY Visual actually does.

Ask with an image

Drop a photo, screenshot or scan into the same conversation. There is no separate tool and no separate upload flow — it is the same chat box.

Text in the wild

Nameplates, serial numbers, error codes on a screen, stamped part markings and handwritten log entries are read and used as retrieval terms.

Drawings and diagrams

Technical drawings, schematics and process diagrams are read structurally, so a question about a callout finds the specification behind it.

Photographed documents

A phone photo of a form or certificate is matched back to the indexed original, so the answer cites the controlled copy rather than the snapshot.

The same citation guarantee

What comes back is still an answer from your documents with the passage attached. The image changes how the question is asked, not how the answer is justified.

Images inherit the same boundary

Uploads are subject to the same permission, residency and retention rules as everything else — including discarding the image once the answer is given.

How it works

From a photograph to a cited answer.

  1. 01

    Upload

    The image arrives in the existing conversation, from a desk browser or a phone on the floor.

  2. 02

    Interpret

    The image is described and any text in it — codes, labels, handwriting — is extracted and treated as retrieval terms.

  3. 03

    Retrieve

    Those terms run through the same hybrid index, under the same access controls as any other question.

  4. 04

    Answer

    The response cites the passages it used, exactly as a typed question would, and abstains on thin evidence in exactly the same way.

  5. 05

    Discard

    Retention for uploaded images is a policy you set — including deleting the image the moment the answer has been returned.

Technical profile

Requires
An AIMY Expert deployment — managed cloud, VPC or GB10
Accepts
JPEG, PNG, HEIC, PDF pages and screenshots
Reads
Printed text, handwriting, nameplates, drawings, diagrams, charts
Retrieval
The same hybrid index — no separate corpus is built
Interfaces
Web app, mobile browser, embeddable widget
Retention
Configurable, including discard-on-answer

Security & data handling

  • Uploads are scoped by the access controls in force before any retrieval runs
  • Configurable image retention, including immediate discard after answering
  • Faces and identifiers can be redacted before the image is processed
  • Runs fully on-premise on GB10 where imagery may not leave the site
  • Every upload recorded in the same audit trail as text conversations
  • No customer imagery used to train third-party models
Questions

The ones procurement always asks.

Something not covered here? Ask directly — you will get a straight answer, including when the answer is that we are not the right fit.

Ask us
Is this a separate product we buy?

No. It is a capability enabled inside your existing AIMY Expert deployment. There is no second system to run, no second index to maintain and no second permission model to keep in step.

Does it work on handwriting?

Usually, and it is honest about when it cannot. Handwritten log entries and annotated drawings are common inputs; where a reading is uncertain, the uncertainty is surfaced rather than guessed past.

What happens to drawings covered by an NDA?

Exactly what happens to the text version: they stay inside the boundary you deployed into. On GB10 that is your own rack, with no outbound path at all.

Can it run air-gapped?

Yes. On GB10 the image interpretation runs on the same resident models as the rest of the platform, so an air-gapped site keeps the capability.

Next step

Put AIMY Visual against your own documents.

We evaluate on a sample of your own corpus, never a canned dataset. That is the only way to tell whether retrieval will hold up on your content.