Skip to content

Assurance & evidence

Enterprise
An advanced capability of AIMY Expert. Enabled on an existing deployment — there is no separate system to buy or run.

AIMY Guard

Proof that it is still answering correctly.

Continuous evaluation, drift detection and audit-ready reporting over your AIMY Expert deployment — so "is it still answering correctly?" has a number behind it rather than an opinion.

  • Answer quality tracked as a release metric, not a feeling
  • Regressions caught before your users find them
  • Evidence packs an auditor can actually read

What deployments look like

Every change
evaluated before it reaches users
a gate, not a report
Per department
quality tracked separately, because corpora differ
by design
Generated
audit evidence, rather than assembled by hand
from the run history
The problem

The pilot was measured. The production system is not.

Answer quality was proven once, during evaluation, against a corpus that has changed a hundred times since. Documents were superseded, permissions moved, a model was updated. None of it is being measured, and the first sign of a problem will be a user who quietly stops trusting the answers.

  • Corpora drift constantly, and retrieval quality drifts with them.
  • Prompt and model changes ship on the strength of a few spot checks.
  • Nobody can say whether the assistant is better or worse than last quarter.
  • Audit and model risk ask for evidence, and someone assembles it by hand.
Capabilities

What AIMY Guard actually does.

Living evaluation sets

The questions your subject-matter experts actually care about, kept as a versioned suite and run continuously — not once, at the pilot.

Groundedness regression

Each configuration change is scored for citation accuracy, retrieval recall and refusal behaviour before it reaches a single user.

Drift detection

Alerts when quality on a corpus falls, so a stale, moved or broken source is found by the system rather than reported by a frustrated user.

Per-department reporting

Quality is reported per corpus and per audience, because a healthy average routinely conceals one department being badly served.

Evidence packs

Model risk, audit and board reporting generated from the run history instead of assembled the week before a review.

Cost and usage visibility

Where the spend goes, which departments drive it, and which query patterns are expensive relative to the value they return.

How it works

From your experts' questions to a gate on every change.

  1. 01

    Capture the questions

    Your experts' questions, the ones from the original scoping, and real queries that went badly all become part of one versioned evaluation set.

  2. 02

    Score continuously

    The suite runs on a schedule and on every change — to the corpus, the prompts, the retrieval configuration or the model.

  3. 03

    Gate the change

    A change that regresses groundedness or recall beyond your threshold does not ship. The gate is the point; the report is a by-product.

  4. 04

    Alert on drift

    Between changes, falling quality on a corpus raises an alert with the specific questions that started failing.

  5. 05

    Produce the evidence

    Run history becomes the audit pack — what was tested, when, against which version, with what result.

Technical profile

Requires
An AIMY Expert deployment — managed cloud, VPC or GB10
Measures
Groundedness, citation accuracy, retrieval recall, refusal behaviour, latency
Cadence
On every configuration change, plus a schedule you set
Scope
Per corpus and per department, not a single global score
Reporting
Dashboards plus exportable evidence packs
Deployment
Managed cloud · Customer VPC · On-premise on GB10

Security & data handling

  • Evaluation runs inside the same boundary as the deployment it measures
  • Evaluation sets and results are customer data, never shared or pooled
  • Reports respect the same permission model as the corpora they cover
  • Evidence packs are exportable in full, with no vendor dependency to read them
  • Runs entirely on-premise on GB10, including the evaluation history
  • Immutable run history suitable for model risk and audit review
Questions

The ones procurement always asks.

Something not covered here? Ask directly — you will get a straight answer, including when the answer is that we are not the right fit.

Ask us
Is this a separate product we buy?

No. It is a capability on your AIMY Expert deployment. It measures that deployment specifically — the same corpora, the same permissions, the same departments.

Who writes the evaluation questions?

Your subject-matter experts, with us facilitating. Questions written by the people who will judge the answers are the only ones worth measuring against; a generic benchmark tells you nothing about your corpus.

What happens when a change fails the gate?

It does not ship, and you get the specific questions that regressed along with the retrieved evidence for each. Failures are diagnostic, not just a red light on a dashboard.

Can it evaluate systems we did not build with you?

That is an advisory engagement rather than this capability. AIMY Guard is built into the AIMY platform; assessing someone else's AI system is work we do, but it is scoped separately.

Next step

Put AIMY Guard against your own documents.

We evaluate on a sample of your own corpus, never a canned dataset. That is the only way to tell whether retrieval will hold up on your content.