Assurance & evidence
AIMY Guard
Proof that it is still answering correctly.
Continuous evaluation, drift detection and audit-ready reporting over your AIMY Expert deployment — so "is it still answering correctly?" has a number behind it rather than an opinion.
- Answer quality tracked as a release metric, not a feeling
- Regressions caught before your users find them
- Evidence packs an auditor can actually read
What deployments look like
- Every change
- evaluated before it reaches users
- a gate, not a report
- Per department
- quality tracked separately, because corpora differ
- by design
- Generated
- audit evidence, rather than assembled by hand
- from the run history
The pilot was measured. The production system is not.
Answer quality was proven once, during evaluation, against a corpus that has changed a hundred times since. Documents were superseded, permissions moved, a model was updated. None of it is being measured, and the first sign of a problem will be a user who quietly stops trusting the answers.
- Corpora drift constantly, and retrieval quality drifts with them.
- Prompt and model changes ship on the strength of a few spot checks.
- Nobody can say whether the assistant is better or worse than last quarter.
- Audit and model risk ask for evidence, and someone assembles it by hand.
What AIMY Guard actually does.
Living evaluation sets
The questions your subject-matter experts actually care about, kept as a versioned suite and run continuously — not once, at the pilot.
Groundedness regression
Each configuration change is scored for citation accuracy, retrieval recall and refusal behaviour before it reaches a single user.
Drift detection
Alerts when quality on a corpus falls, so a stale, moved or broken source is found by the system rather than reported by a frustrated user.
Per-department reporting
Quality is reported per corpus and per audience, because a healthy average routinely conceals one department being badly served.
Evidence packs
Model risk, audit and board reporting generated from the run history instead of assembled the week before a review.
Cost and usage visibility
Where the spend goes, which departments drive it, and which query patterns are expensive relative to the value they return.
From your experts' questions to a gate on every change.
-
01
Capture the questions
Your experts' questions, the ones from the original scoping, and real queries that went badly all become part of one versioned evaluation set.
-
02
Score continuously
The suite runs on a schedule and on every change — to the corpus, the prompts, the retrieval configuration or the model.
-
03
Gate the change
A change that regresses groundedness or recall beyond your threshold does not ship. The gate is the point; the report is a by-product.
-
04
Alert on drift
Between changes, falling quality on a corpus raises an alert with the specific questions that started failing.
-
05
Produce the evidence
Run history becomes the audit pack — what was tested, when, against which version, with what result.
Technical profile
- Requires
- An AIMY Expert deployment — managed cloud, VPC or GB10
- Measures
- Groundedness, citation accuracy, retrieval recall, refusal behaviour, latency
- Cadence
- On every configuration change, plus a schedule you set
- Scope
- Per corpus and per department, not a single global score
- Reporting
- Dashboards plus exportable evidence packs
- Deployment
- Managed cloud · Customer VPC · On-premise on GB10
Security & data handling
- Evaluation runs inside the same boundary as the deployment it measures
- Evaluation sets and results are customer data, never shared or pooled
- Reports respect the same permission model as the corpora they cover
- Evidence packs are exportable in full, with no vendor dependency to read them
- Runs entirely on-premise on GB10, including the evaluation history
- Immutable run history suitable for model risk and audit review
The ones procurement always asks.
Something not covered here? Ask directly — you will get a straight answer, including when the answer is that we are not the right fit.
Ask usIs this a separate product we buy?
No. It is a capability on your AIMY Expert deployment. It measures that deployment specifically — the same corpora, the same permissions, the same departments.
Who writes the evaluation questions?
Your subject-matter experts, with us facilitating. Questions written by the people who will judge the answers are the only ones worth measuring against; a generic benchmark tells you nothing about your corpus.
What happens when a change fails the gate?
It does not ship, and you get the specific questions that regressed along with the retrieved evidence for each. Failures are diagnostic, not just a red light on a dashboard.
Can it evaluate systems we did not build with you?
That is an advisory engagement rather than this capability. AIMY Guard is built into the AIMY platform; assessing someone else's AI system is work we do, but it is scoped separately.
Sectors where AIMY Guard is deployed
Put AIMY Guard against your own documents.
We evaluate on a sample of your own corpus, never a canned dataset. That is the only way to tell whether retrieval will hold up on your content.