Evidence that can be inspected.
We test models and agents to help organisations understand their behaviour before or during deployment. Together, we define the system, the setting and the question to investigate. You receive a report explaining what happened, the evidence behind the findings and the limits of the conclusions.
A question grounded
in your setting.

Developers
AI developers preparing a model or agent for deployment.
Deployers
Deployers assessing a system in a specific operating setting.
Public institutions
Public institutions commissioning independent technical evidence.
Research partners
Research partners studying a defined safety question.
From the question
to a traceable report.
The system, the setting and the question define the evaluation.
Scope
Define the system, operating context, risk question and decision boundary.
Design
Build scenarios and controls that reflect the target setting. Fix the evaluation criteria before the run.
Execute
Run the system under controlled conditions. Preserve outputs, tool actions and relevant state changes.
Report
Separate observed results from interpretation. Record limitations and identify the next test when the evidence is incomplete.
Independent evaluation
- System and configuration
- Evaluation conditions
- Observed results
- Evidence and limitations
Findings you can
inspect and use.
A report documenting the system and configuration tested, the evaluation conditions, the observed results and the limits of the conclusions, with links to the preserved evidence.
Evaluation makes behaviour and limitations inspectable. It does not provide a blanket certification of safety or regulatory approval.