The results table
Each row is a single test run.
Click any row to open the full test detail panel.
The test detail panel
For each test, you can see:The full prompt
The full prompt
Exactly what was sent to the test model.
The model's full response
The model's full response
The unedited reply from your test model.
Each judge's verdict
Each judge's verdict
Verdict, confidence, and reasoning — for all three judges, side by side.
The consensus calculation
The consensus calculation
A breakdown of how the voting method combined the three judges into a final verdict.
Human override (if any)
Human override (if any)
Who overrode the verdict, when, the new verdict, and the reasoning they gave.
Token usage and timing
Token usage and timing
How many tokens each judge used, and how long the whole test took.
Filters
Use the filter bar at the top of the Results tab to narrow results.
You can combine filters. The URL updates as you filter, so you can bookmark or share a filtered view.
Effective verdict vs final verdict
This distinction matters:- Final verdict = what the AI judges agreed on
- Effective verdict = what is actually used in dashboards and reports
- Dashboards show the effective verdict — so they reflect human judgment when available
- Risk assessments count failures based on the effective verdict
- Compliance reports use the effective verdict
Exporting results
From the Results tab toolbar:- Export CSV — flat table of all currently filtered results
- Export JSON — full structured data including judge reasoning and metadata
- Generate compliance report — see Compliance Reports
Sorting and grouping
- Click any column header to sort by it (default: most recent first).
- Group by category, severity, or verdict using the Group by dropdown.