Two selected outcomes from March 2026, judged on transcript analysis. These historical provider records are separate from the current offline replay and operating guide.
Two battles worth opening.
The judge compared the complete answer sets against the same 100-point rubric. Select a battle to inspect its decision.
Turn a messy transcript into a useful handoff.
Extract the objective, facts, decisions, implementation details, constraints, next steps, and open questions.
Every candidate sees the task instructions and one source transcript at a time. The battle judge receives the complete comparison across all ten examples.
Read the full original task instructions
The original rubric
These are the rubric’s available points. The preserved decision does not report numeric category scores.
One input. Both original answers.
Choose any of the ten examples. Outputs remain literal text, including their original headings and formatting.
What the judge actually decided.
One decision across the complete pair of answer sets. Read the original reasoning, then inspect exactly what the judge received.
The judge’s original reason
The judge supplied this confidence value; it is not calibrated accuracy.
Read the judge instructions
Read the full assembled judge request
The task, weighted rubric, and all source-and-answer pairs as preserved in the selected record. Any privacy substitutions are identified below.
Read the original decision response
A record you can inspect.
These two preserved comparisons come from the historical Transcript Model Evaluator. They show actual candidate responses and recorded judging, rather than the fixed responses used in the current offline replay.
Opening this page makes no model calls. The data is included in the page for offline reading and supplied separately for inspection.
Reviewed data · SHA-2564da4aedf2c98d571e00b9f1ec4ff19cd1189e1771a51a71136dd09a9c281c415