90/ 100
HOLD 1 of 6 below your 85% bar
Repeated-run reliability is the gap to close before release.
Evaluate insurance agents across claims and underwriting. These six instruments are illustrated with a commercial property underwriting case; all figures are illustrative.
Six measures, scored against acceptance criteria you set. One number for readiness, and the rows that hold it back.
90/ 100
HOLD 1 of 6 below your 85% bar
Repeated-run reliability is the gap to close before release.
From claims coverage to underwriting referrals, decisions depend on your policies and authority rules. We build private benchmarks against that standard.

Refer any risk above 25% storage occupancy.
REFERAccept up to 40% storage when sprinklered.
ACCEPTBoth answers are correct. A generic benchmark can only grade one of them.
Every failed run is traced step by step and classified, so a low score turns into a specific fix.
Storage occupancy of 38% exceeded the carrier's 25% referral threshold. The agent read the figure and did not apply §4.7.
Verified failure trajectories become targeted post-training examples, then the agent runs the same benchmark again.
Compare models in the same environment on reliability and total cost. Optimise for cost per successful case, not per token.
8×
lower cost per successful case than the frontier model, above your 90% bar.
Each run keeps the case, environment and model versions, trajectory, guideline checks, evidence, tools, and outcome. Human underwriters keep judgment and authority over consequential decisions.