Underwriting Benchmark
Underwriting is a workflow, not an answer.
The file changed. Did the agent follow through?

An agent can spot a missing document, explain a referral, and still leave the underwriting file unfinished. We built the Underwriting Benchmark to test whether it actually carries the work through.
Across our first twelve configurations, the highest mean diagnostic score was 71.75 out of 100. The results capture how much of the required work each configuration completed under the rubric.
Model performance
View scores
| Configuration | Score /100 |
|---|---|
| Claude Fable 5.1 | 69.50 |
| Claude Opus 5 | 67.50 |
| DeepSeek V4.1 Flash | 64.50 |
| Gemini 3.8 Flash | 52.00 |
| GLM-5.3 | 48.75 |
| GLM-5.3 Flash | 49.21 |
| GPT-5.6 Sol | 70.50 |
| GPT-6 Astra | 70.75 |
| Grok 4.7 | 71.75 |
| Kimi K3 | 68.00 |
| MiniMax M3 | 39.25 |
| Qwen 3.8 Max | 43.89 |
What the agent has to do
Each task starts with fresh simulation state, case documents, and explicit rules. The work ranges from assembling submissions and reconciling loss runs to checking property evidence, creating referrals, and preparing a quote. Focused tasks isolate individual capabilities, while a composite workflow brings several together.
Agents use controlled tools to read evidence and change records. Scripted brokers and approvers respond when the workflow prerequisites are met. The evaluator checks findings, citations, final state, and audit history. Saying “I created the referral” earns no credit for a referral that does not exist.
The rubric gives each task five weighted checks, covering the relevant facts, rule application, supporting evidence, and required actions. Each check has an explicit condition for earning credit. Passing checks contribute to the diagnostic score; full success also requires the correct final state and no critical failure. Scoring is deterministic, with no LLM judge.
New evidence changes the next action.
- 01 / INTAKE$650k
Building exposure. Part of the loss history is missing.
- 02 / BROKER RESPONSE+$250k
The requested history arrives with a revised contents schedule.
- 03 / AUTHORITY CHECK$900k
Updated exposure exceeds the fictional $800k quote ceiling.
- 04 / HANDOFFApproval
Obtain approval for the current revision, then prepare the quote for review.
Rubric breakdown
For the illustrative workflow above, these five checks follow the composite task’s scoring structure. Other tasks have their own checks and weights.
- 20 points
Complete the evidence
Obtain the required loss history and link the documents to the correct entity and risk.
- 20 points
Apply the latest revision
Incorporate the revised schedule and recalculate the current exposure.
- 25 points
Respect approval authority
Request and validate external approval for the current risk, revision, amount, and action.
- 20 points
Prepare the correct quote
Use the current rating response and preserve the quoted premium, currency, and full policy term.
- 15 points
Preserve the audit trail
Keep evidence and approval history traceable, and leave coverage unbound.
100 available points. This example is unscored. A full pass also requires the terminal state and no critical failure.
Where the results get interesting
Grok 4.7 led with a score of 71.75 out of 100. GPT-6 Astra followed at 70.75 and GPT-5.6 Sol at 70.50, placing the top three within 1.25 points of each other. Claude Fable 5.1 scored 69.50, close behind.
Kimi K3 scored 68.00, Claude Opus 5 scored 67.50, and DeepSeek V4.1 Flash scored 64.50. Gemini 3.8 Flash reached 52.00, while both GLM configurations scored just below 50. Qwen 3.8 Max and MiniMax M3 recorded 43.89 and 39.25 respectively.
Complete workflows appeared in submission assembly, missing-information completion, and deductible exceptions. Across the evaluation, even the highest-scoring model left some rubric requirements unmet. The overall scores show how much of the required work each model completed; reviewing the individual checks helps identify where it fell short.
What this snapshot can tell us
The rubric makes the scoring conditions explicit so reviewers can examine what each check rewards and whether its weight reflects the importance of the work. Independent underwriting review remains pending.
Some episodes have no valid diagnostic score. An earlier GLM-5.3 Flash run used a fixture missing a new-business qualifier. Qwen 3.8 Max produced malformed tool-call arguments. We keep those exceptions visible rather than assigning them zero.
Next comes reviewing the rules and failed checks, inspecting traces, and running more cases under matched settings. That work will help separate incomplete execution from valid answers the grader may have rejected.