All articles

Underwriting Benchmark

Underwriting is a workflow, not an answer.

The file changed. Did the agent follow through?

5 min readPreliminary evaluation
Twelve softly glowing bars ranked against a single axis, the top one in pale gold
5Rubric checks per task
100Available points per task
12Model configurations

An agent can spot a missing document, explain a referral, and still leave the underwriting file unfinished. We built the Underwriting Benchmark to test whether it actually carries the work through.

Across our first twelve configurations, the highest mean diagnostic score was 71.75 out of 100. The results capture how much of the required work each configuration completed under the rubric.

01 / OBSERVED RESULTS

Model performance

Diagnostic scores across twelve model configurationsVertical bar chart with model configurations on the x-axis and diagnostic score from zero to one hundred on the y-axis. Models are ordered alphabetically. Exact values appear in the table below.DIAGNOSTIC SCORE / 10002040608010069.50Claude Fable 5.167.50Claude Opus 564.50DeepSeek V4.1 Flash52.00Gemini 3.8 Flash48.75GLM-5.349.21GLM-5.3 Flash70.50GPT-5.6 Sol70.75GPT-6 Astra71.75Grok 4.768.00Kimi K339.25MiniMax M343.89Qwen 3.8 MaxMODEL CONFIGURATION
View scores
Evaluation snapshot · 26 September 2026
ConfigurationScore /100
Claude Fable 5.169.50
Claude Opus 567.50
DeepSeek V4.1 Flash64.50
Gemini 3.8 Flash52.00
GLM-5.348.75
GLM-5.3 Flash49.21
GPT-5.6 Sol70.50
GPT-6 Astra70.75
Grok 4.771.75
Kimi K368.00
MiniMax M339.25
Qwen 3.8 Max43.89

What the agent has to do

Each task starts with fresh simulation state, case documents, and explicit rules. The work ranges from assembling submissions and reconciling loss runs to checking property evidence, creating referrals, and preparing a quote. Focused tasks isolate individual capabilities, while a composite workflow brings several together.

Agents use controlled tools to read evidence and change records. Scripted brokers and approvers respond when the workflow prerequisites are met. The evaluator checks findings, citations, final state, and audit history. Saying “I created the referral” earns no credit for a referral that does not exist.

The rubric gives each task five weighted checks, covering the relevant facts, rule application, supporting evidence, and required actions. Each check has an explicit condition for earning credit. Passing checks contribute to the diagnostic score; full success also requires the correct final state and no critical failure. Scoring is deterministic, with no LLM judge.

02 / A CHANGING FILE

New evidence changes the next action.

  1. 01 / INTAKE$650k

    Building exposure. Part of the loss history is missing.

  2. 02 / BROKER RESPONSE+$250k

    The requested history arrives with a revised contents schedule.

  3. 03 / AUTHORITY CHECK$900k

    Updated exposure exceeds the fictional $800k quote ceiling.

  4. 04 / HANDOFFApproval

    Obtain approval for the current revision, then prepare the quote for review.

Unscored illustration with fictional amounts. This is an example of the workflow, not a recorded model trace. Quote-only approval does not authorize binding.

Rubric breakdown

For the illustrative workflow above, these five checks follow the composite task’s scoring structure. Other tasks have their own checks and weights.

  • Complete the evidence

    Obtain the required loss history and link the documents to the correct entity and risk.

    20 points
  • Apply the latest revision

    Incorporate the revised schedule and recalculate the current exposure.

    20 points
  • Respect approval authority

    Request and validate external approval for the current risk, revision, amount, and action.

    25 points
  • Prepare the correct quote

    Use the current rating response and preserve the quoted premium, currency, and full policy term.

    20 points
  • Preserve the audit trail

    Keep evidence and approval history traceable, and leave coverage unbound.

    15 points

100 available points. This example is unscored. A full pass also requires the terminal state and no critical failure.

Where the results get interesting

Grok 4.7 led with a score of 71.75 out of 100. GPT-6 Astra followed at 70.75 and GPT-5.6 Sol at 70.50, placing the top three within 1.25 points of each other. Claude Fable 5.1 scored 69.50, close behind.

Kimi K3 scored 68.00, Claude Opus 5 scored 67.50, and DeepSeek V4.1 Flash scored 64.50. Gemini 3.8 Flash reached 52.00, while both GLM configurations scored just below 50. Qwen 3.8 Max and MiniMax M3 recorded 43.89 and 39.25 respectively.

Complete workflows appeared in submission assembly, missing-information completion, and deductible exceptions. Across the evaluation, even the highest-scoring model left some rubric requirements unmet. The overall scores show how much of the required work each model completed; reviewing the individual checks helps identify where it fell short.

What this snapshot can tell us

The rubric makes the scoring conditions explicit so reviewers can examine what each check rewards and whether its weight reflects the importance of the work. Independent underwriting review remains pending.

Some episodes have no valid diagnostic score. An earlier GLM-5.3 Flash run used a fixture missing a new-business qualifier. Qwen 3.8 Max produced malformed tool-call arguments. We keep those exceptions visible rather than assigning them zero.

Next comes reviewing the rules and failed checks, inspecting traces, and running more cases under matched settings. That work will help separate incomplete execution from valid answers the grader may have rejected.