Inside an insurance evaluation.

Evaluate insurance agents across claims and underwriting. These six instruments are illustrated with a commercial property underwriting case; all figures are illustrative.

Measure what matters in production.

Six measures, scored against acceptance criteria you set. One number for readiness, and the rows that hold it back.

Production readiness

90/ 100

HOLD 1 of 6 below your 85% bar

Repeated-run reliability is the gap to close before release.

MeasureScore against an 85% bar
  1. Workflow successDid the agent complete the full task?
    88%
  2. Guideline adherenceDid it apply your appetite and authority rules?
    96%
  3. Evidence groundingWere findings supported by permitted evidence?
    89%
  4. Referral accuracyDid it escalate at the right point?
    94%
  5. Repeated-run reliabilityWas behaviour stable across repeated runs?
    84%
  6. Tool efficiencyDid it avoid unnecessary tool actions?
    87%
  7. Cost per successful caseWhat does a correctly completed workflow cost?
    $0.74

Your guidelines, not a generic benchmark.

From claims coverage to underwriting referrals, decisions depend on your policies and authority rules. We build private benchmarks against that standard.

Overhead view of a commercial property case file: aerial photograph, floor plan, ledger pages
Same submission, two carriers
Occupancy
Warehouse, 38% storage
Construction
Non-combustible
Protection
Sprinklered
  • Carrier AGuideline §4.7

    Refer any risk above 25% storage occupancy.

    REFER
  • Carrier BGuideline §6.2

    Accept up to 40% storage when sprinklered.

    ACCEPT

Both answers are correct. A generic benchmark can only grade one of them.

Know exactly where the agent failed.

Every failed run is traced step by step and classified, so a low score turns into a specific fix.

Failure distribution by type
  1. Guideline interpretation
    31%
  2. Missing evidence
    21%
  3. Incorrect referral
    16%
  4. Premature decision
    12%
  5. Tool misuse
    9%
  6. Unnecessary follow-up
    7%
  7. Other
    4%
Case UW-1842, run 02, agent trace
  1. Read submissionPASS
  2. Read loss historyPASS
  3. Review property inspectionPASS
  4. Check appetitePASS
  5. Apply guideline §4.7FAIL
  6. Recommendation: ACCEPTFAIL

Storage occupancy of 38% exceeded the carrier's 25% referral threshold. The agent read the figure and did not apply §4.7.

Failure type
Guideline omission
Expected
Refer
Observed
Accept

Turn failures into training signal.

Verified failure trajectories become targeted post-training examples, then the agent runs the same benchmark again.

  1. 01Your workflow
  2. 02Agent runs
  3. 03Verified failures
  4. 04Training data
  5. 05Post-training
  6. 06Regression tests
  7. 07Improved agent

Use the smallest model that clears the bar.

Compare models in the same environment on reliability and total cost. Optimise for cost per successful case, not per token.

Reliability against cost per successful case
75%80%85%90%95%100%$0.00$0.50$1.00$1.50$2.00$2.50$3.00$3.50Your bar, 90%Frontier modelMid-size modelSmall modelSmall + post-training
ModelReliability / Cost
  1. Frontier model96%$3.14
  2. Mid-size model93%$1.12
  3. Small model81%$0.29
  4. Small + post-training92%$0.38

8×

lower cost per successful case than the frontier model, above your 90% bar.

A record behind every result.

Each run keeps the case, environment and model versions, trajectory, guideline checks, evidence, tools, and outcome. Human underwriters keep judgment and authority over consequential decisions.

  • TraceableEvery decision links to the evidence and rule behind it.
  • VersionedAgent, model, guidelines, and environment are pinned.
  • ReviewableUnderwriters and auditors can replay any run.
Evaluation receiptUW-10428
PASS
Environment
Commercial Property v0.4
Agent
UW-Agent-v17
Model
UW-8B-03
Guidelines
Property-v28
Recommendation
REFER
Evidence checks
6 / 6
Guideline checks
4 / 4
Tool calls
9
Human escalation
Correct
Cost
$0.68

Run this evaluation on one of your own workflows.

Build a pilot