Benchmark · Insurance claims
ClaimsBench: Testing AI on Insurance Claim Decisions
Thirty fictional property claims test how an agent investigates a loss, applies coverage, reconciles payment, and resolves disputes.

Overview
The first release of ClaimsBench has 30 different claims tasks. They cover a variety of losses, including wind, theft, fire/smoke, 1st party, frozen pipe, building collapse, renters, a water discharge with 4 covered line items, and an 8 line item condo/earthquake loss. All the claims are fictional and are based on insurance policy forms published by the Oklahoma Insurance Department, such as the AP2 Deluxe Homeowners Policy form plus a stand-alone earthquake policy for some tasks. No customer claim files are included in ClaimsBench. The goal is to create claims that are interesting and representative of common types of claims that are filed, not to provide coverage answers to actual people with losses. While the Insurance Department forms are helpful, the actual insurance contract that a customer signs controls coverage. Copies on Insurance Department websites should not be used to answer coverage questions in actual claims.
Each task starts with a new claim to be handled. The agent conducts an initial claim interview to learn the basics of the claim and organizes the claim into five information groups. Then, the agent investigates the loss, develops findings and creates an item-decision worksheet to make an interim claim payment decision. New information surfaces and the agent must decide what, if any, changes to make to their original decision. You cannot "unpay" money, so the agent must explain why they are making changes while keeping original claim payments. Sometimes, new information will support the original decision, so agents shouldn’t feel obligated to make changes. Next, the agent speaks with the claimant to resolve any disputed items. Lastly, the agent must prepare memos for claim closure and potential reconsideration if warranted.
The chart below summarizes these five workflow phases: initial claim, investigation, interim resolution, final resolution, and reconsideration.
Five workflow phases
At each phase, the agent will develop, modify, or abandon an investigation, a set of findings, an item-decision worksheet, a reconsideration memo, and/or a closeout memo. These artifacts make the decisions and their supporting evidence available for review.
Scoring considers both artifact completion and the accuracy and usefulness of its contents. Five reward components contribute 20% each: investigation and intake, interim adjudication, evidence-based revision, coverage and financial result, and dispute and claim closure.
These components overlap across workflow phases. Payment accuracy, for example, matters in interim adjudication, evidence-based revision, and the final financial result. We also track strict completion of all required artifacts separately; it is not currently weighted in the final score.
Five equal-weight reward components
Preliminary estimates
The chart below preserves our preliminary score estimates for six model versions. These are projections, not measured benchmark results. The six models have not been evaluated together under the current scoring rubric; earlier evaluations used previous versions of the ClaimsBench toolset.
A score alone does not explain a model’s failures. Evaluation traces and completed artifacts are needed to understand which evidence it used, which decisions it supported, and where it made mistakes.
Projected overall reward
Next steps
To discuss a private claims evaluation, bring a defined workflow, policy forms, authority rules, and the evidence your team uses to review decisions.
Contact us if you're interested in ClaimsBench in your organization.