All articles

Research · Underwriting agents

Would Your Underwriting Agent Make the Same Decision Twice?

Why one correct recommendation is not enough for agentic underwriting.

Applied WorldSeptember 24, 20268 min read
Five glowing paths leave one point; three arrive together at a green node, two drift to a separate red node
Five runs of the same case. Three hold the line; two drift.

An underwriting agent can reach the right answer once and still be unready for production. It may overlook a required check, invent a rule, or make a different recommendation when the same submission is run again. A single accuracy figure hides all three problems.

Underwriting is a workflow, not a prediction

A commercial submission rarely arrives as a neat question with all the facts attached. Someone has to read the documents, reconcile conflicting details, enrich the risk, ask for missing information, check appetite, consult the current guidelines, and decide whether the case needs referral. Only then is a recommendation useful. If an agent recommends referral by chance while missing the rule that required it, the final label may be right, but the work is not.

Underwriting tools are beginning to cover more of that sequence, from intake and triage to draft decisions. The product question is therefore changing. It is not just whether a model knows insurance terminology. It is whether an agent can carry a case through the insurer's actual process, using the allowed evidence and tools, while leaving a decision an underwriter can review.

What the research actually tests

UNDERWRITE, a commercial-underwriting agent benchmark developed with insurance experts, puts models into multi-turn tasks with underwriting guidelines, data tools, and a simulated user who does not hand over every fact at once. Its researchers evaluated 13 frontier models. They found that the most accurate models were not always the most efficient, and that models could hallucinate domain information despite having access to tools. Their repeated-run analysis reported a drop of up to 20% in answer correctness under a pass^k-style measure. That finding belongs to the published study, not to an Applied World evaluation.

The public sample of underwriting traces shows what these tasks involve: appetite checks, product recommendations, limits, deductibles, small-business eligibility, and business classification. These are multi-step conversations, not isolated multiple-choice questions. The researchers also note an important limit: the scenarios use synthetic and public information rather than real underwriter-agent logs.

Other benchmarks measure useful pieces of the problem. InsQABench tests insurance knowledge, database questions, and clause questions in a Chinese-insurance setting. INS-MMBench tests visual and multimodal insurance tasks. Both matter if an agent must read clauses or images, but success on those components does not establish that it can finish a carrier's underwriting workflow.

One submission, five runs

The earlier τ-bench introduced pass^k to examine whether an agent succeeds consistently over repeated attempts at the same task. For underwriting, that is a practical question: would the same case receive the same defensible treatment tomorrow?

Consider a commercial property submission that crosses a mandatory referral threshold in guideline §4.7. The case and rules stay fixed. In this illustrative example, the agent refers it correctly three times and misses the rule twice:

Illustrative repeated-run example; not measured customer or model performance.
RunRecommendationGuidelineResult
01REFER§4.7 appliedPASS
02ACCEPT§4.7 missedFAIL
03REFER§4.7 appliedPASS
04REFER§4.7 appliedPASS
05ACCEPT§4.7 missedFAIL

A single run might have caught only a success. Repeating the case shows that the agent is unstable, and the traces narrow down why: the referral rule was absent from the failed runs. The five rows are an illustration of the method, not a benchmark result or a model ranking.

There is no universal underwriting answer key

Two carriers may reasonably treat the same risk differently. Their appetite, delegated authority, exclusions, portfolio priorities, and referral thresholds may not match. Even within one carrier, a guideline revision can change the correct action. A public benchmark is useful for comparing general capabilities, but it cannot by itself establish compliance with one insurer's current rules.

That is why a production evaluation needs a private, carrier-specific case set. Each case should fix the submission, permitted evidence, tool state, applicable guideline version, expected decision, and escalation path. Expert review is especially important when the right response is to ask a broker for more information rather than force a premature accept-or-decline decision.

Score the path as well as the decision

A useful scorecard separates different failure modes instead of compressing them into one accuracy number:

Proposed underwriting-agent evaluation dimensions.
MeasureQuestion
Workflow successWas the entire task completed correctly?
Guideline adherenceWere the carrier’s appetite and authority rules applied?
Evidence groundingCan material findings be traced to the permitted evidence?
Referral accuracyWas the case escalated when required, and only then?
Repeated-run reliabilityDoes the same case receive a defensible result across runs?
Tool efficiencyWere searches and system actions necessary and proportionate?
Cost per successful caseWhat did a correctly completed workflow actually cost?

The trace should show which document was opened, which fact was extracted, which tool returned it, what rule was applied, when the agent asked for more information, and why it recommended acceptance, decline, or referral. An unsupported conclusion, a missed threshold, and an unnecessary escalation need different fixes. They should not disappear behind the same final-answer score.

Reliability has a cost

The cheapest model call is not necessarily the cheapest completed case. Retries, redundant searches, corrections, and underwriter review all consume time and money. Compare systems on cost per successful workflow: include the direct run cost and the operational cost of failures or human intervention, using the same case mix and tools. That can favor a more expensive model that finishes reliably, or a smaller model improved for a narrow workflow. The answer has to come from measurement, not token price alone.

A verified failure can also become useful training data. If reviewers confirm that the agent missed a referral rule, preserve the trace, correct the behavior, then replay the case and a held-out set. The test is whether the change improves the workflow without creating a new error elsewhere.

Make the result reviewable

Governance gives this work another purpose. The NAIC's model bulletin discusses insurers' governance, testing, validation, documentation, and oversight of AI systems; its application depends on adoption and applicable law in each jurisdiction. EIOPA's 2025 opinion addresses AI governance and risk management for European insurance supervision. Neither source says a benchmark score can replace human responsibility for an underwriting decision.

An evaluation receipt should therefore preserve the case and environment versions, model and agent versions, guideline version, evidence reviewed, tool calls, referral steps, final recommendation, and outcome. It lets an underwriter or reviewer see what happened, compare one release with the next, and challenge a result that looks correct for the wrong reason.

The production question is not whether an agent can get one submission right. It is whether it can work through the insurer's process repeatedly, under its rules, at an acceptable cost, with enough evidence left behind for a person to trust the decision.