Research

Ideas and evidence from the work of making AI agents reliable.

Twelve softly glowing bars ranked against a single axis, the top one in pale gold
Benchmark

Underwriting Is a Workflow, Not an Answer

Twelve model configurations evaluated on evidence, rule application, and completed work. A first look at the Underwriting Benchmark.

Read article