This paper studies reliability in Leni, a production enterprise agent, across SpreadsheetBench Verified, BullshitBench v2, and the GAIA validation split. The full system improves over its frontier base model by 11.0, 7–10, and roughly 15 percentage points respectively. Its decomposition attributes most of the gain to scaffolding, routing, and lightweight specialist models. The isolated verification loop adds only 1.5 points, but its rescues are concentrated near the top of the score distribution. Instrumentation reports a verifier catch rate of about 0.20, a fix rate of 0.75, and no observed false-alarm regressions.
No heat snapshots are available in the last 24 hours.