AgentLens introduces a production-assessed benchmark for interactive coding agents. Instead of reducing each run to a binary task result, it evaluates the full trajectory: instruction following, tool use, self-verification, error recovery, and communication. The benchmark combines formal verification where objective checks exist with LLM-written trajectory reviews and side-by-side comparisons. The authors report using it to diagnose agent behavior, compare successive versions of their own agent, and detect product regressions in a nightly evaluation pipeline. The benchmark is released as open source.
No heat snapshots are available in the last 24 hours.