AgentCompass is presented as an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. It separates evaluation into three independent components: Benchmark, Harness, and Environment, allowing researchers to compose evaluations without reimplementing execution logic. The system also includes a fault-tolerant asynchronous runtime and trajectory-analysis tools intended to expose subtle failures such as reward hacking. According to the abstract, it natively supports more than 20 benchmarks spanning five capability dimensions, with reproducibility and scalable experimentation as its main goals.
No heat snapshots are available in the last 24 hours.