OmegaUse-OfficeVal is a benchmark for evaluating LLM agents on long-horizon office-suite workflows with task-level economic grounding. It contains 100 practitioner-derived tasks adapted through a privacy-preserving process, requiring 2.32 hours of human labor on average. Each task includes human labor time and a task-price proxy, enabling comparisons between human and inference costs. Code-based verifiers are built from fine-grained rubrics. The abstract reports that evaluated frontier LLMs are substantially faster and cheaper than humans, but still below human-level deliverable quality. The dataset and code are open-sourced.
No heat snapshots are available in the last 24 hours.