GDPevo is an evolution-native benchmark for evaluating agents that update persistent state from prior experience. It decomposes enterprise workflows into atomic business rules, distributes rules across training tasks, and recombines them into held-out tests, making performance gains more attributable while reducing contamination risk. V1 contains 120 tasks across 12 groups covering CRM, ERP, finance, healthcare, legal, and data-centric workflows. Across four agents and four supervision types, self-evolution improved held-out accuracy by up to 16.44 percentage points. However, the best evolved agents remained well below the 91.6% fully informed oracle ceiling.
No heat snapshots are available in the last 24 hours.