The paper introduces Autonomous Policy Evolution, a controlled setting in which a harness-model agent repeatedly edits an executable policy system while operating under a fixed interaction budget. Its benchmark, EvoPolicyGym, consists of compact interactive reinforcement-learning environments and evaluates how policies improve through iterative feedback. In addition to final task scores, it records trajectory-level diagnostics, including budget allocation, feedback-to-tuning conversion, and discovery of task-appropriate mechanisms. According to the abstract, GPT-5.5 achieves the strongest aggregate rank and places in the top two on all 16 environments.
No heat snapshots are available in the last 24 hours.