This paper introduces OpenAgent, a setting for evaluating tool-use agents under distribution shifts in queries, actions, observations, and domains. In a controlled sandbox, the authors organize environmental changes into four levels: Perception, Interaction, Reasoning, and Internalization. Their reported analysis finds that agents trained with both supervised fine-tuning and reinforcement learning degrade to varying degrees when facing open-world shifts. They also propose Perturbation-Augmented Fine-Tuning, a disturbance-based SFT intervention intended to improve robustness and utility in less predictable environments. The abstract states that code will be released.
No heat snapshots are available in the last 24 hours.