This paper presents an environment-free method for generating training trajectories for API-calling agents. Given only API specifications, an LLM creates diverse tasks, a teacher agent attempts to solve them, and another LLM simulates stateful API responses conditioned on the task and interaction history. A judge model filters the resulting trajectories. The authors evaluate the approach on AppWorld and OfficeBench, covering information retrieval and state-changing tasks, and report significant gains after fine-tuning on the synthetic data.
No heat snapshots are available in the last 24 hours.