SKT introduces a verified data-synthesis pipeline for training language-model agents to use reusable skills. Starting from 2,000 public skills, it creates single-skill and multi-skill task configurations, applies rule-based and agent-based verification, repairs failed generations with feedback, and retains trajectories that substantially use every required skill. The resulting dataset contains 4,000 task packages and 27,164 verified trajectories. The authors also build SkillEval from a disjoint test pool. According to the paper’s abstract, supervised fine-tuning on SKT trajectories improves skill-use performance across models, benchmarks, and agent harnesses, with gains linked to verification quality and broader skill coverage.
No heat snapshots are available in the last 24 hours.