SkillCoach introduces a self-evolving rubric framework for evaluating how LLM agents use reusable skills such as SOPs, domain rules, tool workflows, scripts, and validation routines. It derives skill-grounded process rubrics from real rollouts and scores trajectories on four dimensions: skill selection, skill following, skill composition, and skill-grounded reflection. The framework keeps the external verifier as a separate outcome signal, distinguishing sound execution from accidental success. The paper reports that evolved rubrics reveal failures hidden by final accuracy and provide stronger supervision for selecting training trajectories than outcome-only filtering.
No heat snapshots are available in the last 24 hours.