AutoTrainess presents an LM agent for autonomous post-training. It exposes planning, data preparation, training, evaluation, and experiment logging through agent-computer interfaces, while encoding human experience as explicit workflows, rules, and execution constraints. On PostTrainBench, GPT-5.4 (Codex) with AutoTrainess reaches an average score of 26.94, compared with 23.21 for a CLI-only baseline. The approach also transfers across models and harnesses: DeepSeek-V4-Flash with OpenCode improves from 12.13 to 19.58. The abstract does not provide detailed ablations or full evaluation methodology.
No heat snapshots are available in the last 24 hours.