This paper conducts a controlled comparison of Self-Refine, Best-of-N, and Debate against task-only and chain-of-thought baselines across five LLM backbones and three domains: competitive programming, chess puzzles, and mathematics. All methods receive the same GEPA optimization budget and are tested on difficulty-stratified items. Averaged within benchmarks, orchestration improves accuracy by at most 4.6 percentage points over optimized CoT and 4.5 points over task-only inference, while using roughly 2–4 times the mean total tokens. Gains do not consistently grow with task difficulty and vary strongly by backbone model.
No heat snapshots are available in the last 24 hours.