TRACE-Router addresses the mismatch between per-call LLM routing and delayed, task-level feedback in long-horizon agent workflows. At task admission, a contextual bandit selects one backend and pins all subsequent calls to it. After completion, the router updates its policy using the terminal reward, jointly reflecting accuracy and latency, without explicitly estimating task complexity. The paper reports improved accuracy-latency trade-offs across three agentic benchmarks. On tau2-Bench, it exceeds latency-matched interpolation between individual models by 7–8 accuracy points; on Terminal-Bench, it reports 7.1 higher accuracy points than the strongest single-model baseline with 36% lower latency.
No heat snapshots are available in the last 24 hours.