The paper introduces Multi-Head Latent Control, a lightweight post hoc layer that reads hidden-state trajectories from a frozen LLM or VLM. A Capability Head predicts whether the current model should solve an instance or defer to a stronger model, while a Resolution Head selects clarification, tool use, abstention, or direct answering. According to the abstract, routed systems reduce large-model usage by up to 90.7% on AndroidWorld and by 27–53% across benchmarks, while retaining most large-model performance. Tool-use decisions improve by up to 158% relatively, with 65.5% fewer missed required tool calls.
No heat snapshots are available in the last 24 hours.