This paper trains a small language model to jointly select retrieval agents and generate structured downstream tool parameters. The method uses progressive supervised fine-tuning followed by reinforcement learning with a hierarchical reward based on retrieval relevance and query-agent topic alignment. On a targeted subset of agent-query mismatches, it reaches 0.918 NDCG@10, versus 0.539 for Amazon Nova Lite and 0.490 for Claude Haiku 4.5. Overall mean NDCG@10 is 0.771, with 120.1 ms mean selection latency, an 82.4% reduction versus Nova Lite.
No heat snapshots are available in the last 24 hours.