Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach

First seen · 7/15/2026, 07:11 PMLatest activity · 7/15/2026, 07:11 PM

This paper trains a small language model to jointly select retrieval agents and generate structured downstream tool parameters. The method uses progressive supervised fine-tuning followed by reinforcement learning with a hierarchical reward based on retrieval relevance and query-agent topic alignment. On a targeted subset of agent-query mismatches, it reaches 0.918 NDCG@10, versus 0.539 for Amazon Nova Lite and 0.490 for Claude Haiku 4.5. Overall mean NDCG@10 is 0.771, with 120.1 ms mean selection latency, an 82.4% reduction versus Nova Lite.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/15, 07:11 PMnot independentRepresentative
    SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach