The paper introduces Multi-Agent Protocol Distillation (MAPD), a joint distillation and reinforcement-learning framework for agentic search. An offline multi-agent system decomposes queries, retrieves evidence, repairs failed searches, and converts exploration traces into a structured JSON protocol containing task type, reasoning plan, and extractive grounding facts. The protocol is exposed only to a privileged branch of the student policy during training, providing dense supervision without requiring teacher logits or tokenizer alignment. Across seven QA benchmarks, MAPD reports average success rates of 39.4% for Qwen3-1.7B and 44.4% for Qwen3-4B.
No heat snapshots are available in the last 24 hours.