MADA-RL is a post-training framework for compact models of up to 4B parameters. It assigns generator and critic roles to specialized agents and trains them with a counterfactual critic advantage: the critic reward is compared against the generator ensemble’s per-instance accuracy. Using LoRA adapters, the method improves DeepSeek-R1-Distill-Qwen-1.5B from 39.9% to 41.9% across five mathematical reasoning benchmarks, with 16 times fewer trainable parameters than fully fine-tuned baselines. The paper reports that MADA-RL remains below DeepScaleR and STILL-3, while requiring additional multi-round inference.
No heat snapshots are available in the last 24 hours.