MemSFT separates domain specialization from backbone updates by training a plug-and-play parametric memory to imitate a non-parametric retriever over domain data. A learned router fuses memory and backbone output distributions at each decoding step, enabling selective use of domain expertise. The paper evaluates biology, geoscience, and law settings with models from Qwen3-8B to Qwen3-235B-A22B, reporting improved domain performance with negligible general-task degradation, while full SFT causes substantial forgetting. The provided abstract does not include exact scores, baselines, training costs, or ablation details.
No heat snapshots are available in the last 24 hours.