Wnuan is a three-stage post-training pipeline for adapting models to proprietary enterprise knowledge: task-oriented supervision is generated from documents, supervised fine-tuning uses general-data replay, and reinforcement learning targets residual errors. According to the supplied abstract, the main 32B route improves acceptable-answer rate on the 707-question WnuanBench from 52.76% before adaptation to 80.06% after SFT and 91.51% after RL. The reported gain comes with a 5.17-point average decline on general benchmarks, concentrated in instruction following.
No heat snapshots are available in the last 24 hours.