This paper proposes inference-time consensus decoding as a defense against hidden behaviors introduced by fine-tuning. A separate reference model is fine-tuned on each data source, and their next-token distributions are aggregated during decoding. The authors introduce a token-wise minimum decoder and a base-relative variant that falls back to the base model when sources move in opposite directions. They report suppression of source-specific behavior across controlled poisoning, subliminal learning, and emergent misalignment tasks, including settings where union training and weight averaging fail to remove the behavior.
No heat snapshots are available in the last 24 hours.