Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning

First seen · 7/26/2026, 07:41 AMLatest activity · 7/26/2026, 07:41 AM

This paper proposes inference-time consensus decoding as a defense against hidden behaviors introduced by fine-tuning. A separate reference model is fine-tuned on each data source, and their next-token distributions are aggregated during decoding. The authors introduce a token-wise minimum decoder and a base-relative variant that falls back to the base model when sources move in opposite directions. They report suppression of source-specific behavior across controlled poisoning, subliminal learning, and emergent misalignment tasks, including settings where union training and weight averaging fail to remove the behavior.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/26, 07:41 AMnot independentRepresentative
    Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning