The paper introduces Stabilized Adapter for self-Report (SAR), a lightweight LoRA adapter that uses only a fine-tuned model and the dataset used to train it. SAR prompts the model to describe hidden behaviors in plain language. Across seven implanted behaviors, it detected every behavior, including broad misalignment that was not predictable from the training data alone. Compared with Introspection Adapters (IA), the closest baseline, SAR preserved positive signal in settings where IA failed and roughly halved hallucinated, consistently incorrect reports. The method is presented as a practical auditing tool for understanding what fine-tuning actually induced.
No heat snapshots are available in the last 24 hours.