This paper studies a minimal black-box adaptation method that learns a single context-independent logit-bias vector and adds it at every decoding step. The intervention requires no weight updates or gradients. Starting from a KL-regularized reinforcement-learning objective, the authors characterize when a fixed bias can approximate an optimal prefix-dependent correction and derive a closed-form inverse-propensity estimator using rollouts, rewards, and token probabilities. The abstract reports gains over base models on mathematical and reasoning benchmarks, with substantially fewer trainable parameters than conventional fine-tuning.
No heat snapshots are available in the last 24 hours.