Jet-Long is a tuning-free, zero-shot method for extending context beyond an open-weight model’s pretraining window. It combines a local window that preserves base RoPE behavior with a long-range window whose rescaling factor adapts to sequence length. Inclusion-exclusion attention merging and an on-the-fly correction rotation make the method inexpensive at inference, while a fused CuTe kernel reportedly reaches up to 1.39x FlashAttention-2 prefill throughput on H100, with no more than 4% single-batch generation overhead. On Qwen3 models at up to 128K context, it improves RULER over the strongest baseline by 4.79, 2.18, and 2.03 percentage points for 1.7B, 4B, and 8B models.
No heat snapshots are available in the last 24 hours.