The paper formulates Transformer-to-hybrid conversion as budget-constrained subset optimization over layers and introduces FlashMorph, or Fast LAyer Selection for Hybrid MORPHing. FlashMorph adds a converted linear-attention branch to every full-attention layer, freezes the model weights, and jointly learns layerwise gates on synthetic long-context retrieval data. A linearization regularizer encourages reliance on linear attention. The learned gates are discretized under a preset full-attention budget, followed by logits distillation and long-context fine-tuning. The authors report stronger hybrid configurations, preserved long-context recall and general benchmark performance, and substantially lower layer-selection cost than prior methods.
No heat snapshots are available in the last 24 hours.