Fusion is a staged adaptive-inference framework for Vision Transformers that coordinates token merging, early exiting, and token pruning. It first merges tokens, then evaluates confidence, and prunes only samples that continue inference. Lightweight routing modules adapt compression strength per input and allow the accuracy-latency trade-off to be changed at inference time without retraining. The abstract reports results on DeiT-S and ImageNet-1k that match or exceed prior adaptive ViT methods at comparable compute budgets, with up to 4x lower calibration error and 48% lower inference energy. Additional tests cover ImageNet-100, CIFAR-100, ImageNette, and multiple ViT backbones.
No heat snapshots are available in the last 24 hours.