ELiTeFormer co-designs hybrid linear attention and ternary linear projections for FPGA-based Transformer inference. The authors report 10x weight compression and 12.8x KV-cache compression versus LLaMA 3, with 31.9% MMLU accuracy, within 3.0 percentage points of BitNet b1.58. Its processing element replaces multiplications in ternary projections with bitmasking, avoiding dedicated DSP blocks. Targeting a Xilinx VCK5000 Versal board through HLS, block-level simulations show 9.6x FFN and 4.4x attention speedups, while end-to-end deployment reports up to 3.9x lower latency and 3.2x better energy efficiency than LLaMA 3 on an NVIDIA A100 at long contexts.
No heat snapshots are available in the last 24 hours.