This paper presents Ventaglio, a runtime-configurable sparse execution unit with RVV extensions for indexed gather-accumulate-scatter operations. The design targets Gustavson-style execution of sparse tensor contractions, avoiding software index decoding and L1-backed indexed memory operations. In a 12nm FinFET implementation, it reportedly accelerates sparse contraction kernels by 6.9--7.4x over optimized RVV baselines with 3.1% area overhead for a tightly L1-coupled vector cluster. On a DuoGPT-pruned LLaMA-3-8B with practical 40--60% dual sparsity, modeled speedups over dense baselines reach 2.40--5.25x for prefill and 2.06--3.16x for autoregressive decoding.
No heat snapshots are available in the last 24 hours.