TurboVLA replaces the conventional vision-to-language-to-action pipeline with a direct vision-plus-language-to-action mapping. It separately encodes visual observations and language instructions, connects them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. The paper reports 0.2B parameters, 31.2 ms inference latency, and 0.9 GB inference VRAM on an RTX 4090, while reaching 97.7% average success on LIBERO. The code is publicly available.
No heat snapshots are available in the last 24 hours.