Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

TurboVLA: A Real-Time Vision-Language-Action Model Running at 32 Hz with Under 1 GB VRAM on an RTX 4090

First seen · 7/30/2026, 12:00 PMLatest activity · 7/30/2026, 12:00 PM

TurboVLA replaces the conventional vision-to-language-to-action pipeline with a direct vision-plus-language-to-action mapping. It separately encodes visual observations and language instructions, connects them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. The paper reports 0.2B parameters, 31.2 ms inference latency, and 0.9 GB inference VRAM on an RTX 4090, while reaching 97.7% average success on LIBERO. The code is publicly available.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/30, 12:00 PMnot independentRepresentative
    TurboVLA: A Real-Time Vision-Language-Action Model Running at 32 Hz with Under 1 GB VRAM on an RTX 4090