The paper introduces SVA (Search, Value, and Act), which uses Monte Carlo tree search in simulation to explore the output distribution of a frozen vision-language-action policy and trains a lightweight Q-value model from empirical returns. At deployment, the frozen VLA proposes multiple actions and the evaluator selects the candidate with the highest uncertainty-regularized value, without simulator access. A diagnostic pass@k study reports success increasing from 33% at pass@1 to 92% at pass@32. The abstract also reports that a 9B VLA outperformed a 27B model by 7 points with 27% lower inference latency in the stated evaluation setting.
No heat snapshots are available in the last 24 hours.