Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

First seen · 9/5/2026, 01:47 AMLatest activity · 9/5/2026, 01:47 AM

Vision-language models (VLMs) are frequently repurposed as reward functions for robotic reinforcement learning, where they are assumed to assign consistent progress scores to semantically equivalent instructions. Introducing ROBORMBENCH—a benchmark containing 2,390 real-robot trajectories and 21,673 verified paraphrases—the authors reveal that both open and proprietary VLMs exhibit severe paraphrase fragility, sometimes flipping the evaluation of identical executions between success and failure. Larger model scales or explicit reasoning do not reliably resolve the issue, underscoring the necessity of trajectory-grounded supervision for dependable robotic reward modeling.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv9/5, 01:47 AMnot independentRepresentative
    Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models