Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

ReferTrack: Referring Then Tracking for Embodied Visual Tracking

First seen · 7/24/2026, 12:00 PMLatest activity · 7/24/2026, 12:00 PM

ReferTrack proposes a “referring-then-tracking” paradigm for embodied visual tracking with a single forward-facing camera. The model first selects a language-referred target from an indexed set of image bounding boxes, then predicts tracking waypoints conditioned on that grounded choice. A sliding queue of historical boxes is encoded through temporal-viewpoint-bbox indicator (TVBI) tokens to preserve target motion cues. The paper reports EVT-Bench success rates of 89.4%, 73.3%, and 74.1% on single-target, distracted, and ambiguity splits, respectively, and describes real-world deployments on legged and humanoid robots.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/24, 12:00 PMnot independentRepresentative
    ReferTrack: Referring Then Tracking for Embodied Visual Tracking