Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

First seen · 7/31/2026, 04:00 AMLatest activity · 7/31/2026, 04:00 AM

This paper identifies “Training-Distribution Hallucination” in world action models: when observations differ visually from training data, pixel-generative future prediction may hallucinate training-domain appearance instead of preserving the current scene. ST-WAM uses DINOv3 as a shared semantic representation for future prediction and history retrieval, while retaining Wan-VAE latents for fine-grained visual dynamics. Its Dual-Space Future Experts jointly predict VAE latents and DINO features, and Current-Anchored Intent Retrieval selects task-relevant evidence from recent semantic history. The abstract presents a diagnosis and architectural proposal, but does not provide quantitative results or full experimental details.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/31, 04:00 AMnot independentRepresentative
    ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts