The paper presents Wan-Streamer v0.3, framing video as a persistent world plus an incoming event stream. The world includes relatively stable visual, acoustic, and voice conditions, while the event stream covers environmental changes, behavior, speech, and other sounds. This framing supports general-purpose pretraining on real video and specialization for real-time full-duplex audio-visual interaction. The model maps multimodal user input to spoken language and behavior actions. It retains the v0.2 operating point: 640×368 video at 25 FPS, 160 ms streaming units, about 200 ms model-side latency, and about 550 ms total interaction latency with a 350 ms bidirectional network budget.
No heat snapshots are available in the last 24 hours.