SenseTime Unveils SenseNova-U1.5: An 8B Native Unified Vision Model Without Encoders or VAEs
First seen · 9/10/2026, 04:00 AMLatest activity · 9/11/2026, 01:59 AM
SenseTime has introduced SenseNova-U1.5, an 8B native unified multimodal model designed to perceive, reason, and generate visual content without external visual encoders or VAEs. The architecture relies on spatially coherent patch reconstruction and natively handles resolutions up to 4K. By consolidating domain-specific experts—covering bilingual typography, infographic rendering, and multi-reference editing—via on-policy distillation, the model directly bridges multimodal understanding into structured visual generation. The team plans to open-source its training pipeline, including fine-tuning and reinforcement learning code.
Event heat · last 24 hours
There are 8 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 08:00; latest heat is 0.