Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
arXiv·Haiwen Diao·Sep 10, 2026, 5:59 PM

SenseTime Unveils SenseNova-U1.5: An 8B Native Unified Vision Model Without Encoders or VAEs

Original title:SenseNova-U1.5: Towards Native Unified Visual Intelligence

Models82

SenseTime has introduced SenseNova-U1.5, an 8B native unified multimodal model designed to perceive, reason, and generate visual content without external visual encoders or VAEs. The architecture relies on spatially coherent patch reconstruction and natively handles resolutions up to 4K. By consolidating domain-specific experts—covering bilingual typography, infographic rendering, and multi-reference editing—via on-policy distillation, the model directly bridges multimodal understanding into structured visual generation. The team plans to open-source its training pipeline, including fine-tuning and reinforcement learning code.

Why it's worth reading

It illustrates a viable path toward end-to-end visual intelligence, bypassing standard vision encoders and latent diffusion components to achieve 4K multimodal generation and reasoning at an 8B scale.

Tags

商汤SenseNova多模态大模型统一视觉模型图像生成计算机视觉开源

Also reported by

  • HuggingFace Daily Papers — SenseNova-U1.5: Towards Native Unified Visual Intelligence

Score breakdown

  • Novelty84
  • Impact81
  • Practicality80
  • Credibility83
  • Timeliness85