This paper frames computer vision as unified multimodal generation. SenseNova-Vision uses natural-language instructions and optional visual prompts to define tasks, regions, views, and decoding conventions. It emits text for symbolic outputs, images for dense spatial predictions, or mixed text-and-image responses for compositional tasks. The model starts from an off-the-shelf pretrained unified multimodal model and is trained mainly on the SenseNova-Vision Corpus, an instruction-response corpus converted from diverse vision annotations. It requires no task-specific prediction heads or architectural changes, and reportedly covers detection, OCR, keypoints, segmentation, depth, surface normals, point maps, and camera pose estimation.
No heat snapshots are available in the last 24 hours.