This paper proposes RINO, or RGB In and RGB Out, a unified formulation for vision tasks. Masks, depth maps, poses, and other structured visual signals are represented as RGB images, while tasks are expressed as RGB-to-RGB image editing. Using a generic image-editing backbone without task-specific fine-tuning, RINO reports zero-shot results on dense understanding tasks such as segmentation and depth estimation, as well as dense-conditioned generation such as pose-to-image synthesis. The authors argue that a shared visual interface may play a role analogous to text in language models. Code is available on GitHub.
No heat snapshots are available in the last 24 hours.