The paper introduces Self-Correcting Coupled Markov Jump Processes (SC-CMJP), where text and image transition rates depend on the other modality’s confidence, weighted by cross-modal attention. A remasking jump can retract commitments when new evidence creates contradictions. Built on this framework, the training-free, single-pass CO₂Jump sampler targets joint image understanding, editing, and visual reasoning. The authors also present three proposed corpora: JEdit-1M, JMaze-200K, and JNono-200K, with in- and out-of-distribution benchmarks. The abstract reports best joint performance and monotonic scaling with denoising steps, but provides no numerical results.
No heat snapshots are available in the last 24 hours.