The paper introduces BRAID, a reinforcement-learning framework that formulates interleaved text-image-text reasoning as a unified Markov decision process. Instead of optimizing text with RL while treating image generation as a supervised surrogate, BRAID uses a shared trajectory-level advantage for both text tokens and image denoising paths, with modality-native policy gradients. A vision-language-model judge additionally scores intermediate images according to their utility for reasoning, providing denser turn-level feedback for long-horizon credit assignment. The authors report consistent improvements over multiple baselines on spatial reasoning and visual perception benchmarks, although the abstract does not provide numerical results or implementation details.
No heat snapshots are available in the last 24 hours.