This paper argues that standard Direct Preference Optimization (DPO) supervises multimodal language models mainly at the final-answer level, leaving early grounding errors insufficiently corrected and allowing them to propagate through later reasoning. It proposes Grounded Context Preference Optimization (Groc-PO), together with the Grounded Context Preference Dataset (GCPD). GCPD organizes preference samples across three stages: object grounding, contextual grounding, and grounded reasoning. According to the abstract, experiments report improvements over standard DPO and other strong baselines in hallucination mitigation, faithful reasoning, and overall reliability. The abstract does not provide model names, benchmark names, result tables, or effect sizes.
No heat snapshots are available in the last 24 hours.