This paper introduces CMDR and CMDR-Bench for multimodal document retrieval queries that require aggregating information across multiple pages, rather than relying only on lexical or semantic matching. It proposes CMDR-Embed, which jointly encodes multiple pages and derives page-level embeddings from a shared contextual representation. The paper also presents CMCL, a contextual multimodal contrastive learning objective designed to balance document-context modeling with page-level discriminability. According to the abstract, experiments show that CMDR-Embed substantially outperforms non-contextual embeddings, although detailed benchmark sizes and metric values are not provided in the supplied summary.
No heat snapshots are available in the last 24 hours.