ClinFusion presents a vision-centric medical MLLM built around a compositional cascaded vision encoder and a Cascade Spatial-Aware Locality Fusion operator for unified 2D and native 3D medical-image understanding. The paper also introduces MedIF-Bench for instruction following and a region-of-interest-grounded metric for evaluating factual, clinically aligned report generation. According to the abstract, ClinFusion outperforms leading open-source medical MLLMs on 20 of 24 benchmarks and exceeds GPT-5.2 and Gemini-3-Flash on 13 of 16 benchmarks. Blinded radiologist evaluation reportedly ranks its reports highest.
No heat snapshots are available in the last 24 hours.