MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
AI Summary
MPIE-Bench evaluates multi-person image editing in contact-heavy actions such as embracing, carrying, and grappling, where fused limbs, missing extremities, and body interpenetration remain common. The benchmark contains 2,500 video-mined editing triplets across 405 scenes, 14 interaction categories, and four contact-density levels. MPIE-Eval introduces mesh-based Anatomy and Interaction axes using a frozen public multi-person mesh reconstruction. Across ten editors, the best reported scores reach only 0.65 for Anatomy and 0.72 for Interaction, while VLM checklist scores exceed 0.95. A five-rater study reports closer alignment with human judgments.
Why it's worth reading
Multi-person editors can receive near-perfect VLM checklist scores while producing visibly fused or interpenetrating bodies; MPIE-Bench offers a geometry-aware way to measure that gap.
Deep Read
1. What happened
Original facts: The paper introduces MPIE-Bench and MPIE-Eval for evaluating text-to-image and personalized editing systems in multi-person contact interactions. The benchmark contains 2,500 video-mined editing triplets covering 405 scenes, 14 interaction categories, and four contact-density levels, C0 through C3.
2. Core tech
Original facts: MPIE-Eval uses a frozen public multi-person mesh reconstruction model and separates evaluation into two axes. Anatomy asks whether every human-like mass can be explained by a complete reconstructed body. Interaction measures whether body penetration and surface distance match the contact requested by the instruction.
Analysis: This separates human-form plausibility from relational contact geometry, addressing errors that global similarity scores and VLM checklists may miss.
3. Key evidence and numbers
Original facts: Across ten editors, the reported mesh scores reach at most 0.65 for Anatomy and 0.72 for Interaction, with the two maxima coming from different models. No single editor is strong on both axes. VLM checklist scores for the same images exceed 0.95. A five-rater study reports that both mesh axes align more closely with human judgment than a zero-shot VLM judge, and rankings remain stable when every weight and threshold is ablated.
4. Why it matters
Analysis: Multi-person interaction is a difficult failure mode for image editing: identities may be preserved while limb ownership, contact points, and occlusion relations are wrong. The reported score gap suggests that semantic VLM evaluation can substantially overestimate quality when anatomical and geometric defects are visually obvious to people.
5. Practical impact
Analysis: Researchers can use the benchmark to compare editors across interaction types and contact densities, then use Anatomy and Interaction as diagnostics for data design, post-training filtering, and model iteration. Product teams could adapt similar geometry checks for generated embraces, carrying actions, grappling, and other contact-heavy outputs.
6. Limitations and uncertainty
Original facts: The abstract specifies a frozen public mesh reconstruction model, five human raters, and ten editors, but does not provide the full editor list, data splits, metric formulas, or per-category results.
Analysis: Mesh reconstruction may itself be affected by occlusion, viewpoint, clothing, and unusual poses. Surface distance is also not necessarily equivalent to action semantics or physical plausibility. The five-rater study provides evidence of closer alignment, but does not establish universal validity across users or applications.
Unverified inference: The abstract alone cannot establish whether MPIE-Bench will become a community standard or how well its metrics transfer to video editing, 3D generation, or real photographic inputs.
7. Original sources
- arXiv abstract page
- Source label:
hf-papers - Publication timestamp:
2026-07-31T04:00:00.000Z