Field Converter: Grounding Soccer Pose Estimation in Metric World Space
Original title:Field Converter: Geometry-Initialized Temporal Residual Refinement for World-Grounded Player Pose Estimation from Soccer Broadcasts
Reconstructing 3D player poses in metric world coordinates from calibrated monocular sports broadcasts is often plagued by depth ambiguity. Field Converter tackles this by grounding player roots via ray-pitch intersection before applying a temporal residual refinement model conditioned on visual and geometric cues. Evaluated across unseen matches, this residual approach slashes root localization error from 49 cm to 10 cm using a temporal convolutional network, delivering a world-space MPJPE of 13.2 cm. Ablations reveal that temporal context consistently trumps the choice of network architecture, leaving airborne player movement as the primary remaining failure mode.
Why it's worth reading
Instead of relying on unconstrained global regression, it demonstrates how anchoring monocular sports pose estimation to pitch geometry before refining temporal residuals delivers robust, metric world-grounded tracking.