Reconstructing 3D player poses in metric world coordinates from calibrated monocular sports broadcasts is often plagued by depth ambiguity. Field Converter tackles this by grounding player roots via ray-pitch intersection before applying a temporal residual refinement model conditioned on visual and geometric cues. Evaluated across unseen matches, this residual approach slashes root localization error from 49 cm to 10 cm using a temporal convolutional network, delivering a world-space MPJPE of 13.2 cm. Ablations reveal that temporal context consistently trumps the choice of network architecture, leaving airborne player movement as the primary remaining failure mode.
There are 8 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 08:00; latest heat is 0.