CrossDepth: Geometry-Constrained Attention for Multi-View Surround Depth Estimation
Original title:CrossDepth: Geometry-Constrained Attention for Generalizable Multi-View Surround Depth Estimation
Surround-view camera rigs on autonomous vehicles typically feature minimal overlap between adjacent frames, forcing models to infer depth largely from isolated monocular appearance cues. CrossDepth addresses the resulting boundary inconsistencies by combining per-pixel, camera-aware ray embeddings with geometry-constrained cross-image attention grounded in rig calibration. Operating purely under self-supervised photometric loss, the architecture demonstrates consistent improvements across the DDAD and nuScenes benchmarks, mitigating perspective discrepancies across cameras without requiring manual depth supervision.
Why it's worth reading
It tackles the persistent seam inconsistencies in surround-view autonomous driving setups using an elegant, self-supervised geometric attention mechanism.