Surround-view camera rigs on autonomous vehicles typically feature minimal overlap between adjacent frames, forcing models to infer depth largely from isolated monocular appearance cues. CrossDepth addresses the resulting boundary inconsistencies by combining per-pixel, camera-aware ray embeddings with geometry-constrained cross-image attention grounded in rig calibration. Operating purely under self-supervised photometric loss, the architecture demonstrates consistent improvements across the DDAD and nuScenes benchmarks, mitigating perspective discrepancies across cameras without requiring manual depth supervision.
No heat snapshots are available in the last 24 hours.