DPNeXt targets the decoder bottleneck in ViT-based multi-task dense prediction for robotics perception. It combines a lightweight multi-scale fusion decoder, dual depthwise-separable inverted bottlenecks, task-specific modularization, and Multi-Task Boundary Guidance (MTBG). According to the paper, DPNeXt-S and DPNeXt-B achieve leading or best reported results among compared methods on Cityscapes, while DPNeXt-B also leads semantic segmentation and depth estimation on NYUv2. DPNeXt-S reduces trainable parameters by 78.6% relative to standard DPT and is reported to have the fastest inference on resource-constrained laptop hardware.
No heat snapshots are available in the last 24 hours.