DA3’s biggest flex isn’t accuracy, it’s universality. One model handles monocular depth, multi-view geometry, pose estimation, and even 3D Gaussian splatting without changing the backbone. Just adaptive cross-view attention + a dual-DPT head. This is what “simple architecture,
DA3’s universality: one model handles depth, geometry, pose, and splatting
By
–
