UrbanOmniDetect targets a common deployment bottleneck in V2X and infrastructure sensing: camera intrinsics may be unavailable, imprecise, or drifting. Instead of lifting 2D detections through a calibrated camera model, a single network predicts the eight projected corners of each object’s 3D box directly from a raw RGB image, with no intrinsics, depth estimation, or ground-plane priors.
The work was presented as an oral paper at the CVPR 2026 DriveX workshop and is paired with UrbanOmniView, a dataset that combines real-world driving data from KITTI, infrastructure camera data from DAIR-V2X, and high-fidelity synthetic data rendered in Unreal Engine 5, which is released as part of the project.
Results
- Invariant to calibration error. Calibration-dependent methods lose more than 80% of their accuracy with a 5% focal-length error. UrbanOmniDetect takes no camera intrinsics as input and is invariant to such errors by construction.
- Strongest on the harder KITTI splits. Without calibration, it reaches 30.71 AP3D and 35.19 APBEV on the Moderate split at IoU ≥ 0.7, ahead of the calibration-dependent baselines below on Moderate and Hard.
- Real-time inference. A single forward pass takes under 11 ms per image on an A100 with TensorRT at 640 × 640.
| Method | AP3D Easy | AP3D Mod. | AP3D Hard | APBEV Easy | APBEV Mod. | APBEV Hard |
|---|---|---|---|---|---|---|
| MonoDGP | 30.76 | 22.34 | 19.02 | 39.40 | 28.20 | 24.42 |
| MonoCon | 26.33 | 19.01 | 15.98 | 34.65 | 25.39 | 21.93 |
| MonoLSS | 25.91 | 18.29 | 15.94 | 34.70 | 25.36 | 21.84 |
| DEVIANT | 24.63 | 16.54 | 14.52 | 32.60 | 23.04 | 19.99 |
| UrbanOmniDetect | 29.61 | 30.71 | 27.76 | 33.86 | 35.19 | 31.38 |
Monocular 3D detection on KITTI at IoU ≥ 0.7. The baselines use camera calibration. UrbanOmniDetect does not.
UrbanOmniDetect-2
I have since released UrbanOmniDetect-2, my independent follow-up to the paper. It turns the pose-only detector into a single hybrid network for detection and 3D cuboids. One forward pass detects all 80 COCO classes and, for every road user, regresses the eight projected corners of its 3D box, on any viewpoint and still without calibration. The same output feeds tracking, a bird's-eye view, and an offline refinement stage that turns per-frame detections into rigid, physically consistent trajectories.
Training mixes 2D and 3D supervision. COCO and VisDrone ground the detector, while KITTI, DAIR-V2X, CDrone, and rendered vehicles teach the cuboids through a masked pose loss. The release includes five model scales from 2.6M to 57.6M parameters, along with training code, a real-time bird's-eye-view pipeline, and tools for producing labels and demo reels from raw footage.