Constellation was recorded by a single camera mounted high above a busy intersection in Manhattan. Its 13,314 annotated frames span 28 time intervals between 2019 and 2023, covering dawn, daylight, rain, fog, and night, as well as physical changes to the street itself: faded markings, an unpaved surface, and repaving. The training and test sets are separated in time, so no interval appears in both.
The camera is deliberately placed high enough that faces and license plates cannot be resolved. The same privacy-preserving vantage point makes pedestrians very small, which puts the benchmark squarely in the small-object regime where contemporary detectors are weakest.
Findings
- Small pedestrians are the bottleneck. Vehicle AP is close to saturation for most architectures, while pedestrian AP varies by more than 25 points between them.
- Off-the-shelf models transfer poorly. The strongest zero-shot open-vocabulary detector, SAM3, reaches 64.6 pedestrian AP. A YOLOv8x trained on VisDrone reaches 27.3, compared with 87.4 when trained on Constellation.
- Domain-aware training closes the gap. Scene-specific augmentations such as synthetic shadows, combined with VisDrone pretraining, raise YOLOv8x to 92.0 pedestrian AP and 95.4 mAP@0.5.
- Intersections drift over time. A model trained only on 2020 data loses 7.1 mAP@0.5 on 2023 footage, and heavy snow (73.7) and heavy rain (82.7) remain the hardest conditions.
| Model | Training | Pedestrian | Vehicle | Mean |
|---|---|---|---|---|
| SAM3 | Zero-shot | 64.6 | 97.8 | 81.2 |
| YOLOv8x | VisDrone only | 27.3 | 88.6 | 57.9 |
| YOLOv8x | Constellation | 87.4 | 98.6 | 93.0 |
| YOLOv8x | + domain-specific augmentations | 90.7 | 98.6 | 94.7 |
| YOLOv8x | + VisDrone pretraining | 92.0 | 98.8 | 95.4 |
AP@0.5 on the Constellation test set, as reported in the IJCV paper.
Edge deployment
The camera is meant to process video on site, so the paper also measures latency on embedded and mobile hardware. On a Jetson Orin Nano with TensorRT, a VisDrone-pretrained YOLOv8n reaches 94.5 mAP@0.5 at 27.5 ms per frame, within a point of the best server model.
| Platform | Runtime | YOLOv8n | YOLOv8x |
|---|---|---|---|
| NVIDIA A100 | TensorRT | 3.4 ms | 7.1 ms |
| Jetson Orin Nano | TensorRT | 27.5 ms | 111.6 ms |
| Rubik Pi 3 | LiteRT | 55.1 ms | 819.0 ms |
| Raspberry Pi 5 | NCNN | 353.2 ms | 2,325.7 ms |
| iPhone 13 Pro Max | WebAssembly | 372.6 ms | 10,233.1 ms |
Mean inference latency per frame over 100 Constellation images.
Release
The dataset is available in YOLO format on Hugging Face under CC BY-NC-SA 3.0, together with the training and evaluation scripts and a model zoo of seven pretrained detectors, including the 95.4 mAP YOLOv8x and the edge-ready YOLOv8n. The benchmark toolkit reproduces the latency measurements across GPU, embedded, and mobile runtimes.