DART turns a promptable frontier vision model into a real-time multi-class open-vocabulary detector. The project targets a practical deployment gap: Strong promptable segmentation models can describe almost anything, but repeated per-class inference is too slow for many real-world systems.
How it works
SAM3 runs its ViT-H backbone once for every class it is asked to find. DART runs the backbone once per image, caches a text embedding for each prompt, and decodes all prompts together in a single batched pass through the encoder-decoder. Both stages are deployed as TensorRT FP16 engines. On video, the backbone encodes the next frame while the current one is decoded.
Results
- 55.8 AP on COCO val2017 across all 80 classes, with no training: DART uses the SAM3 weights as released.
- 15.8 FPS with four classes at 1008 px on a single RTX 4080, against 13.8 FPS when the two stages run back to back.
- Distilled student backbones for tighter budgets. ViT-H pruned to 16 blocks keeps 53.6 AP with a 26.6 ms backbone, and RepViT-M2.3 reaches 38.7 AP with an 8.2M-parameter backbone that runs in 13.9 ms.
The public release includes code, benchmarks, TensorRT deployment paths, distilled student backbones, and Hugging Face weights.