All updates

We scored 20 object detectors on Delhi roads

Take a frame from a Delhi evening. An auto-rickshaw cuts across two lanes, three people share one scooter, and a truck rolls past with SLOW HORN painted on its tailboard. Now ask a modern object detector what it sees.

We asked 20 of them, 904 times each. YOLOv5 through YOLO26, RT-DETR, Faster R-CNN and RetinaNet, all with their published COCO weights and nothing fine-tuned. Then we scored every box they drew against frames our own annotators labelled by hand.

The best model, YOLO26x, scores 34.7. On COCO, the benchmark it was built against, it scores 57.5. Across all 20, the models keep 59% of their COCO score on average, and none keeps more than 68%.

Auto-rickshaws hurt the most. COCO has no class for them, so we gave every model the benefit of the doubt and counted any vehicle box on an auto as a hit. YOLO26x still misses half of the 2,511 autos in our frames. A quarter of them it calls trucks, and about a fifth it calls cars. RT-DETR and Faster R-CNN find about three in four, but they still have to call them something else.

Night costs every model, 7.1 points on average against daytime frames. We built the test set heavy on night on purpose: 465 of the 904 frames. Rain costs 4.8 points, but read that one carefully. 84 of our 110 rain frames are also night frames, so the rain number mixes two problems.

The finding that stuck with us: going from YOLOv5x to YOLO26x, six years of releases, adds 4.3 points on COCO. On our roads, the same step adds 2.1. And 16 of the 20 models score between 30.5 and 34.7 here, while on COCO those same 16 run from 46.7 to 57.5. The progress is real. It just doesn't all travel.

The small models fall furthest. The nano versions, the ones meant for edge devices, keep under half their COCO score. YOLO11n scores 39.5 on COCO and 18.6 here.

Every number above rests on the labels, so here is how we made them. Two annotators labelled each frame on their own, without seeing the other's boxes. They agree on about two-thirds of the boxes. That sounds low until you look at a Delhi night frame: a rider half hidden behind a bus, a scooter parked in shadow, a cluster of headlights that could be two vehicles or three.

So a box counts as ground truth only when both annotators drew it, 5,207 boxes in all. A box only one of them drew becomes an ignore region. A model that finds it gains nothing, and a model that misses it loses nothing.

We also checked every frame's time of day and weather by hand, because the machine tags we started from got it wrong more often than we expected. Plenty of frames tagged as dusk were plainly night. We left out the few frames with no usable picture, and before anything went public we blurred faces, number plates and the numbers people paint on trucks.

Nobody ships a COCO model unchanged, and that's the point of running them that way. It's the starting line every team fine-tunes from. The gap between 57.5 and 34.7 shows how much ground Indian road data has to cover.

This is a preview. Next, someone settles every box only one annotator drew, and scores may move by a point or two. We'd also like to add the models people actually run. If you have one, public or private, we'd be glad to test it, and if you need an evaluation on the roads and conditions you ship into, that's work we do.

See the full leaderboard