Benchmark · preview

Off-the-shelf detectors lose about 40% of their accuracy on Indian roads.

We ran 20 widely used object detectors, unchanged, on 904 dashcam frames from Delhi NCR that two of our annotators labelled independently. The best model scores 34.7. On COCO, the benchmark these models are built against, it scores 57.5.

59%
of their COCO score is what these models keep on Indian roads. YOLO26x, the best here, scores 34.7 against 57.5 on COCO.
50%
of auto-rickshaws go unseen by YOLO26x. No COCO model has a class for them; the ones it does find, it calls trucks or cars.
7.1 pts
lost at night, on average across all 20 models, and 4.8 in rain, against daytime frames.
+2.1 pts
is all that six years of YOLO releases add here, from YOLOv5x to YOLO26x. On COCO the same step is worth 4.3.

Leaderboard

Each model's score on our test set, out of 100: the standard COCO measure of how well its boxes match ours. Differences under one point are a tie.

#ModelScore on Indian roadsKeeps of its COCO scoreAt nightAuto-rickshaws found
1YOLO26x34.760%30.950%
2YOLOv8x33.863%30.054%
3YOLO12x33.661%30.251%
4YOLO11x33.561%29.853%
5YOLO26m32.862%28.154%
6YOLOv5x32.661%28.352%
7YOLOv9e32.558%28.151%
8YOLOv10m31.962%27.849%
9YOLOv9c31.860%28.651%
10RT-DETR-X31.658%29.875%
11YOLO12m31.460%28.250%
12YOLO11m31.361%27.853%
13YOLOv10x31.257%29.249%
14RT-DETR-L30.958%28.375%
15YOLOv8m30.962%27.552%
16Faster R-CNN v230.565%25.874%
17RetinaNet v228.268%23.566%
18YOLOv5s23.755%22.244%
19YOLO11n18.647%17.340%
20YOLOv8n18.249%17.438%
All results, every column
Score
COCO-style mAP, IoU 0.50 to 0.95, on the eight classes a COCO model can name.
COCO
The model's published score on COCO, its home benchmark.
Kept
Our score as a share of its COCO score.
AP50
A looser score: a box counts if it overlaps ours by half.
Near
Only objects at least 32 pixels tall, the closer ones.
Conditions
Checked by hand on every frame; the number under each is how many test frames it covers.
#ModelScoreCOCOKeptAP50NearDay320 framesDusk or dawn119 framesNight465 framesRain110 framesAutos found
1YOLO26x34.757.560%58.135.539.536.730.935.750%
2YOLOv8x33.853.963%56.134.838.234.430.032.354%
3YOLO12x33.655.261%55.434.637.335.930.234.651%
4YOLO11x33.554.761%54.734.537.635.029.834.053%
5YOLO26m32.853.162%53.533.737.535.628.128.954%
6YOLOv5x32.653.261%54.133.637.534.228.332.452%
7YOLOv9e32.555.658%55.333.237.935.928.132.451%
8YOLOv10m31.951.162%52.532.836.933.627.831.849%
9YOLOv9c31.853.060%51.832.835.034.128.628.651%
10RT-DETR-X31.654.858%53.832.634.234.129.832.675%
11YOLO12m31.452.560%51.532.334.735.228.232.850%
12YOLO11m31.351.561%51.432.235.433.127.829.153%
13YOLOv10x31.254.457%51.432.133.931.329.227.249%
14RT-DETR-L30.953.058%53.631.734.232.828.330.075%
15YOLOv8m30.950.262%51.531.835.031.827.530.152%
16Faster R-CNN v230.546.765%55.831.635.833.325.828.274%
17RetinaNet v228.241.568%50.429.134.428.223.526.566%
18YOLOv5s23.743.055%39.224.425.723.922.221.544%
19YOLO11n18.639.547%31.619.121.117.117.319.240%
20YOLOv8n18.237.349%30.918.819.918.117.417.238%

How far each model falls

The same models, on COCO and on Indian roads. The longer the line, the more accuracy a model loses when it leaves the benchmark it was built for.

  • On COCO
  • On Indian roads
  1. YOLO26x
    34.7 / 57.5
  2. YOLOv8x
    33.8 / 53.9
  3. YOLO12x
    33.6 / 55.2
  4. YOLO11x
    33.5 / 54.7
  5. YOLO26m
    32.8 / 53.1
  6. YOLOv5x
    32.6 / 53.2
  7. YOLOv9e
    32.5 / 55.6
  8. YOLOv10m
    31.9 / 51.1
  9. YOLOv9c
    31.8 / 53.0
  10. RT-DETR-X
    31.6 / 54.8
  11. YOLO12m
    31.4 / 52.5
  12. YOLO11m
    31.3 / 51.5
  13. YOLOv10x
    31.2 / 54.4
  14. RT-DETR-L
    30.9 / 53.0
  15. YOLOv8m
    30.9 / 50.2
  16. Faster R-CNN v2
    30.5 / 46.7
  17. RetinaNet v2
    28.2 / 41.5
  18. YOLOv5s
    23.7 / 43.0
  19. YOLO11n
    18.6 / 39.5
  20. YOLOv8n
    18.2 / 37.3
Score from 0 to 100 (COCO-style mAP). Hover or focus a row for its two scores.

Accuracy by condition

Night is where every model struggles most. Rain sits between day and night: 84 of our 110 rain frames were filmed at night, so that bar mixes the two.

  1. Day320 frames
    34.1
  2. Dusk or dawn119 frames
    31.7
  3. Night465 frames
    26.9
  4. Rain110 frames
    29.3
Average score of all models in each condition, out of 100.

What models call an auto-rickshaw

COCO has no auto-rickshaw class, so we count a hit when a model puts any vehicle box on one. Across all 2,511 auto-rickshaws in our frames, this is where they end up.

YOLO26x

YOLOv8x

RT-DETR-X

Faster R-CNN v2

  • Missed
  • Truck
  • Car
  • Bus

From the test set

12 frames from the 904, across the conditions we score. Faces, number plates and numbers painted on vehicles are blurred.

Delhi NCR dashcam frame: Day, clear
Day · clear
Delhi NCR dashcam frame: Day, clear
Day · clear
Delhi NCR dashcam frame: Day, clear
Day · clear
Delhi NCR dashcam frame: Day, clear
Day · clear
Delhi NCR dashcam frame: Day, clear
Day · clear
Delhi NCR dashcam frame: Dusk, clear
Dusk · clear
Delhi NCR dashcam frame: Dusk, clear
Dusk · clear
Delhi NCR dashcam frame: Night, clear
Night · clear
Delhi NCR dashcam frame: Night, clear
Night · clear
Delhi NCR dashcam frame: Night, clear
Night · clear
Delhi NCR dashcam frame: Night, clear
Night · clear
Delhi NCR dashcam frame: Night, rain
Night · rain

How we scored it

Test set. 904 frames from Delhi NCR dashcams: 320 by day, 119 at dusk or dawn and 465 at night, 110 of them in rain. A person checked each frame's time of day and weather by hand, and we left out the few frames with no usable picture. Two annotators labelled each frame without seeing the other's work. A box both of them drew counts as ground truth (5,207 boxes). A box only one of them drew is left out of scoring until someone adjudicates it, so a model is neither rewarded nor punished there.

Classes. Person, bicycle, car, motorcycle, bus, truck, traffic light and animal. A COCO "person" counts for our riders too, since COCO has no rider class. Traffic signs are left out: COCO only knows stop signs.

Models. Each runs with its published COCO weights and its default input size, keeping up to 100 boxes per frame above 1% confidence. Nothing is fine-tuned on our data. Near counts only objects at least 32 pixels tall.

Privacy. We blur faces, number plates and numbers painted on vehicles in every frame we publish.

Preview. The adjudicated test set, where someone settles every box only one annotator drew, will follow. Scores may move by a point or two.

Want your model on the board, or a private evaluation on the conditions you ship into?

Talk to us