What our annotators changed on 5,000 frames
Our open dataset runs on machine labels. We say so on the card, loudly. But at some point you have to find out how wrong the machine is, and the only honest way to do that is to put people in front of the frames.
So we did. Annotators on our own platform reviewed 5,000 frames, each one starting from the model's boxes. They could keep a box, fix it, delete it, or draw one the model never saw.
They drew a lot. By the end, 53% of the final boxes came from people, not the model. Read that again: the model missed more than half of what a careful person marks on a Delhi road frame.
They also deleted 12% of the model's boxes as plain wrong. Things that weren't there, or weren't what the model thought.
Night changes the picture. Annotators mark 3.1 boxes per frame at night and 7.1 in daylight.
And the pace? A median of 17 seconds per frame. That one number decides what the next phase costs, so we watch it closely.
A few calls we made on purpose: e-rickshaws, cycle-rickshaws and mini-trucks each get their own variant, and riders get a box separate from their vehicle. They behave differently in traffic. A model that lumps them together can't plan around them.
Next, we take this process toward 100,000 frames, with people reviewing every single box.