From SD card to dataset: how our pipeline works
People ask how a folder of dashcam video turns into a dataset. It's less magic than it sounds, and more plumbing.
It starts with SD cards. We pull footage off the cameras with a small command-line tool we wrote around rclone, which pushes everything to cloud storage in parallel. Boring, fast, and it means nobody drags files around by hand.
From there, a worker picks up each new folder and runs every clip through a chain of stages. It reads the GPS track, pulls one keyframe per second, then runs detection, segmentation, tracking and scene classification on those frames.
The models: YOLO11m for detection and SegFormer-B5 for segmentation, both fine-tuned on IDD, the Indian Driving Dataset from IIIT Hyderabad. ByteTrack links boxes across frames so each object keeps an ID. CLIP tags each clip's weather, time of day and scene type.
Results land in Postgres with PostGIS, so we can ask questions like which clips ran through a given area at night. The pipeline is cumulative. Drop in a new folder and it adds to the same database without redoing old work.
When we want a release, one export step writes everything out in BDD100K format. That's how 300+ hours of raw 1080p video became the 645,714 frames on Hugging Face.
The weak link is the labels. Models fine-tuned on someone else's cities miss a lot on ours, which is why the next chapter is about people, not models.