Data source
Data from HuggingFace, delivered clean.
6 datasets pulled from HuggingFace's releases, checked field by field and shipped the way you want them — daily, weekly, or hourly, your call.
- 6 datasets
- 2 industrys
- Real rows on request
Automobile Manufacturers
3Stanford Cars Dataset (HuggingFace)
NHTSA vPIC Vehicle Manufacturer and VIN API Data
EPA Fuel Economy Dataset (1984-2026)
Rail Transportation
3HuggingFace Railway Datasets
Eurostat Railway Freight Transport Statistics
World Bank Railways Goods Transported Indicator
Pick a catch, see the rows.
Name any HuggingFace dataset and we send real rows from it — not a screenshot of rows. 1,744 datasets. Pick your catch.
Get a sampleAPI, files, or your warehouse. Daily, weekly, or hourly.
Straight answers about HuggingFace data
How many HuggingFace datasets does Datadory deliver?
Two, across two industries. Under Automobile Manufacturers sits the Stanford Cars image-classification benchmark - 72,472 labeled car images across 196 make-model-year classes. Under Rail Transportation sits a verified shelf of exactly 27 railway collections, from complaint tweets to LiDAR point clouds. Each ships with its own product page, field dictionary and sample rows.
What makes the Stanford Cars corpus useful for robustness testing?
Beyond the canonical split of 8,144 training and 8,041 test images, this edition folds in seven corrupted variants of the test set - contrast, gaussian noise, impulse noise, JPEG compression, motion blur, pixelate and spatter, 8,041 images each. That brings the total to 72,472 rows, so degradation benchmarks run on identical class definitions.
How fresh is HuggingFace data through Datadory?
Your delivery cadence is your call - daily, weekly, or hourly - regardless of how the upstream material behaves. The Stanford Cars corpus is a fixed benchmark whose photography runs circa 1991-2012; the railway collections are static snapshots uploaded between 2022 and 2025. Deliveries stay consistent whatever the cadence you pick.
Can deliveries be cut by manufacturer or modality?
Yes. On the automotive side, class subsets filter cleanly by maker, model, body style or year band because every row carries one 196-way label. On the railway side, scoping by modality - text, imagery, point clouds or tabular statistics - is a standard cut. Specify the scope at request time and every recurring pull applies it identically.