Data source

Data from Hugging Face, delivered clean.

5 datasets pulled from Hugging Face's releases, checked field by field and shipped the way you want them — daily, weekly, or hourly, your call.

  • 5 datasets
  • 4 industrys
  • Real rows on request

Application Software

2
Application Software Global app listings from Google Play and the…

Play and Apple App Store Metrics (Hugging Face - appGoblin)

Application Software United States · Static snapshot collected December 2023 - Jan…

Mac App Store Apps Metadata (Hugging Face - macpaw-research)

Publishing

1
Publishing Global contributor base · Static snapshots pinned per git revision

Hugging Face Datasets Hub - Books Search

Apparel, Accessories & Luxury Goods

1
Apparel, Accessories & Luxury Goods Global · Continuously growing repository

Hugging Face Datasets Hub - Fashion/Apparel Datasets

Electronic Components

1
Electronic Components Global distributor coverage (DigiKey · Point-in-time snapshot of 791 records

Electronic Components Supply Chain & Risk Dataset

Pick a catch, see the rows.

Name any Hugging Face dataset and we send real rows from it — not a screenshot of rows. 1,744 datasets. Pick your catch.

Get a sample

API, files, or your warehouse. Daily, weekly, or hourly.

Straight answers about Hugging Face data

What does Datadory deliver from Hugging Face?

Twenty-seven cataloged collections across 24 industries, averaging 6.48 out of 10 in Datadory's quality system against a 7.81 catalog-wide mean. The standouts: a 3,194,924-row app store metrics table scoring 10, a 61,196-app Mac App Store snapshot scoring 9, and roughly 770 book corpora led by 140,000 Library of Congress volumes. Each ships with a verified field dictionary, worked examples and a coverage statement.

How large is the Hugging Face shelf overall?

The widest single measurement in the catalog: one record counts 1,012,353 public dataset repositories hub-wide and another, researched separately, 1,012,419 - the gap is just growth between two snapshots. Datadory's 27 records are the curated slices: 280 field definitions, 88 verified example rows on file, and sizes ranging from two audio clips to billions of words.

Are all 27 collections the same quality?

No, and pretending otherwise would be the dishonest sell. Scores run from 4 to 10: the app store metrics table earns a 10 with fully populated keyed columns, while the two-file Patois audio set and a one-photo sneaker set earn 4s for empty cards and unstated provenance. Your sample report states each record's score and exactly what documentation is missing before you commit pipeline time.

Can I get several industries in one delivery?

Yes. Because all 27 come through one contract, an application software cut joins cleanly against a publishing or tobacco cut in the same envelope - same normalization, same validation rows, one cadence either way. Name the industries and the sample ships shaped that way, rows and dictionaries included.