Data source

Data from Google Play, delivered clean.

3 datasets pulled from Google Play's releases, checked field by field and shipped the way you want them — daily, weekly, or hourly, your call.

  • 3 datasets
  • 1 industry
  • Real rows on request

What Datadory delivers from Google Play

3
Application Software Google Play Store listings · Static snapshot scraped February 2019 (versio…

Google Play Store Apps (Kaggle - lava18)

Application Software Google Play Store listings (English-language… · Static snapshot collected June 2021

Google Play Store Apps - Extended (Kaggle - gauthamp10)

Application Software Global app listings from Google Play and the…

Play and Apple App Store Metrics (Hugging Face - appGoblin)

Pick a catch, see the rows.

Name any Google Play dataset and we send real rows from it — not a screenshot of rows. 1,744 datasets. Pick your catch.

Get a sample

API, files, or your warehouse. Daily, weekly, or hourly.

Straight answers about Google Play data

What does Datadory deliver from the Google Play source?

One application-software dataset, scoring 8 out of 10 in our catalog: a fixed February 2019 snapshot of 10,841 Android store listings across thirteen documented fields, plus a companion file of 64,295 sentiment-scored user reviews covering 1,074 of those apps.

How many apps and reviews are covered?

10,841 app rows with thirteen fields each, and 64,295 translated review rows across 1,074 apps - 75,136 combined rows. Each app row preserves its own Last Updated date at capture time, so historical listing ages remain computable inside the frozen month.

How clean is the data out of the box?

Famous but imperfect: one row records Category as 1.9, some ratings are missing, duplicate app names exist and Installs mixes types. Datadory runs the deduplication and type-coercion pass before delivery, so those quirks are resolved upstream instead of being discovered mid-analysis.

Does the review file include sentiment labels ready for modeling?

Every one of the 64,295 reviews carries three precomputed columns: a Positive/Negative/Neutral label, polarity from -1 to 1 and subjectivity from 0 to 1. That makes it an instant benchmark corpus - your model scores against labels produced entirely outside your own pipeline.