Datadory notebook
Where can I get App Store data? Every listing on both storefronts, delivered as typed rows
Datadory delivers application software data covering both mobile storefronts in a single schema: Play and Apple App Store Metrics unifies 3,194,924 listings - 1,695,746 Google Play apps beside 1,499,178 Apple App Store apps - with price, installs, ratings, category, developer and ad/IAP monetization flags on every row, while Apple iTunes Search holds the live iOS catalog at twenty-nine documented fields per app and the Google Play web catalog documents the 2-million-listing Android shelf. Frozen research vintages add exact Android install ceilings and laptop-scale training sets. Typed, keyed, field-dictionaried, delivered daily, weekly, or hourly.
1,744 datasets. Pick your catch.
What does “app store data” actually point at?
The query sounds like an errand; the honest answer is a stack. App store data is not one table - it is several records that describe the same two storefronts from different altitudes, and picking the wrong altitude is how app-market projects quietly die. Datadory catalogs 24 primary application software datasets plus one adjacent record, and six of them hold the mobile shelf itself.
Play and Apple App Store Metrics (Hugging Face - appGoblin) is the widest table in the slice and the closest thing the industry has to a whole-market census: 3,194,924 rows - 1,695,746 Google Play apps and 1,499,178 Apple App Store apps - described by the same 15 columns, from identifiers and pricing through installs, ratings, lifecycle dates and per-row capture status. Most public corpora pick a side; this one puts both storefronts in one schema, which turns two half-markets into one measurable market.
Beneath the wide table, three research vintages freeze the shelves for modeling work: the 2.3-million-row Google Play Store Apps - Extended with exact install minimums and maximums, the 10,841-app-plus-64,295-labeled-reviews Google Play Store Apps pair, and the 7,197-app Apple App Store - 10k Apps Dataset with its companion description file. Mac App Store Apps Metadata adds the desktop edge: 61,196 macOS apps across US, German and Ukrainian storefronts.
It is a deep shelf. This industry averages 8.4 out of 10 on our quality rubric against a catalog-wide mean of 7.81 across all 1,744 datasets, with eighteen records scoring eight or higher and five perfect tens.
Which dataset fits which job?
Each app-market workflow maps to a specific record rather than a generic category:
What do app store rows look like once delivered?
Rows, not screenshots - and the grain shows immediately. Three rows exactly as they arrive from the cross-store metrics table: one Google, one Apple, one long-tail Google.
Store: Google Store ID: com.fwbpack.byinfinitybook App Name: FW-BPack By InfinityBook Category: productivity Total Installs: 3 Total Ratings: 0 Released: 2026-02-25 Row Refreshed: 2026-03-15 Crawl Result: success
What the three rows already settle: the same fifteen columns describe both storefronts; Google rows resolve install counts while the Apple row leans on its rating volume; and every row carries the date it was refreshed plus whether the capture succeeded, so bad rows announce themselves instead of poisoning your averages.
From the live iOS record, one captured row beats a paragraph of specification:
And from the Play web catalog, the anchor row of the sample set:app_title : Spotify: Music and Podcasts developer_name : Spotify AB package_id : com.spotify.music category : Music & Audio rating : 4.3 review_count : 36.2M downloads_bracket : 1B+ flags : Contains ads | In-app purchases updated_on : Aug 18, 2026 ```
Which fields carry the analytical weight?
Fifteen columns sound thin until you notice which questions they decide. All fields below arrive documented in the field dictionary that ships with every extract.
| Field | What it holds | Why it matters |
|---|---|---|
| Total Installs | Resolved install count on Google rows; a rating-count-derived popularity figure on Apple rows | Market-sizing denominators - segment by store before comparing distributions |
| Ad Supported / In-App Purchases | Declared monetization flags, populated on Google rows | Category-level monetization-mix baselines without touching a single listing |
| Developer ID | Publisher name on Google, artist ID on Apple | Portfolio mapping and acquisition sourcing inside one store |
| Minimum Installs / Maximum Installs (Extended snapshot) | Exact bracket bounds per app, June 2021 | The only record where install precision survives past the public bracket |
| Translated_Review + sentiment label (lava18 companion) | 64,295 pre-scored reviews joined to 1,074 apps | Sentiment baselines that skip the labeling budget entirely |
The money-and-traction columns deserve emphasis: because price, installs, rating volume and both monetization flags sit on the same row as identity and lifecycle dates, a whole-market segmentation is one group-by rather than a stitching project across four exports.
How much ground does the coverage span, and where do the vintages sit?
- Geographic: global listings from both storefronts, collected primarily against English-language storefront views; the live iOS record selects any storefront country by parameter. The Mac corpus adds dedicated US, German and Ukrainian captures.
- Temporal: the cross-store baseline was swept February through March 2026, with per-row refresh stamps running 2026-02-06 to 2026-03-19, while the
release_datevalues inside reach back to September 2008 and forward past today, because storefronts list pre-release apps. The research vintages sit where their collectors left them: June 2021 for the Extended Android snapshot, February 2019 for the lava18 pair, 2017 for the App Store 10k set, December 2023 - January 2024 for the Mac table. Current-state records read the shelf as it stands. - Granularity: one row per app holding its latest observed metrics. There is no time series inside the tables.
Read that plainly: these are storefronts photographed, not panels. Anything that needs movement - growth rates, churn, seasonality - wants a baseline plus a scheduled delivery that accumulates snapshots over time. That pairing is exactly what Datadory assembles: the vintage lands once as a backfill, then fresh cuts rotate in on the rhythm your models expect, each stamped so no trend chart silently mixes February 2026 with June 2021.
One habit separates teams whose app charts survive review from teams whose do not: recording the vintage of every column. Label the snapshot dates in your warehouse on arrival, because a 2021 install ceiling beside a 2026 rating average will plot happily and mean nothing.
What can't app store data tell you?
Four boundaries define the shelf, and named companions fill them.
- Installs do not translate across stores. Google rows resolve public install brackets into numbers; Apple rows substitute a popularity figure derived from rating counts, and the conversion is not published. Compare distributions within a store, never across both - or let us deliver the two sides pre-segmented so the trap never gets built.
- No history inside the tables. Each app appears once with its latest state. Trend work pairs the baseline with a scheduled feed that accumulates snapshots, which is a delivery decision rather than a search problem.
- No usage or revenue telemetry. Installs measure acquisition, not engagement. Similarweb Top Websites & App Intelligence tracks 4 million apps across 58 countries with up to 37 months of history when engagement estimates matter, and StatCounter Global Stats supplies platform-share series back to January 2009 for backtesting platform-shift theses.
- Consumer stores are not the B2B market. For software sold to companies, Capterra Software Directory holds 86,267 product profiles across 1,025 category pages with verified review counts and feature sub-scores, G2 Software Categories & Product Pages adds 2,239 category pages, and Chrome Web Store - Extensions & Apps enumerates roughly 360,000 browser extensions with user-count badges.
None of these caveats hides at delivery: the limitations ship beside the table, and we flag which ones bite your specific question before you commit.
Who builds on app store data?
Market researchers and consultants size categories from ranked, comparable counts instead of chart anecdotes - 3.19 million rows make a defensible denominator for any “how big is this market” slide. The workflow lives at market researchers use cases.
Investors and quant researchers read publisher portfolios, dormancy patterns and monetization mix as diligence signals, and pair adoption curves with Similarweb app intelligence for engagement. See investors quants use cases.
Data scientists and ML engineers get pre-split baselines: millions of labeled rows for rating prediction, sixty-four thousand pre-scored reviews for sentiment, and full marketing text for embeddings. See data scientists use cases.
How is app store data delivered?
API, files, or straight into your warehouse. Daily, weekly, or hourly - your call. You pick the slice and the shape: flat files for analysts, a feed for running pipelines, landed tables keyed on store ID so both storefronts join cleanly beside anything else in your warehouse.
Delivery absorbs quirks you would otherwise meet one notebook at a time. Storefront-native category slugs arrive mapped so a cross-store group-by does not turn into a taxonomy project. The two install semantics arrive pre-segmented by store. Failed captures arrive flagged rather than silently averaged in. Historical depth lands once as a backfill, then fresh cuts rotate in as new vintages against keyed records instead of overwrites.
Start with a sample: name your categories, stores, install bands, price models or developer sets, and the extract comes back cut to that scope with all fifteen columns intact and the field dictionary attached. Get a sample of app store data - the schema in the sample is the schema you ship against.
Where to go next
Start with the anchors: the Play and Apple App Store Metrics whole-market table, the live Apple iTunes Search and Google Play Store web catalog records, and the three research vintages - Play Extended, Play Store Apps and App Store 10k - plus the desktop Mac App Store metadata corpus.
For the whole slice, the application software data guide maps all 25 pooled records with quality scores and use cases, and the application-software data hub browses the same catalog as products with sample rows and field dictionaries. When you want the ranked form, the best application software datasets list scores the top eight side by side, and the Apple iTunes Search vs Google Play Store web catalog comparison settles which live record answers which question.
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
Play and Apple App Store Metrics (Hugging Face - appGoblin)
15 per row · shared across both storefronts …+12 more
Apple iTunes Search API
Google Play Store (Web Catalog)
privacy_policy_url · similar_apps
Google Play Store Apps - Extended (Kaggle - gauthamp10)
24 per row …+21 more
Apple App Store - 10k Apps Dataset (Kaggle - ramamet4)
currency
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
What is the largest app store dataset?
Play and Apple App Store Metrics (Hugging Face - appGoblin): 3,194,924 rows in one table - 1,695,746 Google Play apps and 1,499,178 Apple App Store apps - each described by the same 15 columns covering price, installs, ratings, category, developer, lifecycle dates and per-row capture status.
Does one dataset really cover both Google Play and the Apple App Store?
Yes, and that is its entire value. Both storefronts share one schema, so a category count, a pricing band or a monetization split runs as a single group-by across 3.19 million rows instead of two reconciled exports. Cross-store joins key on the store identifier column; within-store publisher grouping uses the developer column.
Are install counts comparable between the two stores?
No. On Google rows the install figure is resolved from the storefront's public bracket. On Apple rows it is a popularity figure derived from rating counts, and the exact conversion is not published. Segment by store before comparing distributions - Datadory delivers the two sides pre-segmented so the mismatch never reaches your charts.