Interactive Home Entertainment · Steam Store

Steam Store - Scrapeable Catalog

Datadory delivers steam store scrapeable catalog data covering 261,285 storefront applications as of August 2026 - every record carrying title, full and short descriptions, developer and publisher, genres, categories and tags, platform availability, system requirements, localized prices with nested discount packages, achievement totals and review-score summaries, from 1990s releases through pre-orders dated years ahead.

API, files, or your warehouse. Daily, weekly, or hourly.

Where it covers
Global storefront - pricing localized by country across all supported currencies
How far back
Live catalog reflecting current prices and discounts; release dates span 1990s titles through pre-orders dated years ahead
How fine
One record per application - games, DLC, demos, software - with purchase packages and price tiers nested

What is the Steam Store - Scrapeable Catalog dataset?

Steam Store - Scrapeable Catalog is the Interactive Home Entertainment catalog's commerce layer: the 261,285 applications searchable on Valve's PC storefront as of August 2026, each carried as a structured record rather than a product page. Games lead the set, but one schema holds for games, DLC, demos and software, which is what makes it a market census instead of a bestseller list.

A single row is commerce-grade in a way storefront mirrors rarely manage: full marketing copy in two lengths, developer and publisher, genre and category flags, platform availability, minimum and recommended requirements for all three desktop operating systems, supported languages, achievement totals, content ratings, support contacts, and a release field separating shipped titles from pre-orders dated years ahead. Pricing arrives as nested purchase packages with original and discounted amounts, localized by country across all supported currencies.

Get a sample of this dataset - name a genre, price band or platform, and real rows come back before you commit to a feed.

What does a sample row look like?

Four records from the current catalog, shown exactly as they arrive:

type        : game                   name      : Dota 2
steam_appid : 570                    is_free   : true
developers  : ["Valve"]              publishers: ["Valve"]
platforms   : windows, mac, linux    dlc       : present

review_score_desc : Very Positive    review_score : 8
total_positive    : 5430             total_negative : 689
total_reviews     : 6119

title       : Counter-Strike 2       listing   : Free

title       : Black Myth: Wukong     discount  : -30%
original    : 3,599.00 (INR)         discounted: 2,519.00 (INR)

Three things worth noticing. First, identifier discipline: steam_appid: 570 is the canonical key for Dota 2, and it is the join point to player-count, ownership-estimate and sales-chart sets elsewhere in this catalog. Second, reputation rides along with commerce - the review block pairs a readable label (Very Positive) with its own arithmetic: 5,430 positive against 689 negative of 6,119 recent reviews. Third, the Black Myth: Wukong row is a discount event caught mid-flight - a 30% cut takes the listing from 3,599.00 to 2,519.00 INR, both amounts preserved so discount depth is computable rather than inferred.

What fields does the dataset include?

Thirty verified columns per application, defined below. Three groups do most of the analytical work.

Identity and classification: steam_appid, name, type, developers, publishers, genres, categories. The category vocabulary is behavioral rather than thematic - Single-player, Multi-player, Remote Play, Steam Achievements - the features a buyer actually receives.

Commerce: is_free, package_groups and the platform booleans. Purchase options nest as packages with sub-tiers priced in cents, preserving bundle structure instead of flattening everything to one number.

Reception and logistics: recommendations, achievements, metacritic, ratings, review_score_desc, total_reviews, plus supported languages and per-OS requirements. The long-form text columns (detailed_description, about_the_game) carry HTML, ready to be parsed into clean text corpora.

Sparsity is informative here: optional attributes populate only when they apply. Metacritic scores appear where a score exists, content ratings where a board issued one. Absence is signal about the title, not a gap in the delivery.

What does coverage look like across geography, time and granularity?

Geography: one global storefront. Pricing localizes by country across every supported currency, so the same application carries comparable price points market by market - the substrate for regional-pricing and purchasing-power work without assembling country mirrors yourself.

Time: records reflect live prices and discounts, while release_date reaches from 1990s titles still on sale through pre-orders dated years ahead. Snapshot cadence is your call - daily, weekly or hourly - and every snapshot lands in the same comparable rows.

Granularity: one record per application - games, DLC, demos, software - cuttable by type, genre, category, platform, price band and review tier. Because packages nest inside each record, a discount-depth cut composes cleanly with a genre cut or a geography cut.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly. Pick the channel that fits the workload: an API for tools wanting current rows on demand, files for batch analysis, or direct delivery into your warehouse for teams running joins against their own tables. Cadence is yours too - hourly suits discount tracking through a sale season, daily suits standing market reviews.

Start smaller than a feed. Get a sample of this dataset scoped to the genres, price bands or platforms you care about, confirm the rows answer your question, then scale the cadence.

Who uses this data, and for what?

Price and promotion intelligence. Original and discounted amounts sit side by side in every record, so discount frequency, depth and timing across 261,285 applications become queries rather than manual collection. Publisher pricing teams benchmark their promo calendars against the entire storefront.

Catalog and assortment research. Type, genre, category and platform flags turn the storefront into a supply census: how many co-op titles shipped this quarter, how the Linux-available share moves, where DLC clusters around franchises.

Content-based machine learning. Two lengths of marketing copy, a behavioral category vocabulary, screenshot and trailer references give recommender and classifier models cold-start material before any play data exists.

Localization and readiness review. Supported-language lists, age gates and per-board rating descriptors turn region-readiness screening into a filter instead of a survey.

Which personas get the most value?

  • Competitive Intelligence & Product Teams - tracking discount behavior, category volumes and platform reach as leading indicators of where the PC market is heading.
  • Market Researchers & Consultants - sizing segments by genre, price band, platform and review tier with one consistent 30-column schema.
  • Data Scientists & ML Engineers - a quarter-million text-plus-metadata records for classification, pricing models and recommendation cold-start.
  • Investors & Quants - storefront-wide pricing and supply series for due diligence on studios and gaming funds.
  • Developers & Data-Product Builders - stable app IDs and typed columns behind discovery tools, widgets and catalog integrations.
  • Journalists, Academics & Students - citable, reproducible evidence for reporting and coursework on digital game markets.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - Steam Store - Scrapeable Catalog (definitions verified against live storefront records)
fieldtypedefinitionexample
typeenumStore item class such as game, dlc or demo.game
namestringDisplay title of the application.Dota 2
steam_appidintegerCanonical Steam application ID, stable across re-pricing and re-releases.570
required_ageintegerMinimum age gate for the title.0
is_freebooleanWhether the base application is free to play.true
dlcstringApp IDs of downloadable content packages attached to the title.
detailed_descriptiontextFull store description, delivered as HTML.
short_descriptiontextShort marketing blurb used in listings.
about_the_gametextAbout-section text, the longer pitch variant.
supported_languagestextLanguages with interface and audio support, including headers distinguishing subtitles.
header_imagestringMain capsule image URL for the listing.
pc_requirementsstringMinimum and recommended Windows system requirements.
mac_requirementsstringmacOS system requirements.
linux_requirementsstringLinux system requirements.
developersstringDeveloper names on the record.["Valve"]
publishersstringPublisher names on the record.["Valve"]
package_groupsstringPurchase options, sub-packages and price tiers in cents.
platformsstringBooleans for windows, mac and linux availability.windows, mac, linux
categoriesstringFeature flags such as Single-player, Multi-player, Steam Achievements, Remote Play.
genresstringGenre objects with id and description.
screenshotsstringScreenshot image URLs with path thumbnails.
moviesstringTrailer and gameplay video metadata with webm and mp4 URLs.
recommendationsstringTotal count of positive recommendations.
achievementsstringTotal number of achievements.
release_datestringComing-soon flag plus the release date string.
support_infostringSupport email and URL published for the title.
metacriticstringMetacritic score and product URL where available.
ratingsstringContent rating descriptors per ratings board.
review_score_descstringReview summary label such as Very Positive.Very Positive
total_reviewsintegerTotal recent reviews counted in the review query summary.6119

What teams do with it

  • Price and promotion intelligence Original and discounted amounts in every record make discount depth, frequency and timing computable across the whole storefront.
  • Catalog and assortment research Type, genre, category and platform flags turn the storefront into a supply census, DLC clusters included.
  • Recommendation cold-start Descriptions, categories, screenshots and trailer references give models structure before any play data exists.
  • Regional pricing analysis Country-localized price points across supported currencies support purchasing-power comparisons without stitched mirrors.
  • Localization and compliance screening Language support, age gates and per-board rating descriptors make region-readiness a filtered query.

Questions buyers ask

How many applications does the Steam Store catalog dataset cover?

261,285 searchable applications as of August 2026, counted from the storefront's own result totals. That scale covers games, DLC, demos and application software in one schema, which is what turns the storefront from a shopping destination into a census of the PC market.

Does the dataset include anything besides games?

Yes. One record per application covers games, DLC packages, demos and software alike, with a type field separating them. Bundle structure survives too: purchase options arrive as nested packages with sub-tiers rather than a single flattened price.

How are discounts represented in the data?

As paired amounts, never a bare percentage. The Black Myth: Wukong row shows the pattern: a 30% cut takes the listing from 3,599.00 to 2,519.00 INR with both figures preserved, so discount depth is computable from the row instead of inferred from a label.

What do the review fields tell me?

A human-readable summary label such as Very Positive, backed by the arithmetic: Dota 2's recent window shows 5,430 positive against 689 negative out of 6,119 total reviews. Label and counts together mean sentiment tiers are reproducible, not taken on faith.

How current is the pricing information?

Records reflect the storefront's live prices and discounts at collection time, and delivery cadence is your choice - daily for market reviews, hourly through sale seasons when discount events move fast. Each snapshot lands in the same comparable schema.

How far back does release history go?

Release dates span 1990s titles still on sale through pre-orders dated years ahead, with a coming-soon flag separating shipped from announced. That makes the field usable for both back-catalog age analysis and forward-looking slate tracking.

What can I combine this dataset with?

Stable app IDs join cleanly to player-count, ownership-estimate and sales-chart sets in the same Interactive Home Entertainment catalog, letting you line storefront supply and pricing up against demand-side signals in one analysis.

Notes on this record

  • Commerce-grade depth Thirty verified columns per application, from full description HTML to support contacts and per-OS requirements.
  • Prices you can model Original and discounted amounts preserved side by side, with purchase packages nesting sub-tiers in cents.
  • Discount events stay reconstructable Paired amounts mean a -30% event is arithmetic, not a label - depth and timing survive every snapshot.
  • Reputation rides along Summary labels plus positive, negative and total counts land on the same record as the price.
  • Joins hold up Canonical app IDs connect rows to player-count, ownership-estimate and sales-chart sets in the same catalog.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing