Apparel, Accessories & Luxury Goods

Zalando Apparel & Accessories Catalog Data

Datadory delivers Zalando apparel accessories catalog scrapeable data covering millions of articles from more than 7,000 brands across 29 European markets — brand, EAN/GTIN, category tree, price and discount state, colour, size, material composition, ratings and sustainability badges at variant level. Delivered daily, weekly, or hourly as an API, files, or straight into your warehouse.

API, files, or your warehouse. Daily, weekly, or hourly.

Where it covers
29 European markets
How far back
Continuously current catalogue
How fine
Article / variant level (brand × style × colour × size)

What does a Zalando catalog record look like?

Each record is one purchasable article — brand × style × colour × size — with its live storefront attributes attached. Three rows from the catalog, exactly as they land in your delivery:

Adidas | 4065428123456 | Shoes > Sneakers > Low-top sneakers | €89.95 | black | sizes: EU 42, EU 43, EU 44 | Upper: synthetic; lining: textile; sole: rubber BOSS | 4062547890123 | Men > Clothing > Shirts > Business shirts | €119.00 | light blue | sizes: S, M, L, XL | 100% cotton Zign | 4059372123456 | Women > Bags > Shoulder bags | €49.95 | cognac | sizes: One size | Outer: imitation leather; lining: textile

Every field above resolves to the dictionary below. No nulls dressed up as zeros, no brand names split across three spellings — one row, one article, one EAN.

Which fields does the dataset include?

The core record carries brand, EAN/GTIN, category tree, price with discount state, colour, size run, declared material composition, aggregate customer rating, and any sustainability badge shown on the product page. Nine structured fields per article is enough to answer most assortment and pricing questions without joining a second source.

How wide is the coverage?

Geography: 29 European markets, each served by its own country storefront — Germany, France, Italy, Spain, the UK, the Nordics, Benelux, CEE and more.

Temporal: continuously current catalogue. Pricing history accrues from repeated capture, so a weekly cadence builds a usable discount-depth curve within a quarter.

Granularity: article/variant level — brand × style × colour × size — not brand rollups, not category averages.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Pick the channel that matches how your team already works: REST endpoints for live lookups, flat files for batch analysis, or a direct pipe into Snowflake, BigQuery or Redshift. Cadence is your call — set it once, change it whenever your models change.

Every delivery ships with the full field dictionary, sample rows for validation, and a schema that does not drift underneath you between refreshes.

Who uses this dataset, and for what?

Pricing and promotion analysts at apparel brands benchmark their own price points and discount depth against the largest assortment in European fashion e-commerce, market by market. A 30% markdown in Berlin reads differently when you can see whether the same article is at 15% in Amsterdam.

Assortment and buying teams map category-tree breadth against their own range plans: which brands are adding styles in women's outerwear this season, which colour families dominate each price band, where a 7,000-brand platform chooses to go deep versus wide.

Investment and strategy researchers treat the catalogue as a demand-side signal — active style counts, discount pressure and rating distributions across 29 markets, tracked over time rather than guessed once.

Data science teams use the structured attributes as clean training or enrichment input: category classification from category trees, price-elasticity features from price plus rating, sustainability positioning from badge coverage.

Why Zalando's catalog specifically?

Scale first: 62 million active customers, more than 7,000 brands, roughly EUR 17.6 billion in gross merchandise volume across 29 markets. Breadth like that makes the platform a de facto reference price and assortment index for European fashion.

Structure second: product pages follow one consistent template, so brand, EAN, price state, colour, size and material composition arrive as comparable columns instead of free text pulled out of twenty different layouts. Comparability is the hard part of retail benchmarking; here it comes built in.

And a note on what this dataset is not: it is not a historical archive. The catalogue is live commerce, so history exists only where someone captured it — which is precisely why teams start a capture cadence before they need the curve.

Field dictionary — Zalando apparel & accessories catalog data (article-level)

FieldTypeDefinitionExample
brandstringBrand name of the listed article.Adidas
eanstringEuropean Article Number identifying the specific product variant.4065428123456
pricenumberCurrent selling price in the storefront currency.89.95
colorstringColour family or colour name of the article.black
Additional fields on requestCategory tree, size run, material composition, customer rating and sustainability badges are also captured per article; definitions and examples ship with your sample.

Questions buyers ask

How many brands and articles does the dataset cover?

More than 7,000 brands sell through the platform, and the catalogue runs to millions of articles across 29 European markets. Exact article counts per market shift daily with the live assortment, so coverage is quoted as breadth — brands, markets, categories — rather than a frozen row count.

Does the dataset include pricing history?

Each delivery reflects the current catalogue state; history builds from repeated captures on your chosen cadence. Weekly capture yields a workable discount-depth series inside a quarter, while daily cadence catches flash-promotion windows that weekly snapshots miss.

Can I get sustainability and material attributes per article?

Yes. Material composition is declared on the product page and carried per article, alongside any sustainability badge shown to shoppers. Both fields are part of the standard dictionary, so green-positioning analyses need no separate enrichment pass.

Is country-level comparison possible within one delivery?

Yes. All 29 markets arrive under one schema, keyed consistently by EAN, so the same article can be compared across storefronts without reconciliation work. That single-schema design is what makes cross-market price and assortment gap analysis a query rather than a project.

What granularity does each record represent?

One row per purchasable variant: brand × style × colour × size. Rollups to brand, category or market level are straightforward aggregations from that base, but the atomic unit stays the article, which keeps every downstream cut reproducible.

Datasets that pair with this one

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing