Fashion Product Images Small — Hugging Face (ashraq)
Datadory delivers apparel, accessories & luxury goods data covering Fashion Product Images Small — Hugging Face (ashraq): 44,072 Myntra product rows with gender, masterCategory, subCategory, articleType, baseColour, season, usage, display name and embedded thumbnail in Parquet. Delivered daily, weekly, or hourly via API, files, or your warehouse — request a sample to see the schema against your use case.
API, files, or your warehouse. Daily, weekly, or hourly.
What is Fashion Product Images Small?
It is the small edition of the Myntra-derived fashion catalogue that circulates widely for apparel machine learning: 44,072 products, each described by ten text or numeric attributes plus a resized photograph. Download size is 271,496,441 bytes, expanding to 546,202,015 bytes once decoded, and the whole thing sits in two Parquet shards auto-converted by the Hugging Face Hub under its data/ directory. The card carries the size category 10K–100K with modality tags for image and text.
What makes this copy the convenient one is completeness. A sibling Hub upload (ceyda/fashion-products-small) ships only 42,700 rows and resizes images to a 512-pixel maximum dimension; this version keeps all 44,072 rows intact. Downstream artefacts — including the FashionGPT Space and item-search demos — build on it, and engagement sits at 44 likes and roughly 1,658 downloads.
What does the schema look like in practice?
The tabular half reproduces the Kaggle styles.csv schema unchanged: id joins back to the original release, gender through articleType form a strict three-tier category hierarchy, and baseColour, season and usage add merchandising context that most hand-assembled catalogues leave out. productDisplayName is free text but tightly patterned — colour words appear inside it, so it doubles as weak supervision for colour extraction.
Because the categories are closed sets rather than open strings, the corpus supports label-space experiments directly: train a classifier over the 141 article types, then evaluate whether predicted labels agree with subCategory constraints. Only 86 of the 1,744 datasets in Datadory's catalog ship Parquet, so this record loads into Polars or Dask without a conversion step most peers require.
How fresh is this corpus?
Static. The mirror was last modified on 2022-11-01, and the underlying assortment predates that — product years run roughly 2009–2017. Static snapshots are not rare in this catalog: 181 of 1,744 datasets never update, and research corpora are where they cluster. Plan around a fixed corpus rather than an incremental feed, and pair it with a living complement when recency matters — Lyst's catalogue spans more than 27,000 brands, Zalando's covers 7,000+ brands across 29 markets, and both sit in the same industry brief. For transaction-level demand signal instead of catalogue structure, H&M Group's release carries 31,788,324 purchases across 105,542 articles.
Who uses this data?
Six of Datadory's eight personas clear the relevance bar for this record, ranked here by fit:
- Data Scientists & ML Engineers (relevance 2/3) — stream all 44,072 rows for embedding and retrieval experiments without unpacking archives. Industry-specific patterns live on our data scientists use cases page.
- Developers & Data-Product Builders (relevance 2/3) — load thumbnails straight into an app pipeline with a documented feature schema, no archive handling required; see developers builders use cases.
- Market Researchers & Consultants (relevance 1/3) — inspect Indian online apparel assortment structure when the original Kaggle source is awkward to fetch.
- Competitive Intelligence & Product Teams (relevance 1/3) — review the attribute vocabulary a top marketplace assigns, to calibrate your own product-data checklist; more on competitive intel product teams use cases.
- E-commerce Operators (relevance 1/3) — reference realistic attribute completeness across colour, usage and season when defining intake standards for your own catalogue.
- Journalists, Academics & Students (relevance 1/3) — demonstrate dataset-provenance risk using a popular mirror that declares no licence anywhere on its card.
How does it compare to the other fashion corpora in this slice?
It wins on loading friction and loses on scale and rights clarity. Within its own family: the Kaggle original ships the same 280 MB-scale imagery behind a Kaggle account with an explicitly declared licence; the 44k Fashion Product Images Dataset adds roughly 24.8 GB of high-resolution photography; the ceyda copy trades 1,372 rows for smaller 512-pixel images. Against the wider machine-learning corpora, DeepFashion dwarfs it with 800,000+ annotated images across 50 categories and 1,000 attributes, Fashionpedia brings 46,781 images with 342,182 bounding boxes, and Fashion-MNIST remains the 70,000-image baseline benchmark. None of those load as a single flat table of 44,072 labelled SKUs with a closed attribute vocabulary — this record's defining advantage.
How does this record rate on quality?
Datadory scores it 7 out of 10 on its rubric of field documentation, access reliability and freshness — a band shared by 296 of the 1,744 cataloged datasets, just below the catalog-wide mean of 7.81. Field definitions are marked verified: the feature schema was cross-checked against repository metadata and live row output during research. Sample data exists and is quotable — id 15970, "Turtle Check Men Navy Blue Shirt", Men/Apparel/Topwear/Navy Blue, Fall 2011.
Three things hold the score at 7. The dataset card is a stub documenting nothing beyond the feature schema; no licence is declared anywhere on the card, repo YAML or API metadata, even though the source Kaggle release carries a licence label applied to marketplace-sourced imagery — treat rights verification as part of your procurement step; and streaming mode against the auto-converted shards went untested during research. Engagement metrics suggest the repo stays healthy regardless.
Fashion Product Images Small — Hugging Face (ashraq): specification
| Attribute | Value |
|---|---|
| Rows | 44,072 (single split) |
| Fields | 11 — 10 tabular + 1 embedded image feature |
| Formats | Parquet (two shards, auto-converted) |
| Download size | 271,496,441 bytes; 546,202,015 bytes decoded |
| Coverage | India — Myntra e-commerce catalogue; product years roughly 2009–2017 |
| Granularity | Product level, one row per SKU |
| Source name | Hugging Face Datasets (community user ashraq) |
| Datadory quality score | 7 / 10 (catalog mean 7.81) |
Closed-set vocabularies in the attribute columns
| Field | Distinct values | Example value |
|---|---|---|
| gender | 5 | Men |
| masterCategory | 7 | Apparel |
| subCategory | 45 | Topwear |
| articleType | 141 | Shirts |
| baseColour | 46 | Navy Blue |
| season | 4 | Fall |
| usage | 8 | Casual |
Questions buyers ask
How many rows and fields does this dataset contain?
44,072 rows in a single split, each carrying eleven features: id, gender, masterCategory, subCategory, articleType, baseColour, season, year, usage, productDisplayName and an embedded image. The categorical columns are closed sets — 5 gender values, 7 top-level categories, 45 subcategories, 141 article types, 46 colours, 4 seasons and 8 usage contexts — so the label space is fully enumerable before you write any code.
Is the imagery full resolution?
No — this is deliberately the small edition. Photographs are resized thumbnails packaged alongside the attributes so the whole corpus fits in about 271 MB downloaded and 546 MB decoded, against roughly 24.8 GB for the high-resolution 44k sibling. For garment-classification prototyping, retrieval indexing and schema work the thumbnails are sufficient; production vision systems usually graduate to the full-resolution set once the approach is validated.
Which markets does the catalogue cover?
India. Every row comes from Myntra's e-commerce catalogue, inherited from the source release, so merchandising conventions reflect that marketplace: season tags tuned to Indian retail cycles, usage contexts such as Casual and Ethnic Wear, and colour naming in marketplace style. Treat it as one market's assortment vocabulary rather than a global standard, and expect transfer-learning gains to shrink for markets with very different category trees.
Does the dataset include prices or stock levels?
No. The schema stops at descriptive attributes — category, colour, season, usage and display name. There are no price, discount, inventory or sales columns anywhere in the release, because the source was assembled for computer-vision research rather than commerce analytics. If you need commercial signal, pair it with Lyst's catalogue for luxury price and discount dynamics or H&M's transaction release for purchase behaviour.
Can I use it to train a visual search model?
That is close to its intended use. Each row binds a thumbnail to a complete attribute vector, giving supervised pairs for embedding training and a ready-made query space — filter by gender, subCategory and baseColour before the model ever runs. Downstream artefacts built on this exact copy include the FashionGPT Space and item-search demos, which shows the pipeline holds at production-ish scale despite the modest image size.
How often would Datadory refresh my copy?
Your call: daily, weekly, or hourly, delivered by API, files, or straight into your warehouse. Note that the upstream snapshot itself is static — last modified November 2022 with product years spanning roughly 2009–2017 — so a faster cadence buys you fresher delivery of a frozen corpus, not new rows. If your roadmap needs post-2017 assortment, say so in the sample request and we will scope a living complement from the same industry.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.