Datadory notebook
Garment segmentation dataset with bounding boxes: the 2026 shortlist
A garment segmentation dataset with bounding boxes starts with Fashionpedia on the Hugging Face Hub: 46,781 images carrying 342,182 bounding boxes and instance-segmentation masks across 46 categories and a 294-attribute ontology, annotations free under commercial delivery terms 4.0. DeepFashion contributes 800,000+ research-only images with boxes and landmarks; every retailer catalog in this industry ships none.
1,744 datasets. Pick your catch.
What does the Fashionpedia bounding box dataset actually contain?
What does the Fashionpedia bounding box dataset actually contain?
The Fashionpedia — Detection Datasets on Hugging Face record packages the ECCV 2020 benchmark as a fixed Parquet conversion of exactly 3,478,843,264 bytes. Its 46,781 everyday and celebrity-event photographs carry 342,182 instance-segmentation masks and bounding boxes, split 45,623 train and 1,158 validation. One row per image holds image_id, the decoded PIL image, width, height, and a nested objects struct with bbox_id, category, bbox coordinates and mask area. Coordinates arrive as four float64 values in Pascal VOC order - [xmin, ymin, xmax, ymax] - and observed images run from roughly 232x388 up to 1,024 pixels on the long side.
Annotation density varies the way street photography does: sample rows show image_id 23 as a 682x1024 portrait with 4 objects while image_id 26 packs 15 objects into a 1024x683 frame. Underneath the flat labels sits the expert-built ontology - 27 main apparel categories, 19 apparel parts and 294 fine-grained attributes reaching down to zipper, bead, sequin and tassel.
Two conversion gaps matter before training. The ClassLabel flattens those 27 categories and 19 parts into a single list of 46 names, so confirm which level an index refers to before computing per-class metrics. And none of the 294 attributes appear as columns - per-mask attribute labels require joining the original COCO-style JSON back on image and box identifiers.
How do DeepFashion's boxes compare to Fashionpedia's masks?
How do DeepFashion's boxes compare to Fashionpedia's masks?
DeepFashion Dataset (CUHK MMLab) (quality 9) wins on sheer volume: 800,000+ web-sourced shop and consumer photographs annotated with 50 clothing categories, 1,000 attributes, bounding boxes and landmarks, plus 300,000+ cross-pose pairs. Its granularity is one record per clothing image with item-level annotations - a single garment centered in frame - which suits classification, attribute prediction and retrieval better than scene parsing. Four benchmarks organize the release: Category and Attribute Prediction alone spans 289,222 images, In-shop Clothes Retrieval covers 52,712 images over 7,982 items, and Consumer-to-shop Retrieval and Landmark Detection complete the set, with extensions through 2022 (DeepFashion-MultiModal).
Fashionpedia inverts every one of those properties. Fewer images (46,781), but multiple garments and accessories per unconstrained frame, pixel-accurate outlines instead of rectangles, and a taxonomy that names trims rather than just garment types. Choose DeepFashion when the task is recognizing what a garment is; choose Fashionpedia when the task is finding where every garment sits in a cluttered photo.
The contractual difference is just as sharp. DeepFashion requires a signed Release Agreement, bars reproduction, sale and commercial exploitation of images or derived data, forbids redistribution beyond internal copies at a single site, reserves termination of access, and mandates citation of the CVPR 2016 paper. Fashionpedia needs no paperwork at all.
Is any garment detection data here cleared for commercial use?
Is any garment detection data here cleared for commercial use?
The pattern repeats industry-wide: 20 of the 26 primary datasets here are free to access, mirroring 1,437 of the 1,744 datasets Datadory catalogs overall, yet few pair free access with commercial clearance on the pixels themselves.
Can retail product catalogs substitute for detection annotations?
Can retail product catalogs substitute for detection annotations?
Where to go next
Where to go next
This page is the detection-and-segmentation chapter of Datadory's apparel coverage. Continue with:
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
Fashionpedia — Detection Datasets on Hugging Face
DeepFashion Dataset (CUHK MMLab)
Fashion-MNIST - Zalando Research Apparel Image Dataset
Fashion Product Images Dataset (44k) — Kaggle
Lyst Fashion Product Catalog Data
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
Does DeepFashion include segmentation masks?
No. DeepFashion annotates its 800,000+ images with 50 clothing categories, 1,000 attributes, bounding boxes and landmarks - not instance masks. When per-pixel garment outlines are the requirement, Fashionpedia's 342,182 masks are the only option among this industry's 26 primary datasets; DeepFashion remains the pick for attribute breadth and cross-pose retrieval.
How are Fashionpedia's 342,182 bounding boxes split?
They sit across 45,623 training images and 1,158 validation images, with the test split left empty - so evaluation runs on validation and its 8,781 boxes. Each Parquet row nests objects containing bbox_id, a category index from the 46-value ClassLabel, Pascal VOC-order coordinates and mask area.