Homebuilding · Hugging Face
Hugging Face Datasets - Furniture Search - 77 Community Corpora
Datadory delivers hugging face datasets furniture search 77 datasets data: every community corpus the Hub's 'furniture' query surfaced at research time - image-classification sets, robotics manipulation episodes, synthetic renders and interior-design imagery, some above one million rows - normalized into catalog rows with verified field definitions, delivered daily, weekly, or hourly.
API, files, or your warehouse. Daily, weekly, or hourly.
What is the Hugging Face Datasets - Furniture Search?
A discovery layer over a community machine-learning hub, not a subject dataset - and the distinction decides how you use it. Where a conventional dataset hands you observations in one shared table, this record hands you the 77 repositories the Hub's 'furniture' query returned across three result pages at the August 21, 2026 research pass. Results sort by trending and filter along four axes: nine modalities (3D, audio, document, geospatial, image, tabular, text, time-series, video), row-count bands from under 1K to over 1T, file format, and type flags separating benchmarks from traces.
The named matches show how wide the quality range runs. minhoheo/furniture-test recorded four trailing downloads when observed; Francesco/furniture-ngpea carried two likes and a March 30, 2023 revision. At the other end, IPEC-COMMUNITY/furniture_bench_dataset_lerobot - a LeRobot-format corpus of furniture-assembly robot episodes in the 1M<n<10M band, tagged tabular, timeseries and video at once - showed 9,782 trailing downloads and a February 19, 2025 revision stamp. Falah/chairs_furniture, Arkan0ID/furniture-dataset and disham993/Synthetic_Furniture_Dataset fill the middle of the range.
Within Datadory's catalog of 1,744 datasets across 159 viable industries this record scores 6/10 for quality with field definitions verified against live responses during research. Its role on the homebuilding data hub is singular: the only slice reaching model-training material, where every Census and trade-group neighbor measures the built environment instead.
What do sample rows look like?
Three ranked results from the August 2026 research pass, flat exactly as the catalog fields arrive:
# one catalog row per matched repository - the triage card in flat form
# Row 1 - a minimal smoke-test upload
id : minhoheo/furniture-test
author : minhoheo
lastModified : 2023-03-16T15:48:21Z
downloads : 4
likes : 0
# Row 2 - an early community experiment
id : Francesco/furniture-ngpea
author : Francesco
lastModified : 2023-03-30T09:12:40Z
likes : 2
# Row 3 - the heavyweight of the match set
id : IPEC-COMMUNITY/furniture_bench_dataset_lerobot
downloads : 9782
tags : task_categories:robotics | size_categories:1M<n<10M
modality:tabular | modality:timeseries | modality:video
# the funnel in two numbers
query_matches : 77
result_pages : 3Read the spread, not the individual values. The first two rows are the floor of the range: single-digit engagement, spring-2023 revisions, uploads that exist to prove a pipeline works. The third row is the ceiling: five-figure downloads, and a tag set doing the work of a second schema - task category, size band and three modality declarations compressed into one text field. Triage happens on those tags before any underlying file is opened, which is why the field dictionary treats tags as substance rather than decoration.
Rows reflect the August 2026 review pass; a fresh pull arrives pinned and documented so your counts reconcile against a dated snapshot.
What fields does the dataset include?
Eight fields complete the core schema, verified against live responses during the research pass - nothing inferred from names alone. They fall into four families. Identity: id, the author/name handle every join starts from, and author, which separates research labs from individual accounts. Attention: downloads and likes, the usage signals that sort a seventy-seven-row sweep into worth-a-pipeline and worth-skipping. Freshness: lastModified, the revision stamp whose sampled values run March 2023 through February 2025. Classification and substance: tags, the collapsed vocabulary of task, modality, size band and reuse declarations; gated / private, the restriction flag; and description, the card text describing structure and provenance.
Beyond the eight-column core, the rest is shaping - and shaping folds in on request rather than padding every extract:
What does coverage look like across geography, time and granularity?
Geography - global by construction. Contributors upload from anywhere and no geographic restriction applies, but the records carry no country column either, so any regional attribution needs an outside join. Treat geography as a dimension you bring to the corpus, not one you read off it.
Temporal - layered rather than single-dated. Uploads across the match set begin in 2023, and sampled revision stamps run from March 16, 2023 to February 19, 2025, so vintage labeling matters row by row. The seventy-seven figure itself is a point-in-time reading taken August 21, 2026: repositories appear and vanish as contributors push and withdraw material, which is why delivered sets are pinned and documented at sampling time.
Granularity - one catalog entry per dataset, each opening into per-file detail beneath it. There is no shared observation table underneath, and row counts span four orders of magnitude, from under-1K test sets to the 1M<n<10M band the assembly corpus occupies. Pairing the corpus view with a commercial counterpart - the per-SKU rows of the IKEA Global Product Catalog - is how training material and retail ground truth end up on the same axis.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Who uses this data, and for what?
- Furniture vision models - image-classification and detection sets supply labeled chairs, sofas and tables for training and evaluation without commissioning a photo shoot; the workflow continues on our ml model training page.
- Robotics manipulation research - the LeRobot-format assembly corpus contributes million-band episode logs carrying timeseries and video modalities, a ready substrate for policy-learning benchmarks.
- Interior-design and listing enrichment - interior imagery and synthetic renders back staging experiments, category-page art tests and listing-photo augmentation for retail teams.
- Synthetic-to-real pretraining - render-heavy corpora pretrain models cheaply before a fine-tune on real photographs.
- Label-taxonomy benchmarking - community tag vocabularies serve as an outside yardstick for retail category trees, surfacing where crowd taxonomies merge or split classes differently.
- Teaching and coursework - small typed corpora put genuine computer-vision material in front of students at notebook scale.
Which personas get the most value?
Data Scientists & ML Engineers lead fit: training-ready corpora with typed catalog rows above them and a five-figure-download flagship for benchmarking - see data scientists use cases. Developers & Data-Product Builders ship demo and evaluation features on corpora whose schemas document themselves - see developers builders use cases. E-commerce Operators borrow imagery for listing and category experiments sourced outside their own catalogs - see e-commerce operators use cases. Journalists, Academics & Students get named uploaders, revision stamps and download counts, which is what makes a citation checkable. Market Researchers & Consultants read the tag vocabulary as consumer-taxonomy evidence; Competitive Intel & Product Teams compare community corpora against proprietary content to find gaps. Start from the best homebuilding datasets ranking to see where this record lands among its neighbors.
What should I know before requesting a sample?
Four honest caveats.
First, curation is the product. A name-only query drags in near-zero-download uploads and SEO-shaped aggregator entries alongside genuinely useful corpora; Datadory filters at packaging time and documents what was excluded rather than passing the raw count through as substance.
Second, reuse declarations travel per repository and are absent on a meaningful share of the match set. Every delivered row carries the declaration status of its corpus explicitly, so rights questions surface before integration rather than after production has memorized the wrong assumption.
Third, the count drifts. Seventy-seven is the August 21, 2026 reading; the assignment brief itself cited seventy-six. Any pipeline consuming this record should treat the number as dated evidence, not a stable denominator.
Fourth, geography is absent by design. No country column exists on these rows, so market-mapping work needs an enrichment join - tell us the dimension you need and it rides along with the sample.
Notes and related datasets
Scope note - 'seventy-seven datasets' counts the Hub's search results at research time, not seventy-seven vetted industry sources. After removing low-engagement and promotional uploads, the usable subset is materially smaller, and the curation reflects that rather than the raw figure.
Pairing note - this record answers 'what exists for teaching a machine', while the Data.gov Catalog - Furniture Search answers 'what government knows about the trade'. Run together, safety science and training corpora land in one delivery.
Fitness note - nothing here supports housing-market measurement. For permits, sales and builder sentiment, the statistical spine on the rail below owns those questions entirely; this slice trains models, not forecasts.
Where to go next - the rail pairs the community-corpus view with its government-side, retail-side and statistical neighbors.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| Field | Type | Definition | Example |
|---|---|---|---|
id | string | Dataset identifier in author/name form on the Hub - the stable key every downstream join and per-file reference resolves against. | IPEC-COMMUNITY/furniture_bench_dataset_lerobot |
author | string | Account or organization that uploaded the dataset; doubles as an provenance marker separating individual tinkerers from research labs. | IPEC-COMMUNITY |
lastModified | datetime | Timestamp of the most recent revision to the repository - the freshness signal that lets stale uploads identify themselves. | 2025-02-19T01:52:26.000Z |
downloads | integer | Trailing download count shown on the result card; the closest thing this collection has to a usage-weighted relevance score. | 9782 |
likes | integer | Community like count recorded on the card; a weak but real signal alongside downloads. | 43 |
tags | text | Classification tags covering task category, modality, row-count size band, region and the reuse declaration each uploader chose. | size_categories:1M<n<10M |
gated / private | boolean | Whether the dataset requires accepting the uploader's conditions before use, or is restricted to its owner outright. | false |
description | text | Dataset-card text describing structure, collection method and provenance. | This dataset was created using LeRobot... |
Additional fields | - | Folded under "additional fields on request": per-file inventories, tag vocabularies unfolded one dimension per column, full dataset-card text, curated cuts and cross-catalog joins. | on request |
What teams do with it
- Furniture vision models Image-classification and detection sets supply chair, sofa and table labels for training and evaluating recognition pipelines without assembling a photo shoot.
- Robotics manipulation research The LeRobot-format assembly corpus contributes million-band episode logs with timeseries and video modalities - a ready substrate for policy-learning benchmarks.
- Interior-design and listing enrichment Interior imagery and synthetic renders support staging experiments, category-page art tests and listing-photo augmentation studies for retail teams.
- Synthetic-to-real pretraining Render-heavy corpora pretrain models cheaply before a fine-tune on real photographs, cutting collection cost on the front of the pipeline.
- Label-taxonomy benchmarking Community tag vocabularies give analysts an outside yardstick for retail category trees - where crowd taxonomies split or merge classes differently, merchandising assumptions surface.
- Teaching and coursework Small typed corpora put real computer-vision material in front of students at notebook scale, with the same columns they will meet in production.
Questions buyers ask
What does hugging face datasets furniture search 77 datasets data contain?
Seventy-seven community-uploaded machine-learning repositories that answered the Hub's 'furniture' query across three result pages at the August 2026 research pass: image-classification sets, robotics manipulation corpora, synthetic renders and interior-design imagery, ranging from single-digit-download test uploads to a furniture-assembly corpus above one million rows.
Why might other sources cite 76 datasets instead?
Because the number is a live reading, not a constant. The original assignment cited 76, while both the web interface and the programmatic listing returned 77 when checked on August 21, 2026 - repositories appear and vanish as contributors push and remove material, so any count carries its observation date.
What is inside the furniture_bench robotics corpus?
IPEC-COMMUNITY's LeRobot-format dataset of furniture-assembly robot episodes sits in the 1M<n<10M row band, tagged with tabular, timeseries and video modalities simultaneously, and carried 9,782 trailing downloads with a February 19, 2025 revision stamp at research time - by far the most-used entry in the match set.
Are all seventy-seven matches image collections?
No. The query pulls robotics episode logs, tabular and time-series tables, text corpora and synthetic renders alongside the photo sets, which is why the Hub exposes a nine-way modality filter - 3D, audio, document, geospatial, image, tabular, text, time-series and video - over these very results.
How current are the rows in this collection?
Each row carries its own revision stamp. Among sampled entries those run from March 16, 2023 to February 19, 2025, and uploads across the wider match set begin in 2023. The catalog rows themselves were re-observed during the August 2026 research pass, so delivered counts reconcile against a dated snapshot.
Can a sample be scoped to one modality or size band?
Yes - that is what scoping is for. Name the modality, the size band or specific repositories such as the LeRobot assembly corpus, and the sample returns catalog rows and file inventories for exactly that cut, shaped to your production columns rather than the full seventy-seven-row sweep.
Notes on this record
- Scored 6/10 Datadory scores this record six of ten on its rubric against a cross-catalog mean of 7.81 across 1,744 datasets - strong field verification, docked for uneven curation across the match set.
Datasets that pair with this one
- Data.gov Catalog - Furniture Search The government-side complement - 30 federal, state and local records matching the same query, from NIST fire tests to procurement lines.
- IKEA Global Product Catalog The commercial counterpart - per-SKU price, dimension and availability rows across storefronts, where these corpora supply the training material.
- US Census Building Permits Survey (BPS) The industry spine - units authorized by permit from nation down to place, measuring the market these models depict.
- Best homebuilding datasets Where this record ranks - eighth of ten - against the vertical's statistical spine and retail records.
- homebuilding data hub All primary datasets in this industry, ranked and cross-linked.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.