US Census Retail Historical Data Portal

Datadory delivers us census retail historical data portal data: the Census Bureau's complete retail-release archive - advance monthly reports opening October 1953, quarterly e-commerce reports from late 1999, benchmark revisions since 1996 and SIC-era state workbooks - every clothing-store series normalized into one queryable table, delivered daily, weekly, or hourly.

What is the US Census Retail Historical Data Portal?

One front door, seven decades of American retail. The US Census Retail Historical Data Portal is the historical archive behind the Census Bureau's retail indicators program - four distinct layers spanning the advance reports, the e-commerce supplement, the benchmark revisions and the pre-NAICS era. Together they hold roughly 1,010 advance monthly retail trade reports opening in October 1953, 107 quarterly e-commerce reports starting Q4 1999, 109 annual-revision and benchmark files from 1996 and 22 SIC-based state and area workbooks covering 1986-1996.

For apparel the payoff is the stretch nothing current reaches: clothing-store history recorded under the old Standard Industrial Classification, where the lines run as SIC 56 men's and boys' clothing stores and SIC 57 clothing accessories instead of today's NAICS 448 branch. Spliced onto the modern series, that gives a clothing-demand model retail cycles reaching back to the 1950s - recessions, the mall boom, the e-commerce inflection - rather than the post-1992 slice everyone else trains on.

Datadory delivers the whole stack as one normalized table keyed by kind of business and period, so the SIC-to-NAICS seam becomes a join you control instead of a decade of file archaeology. The apparel retail data hub shows how this archive sits beside the industry's transactional and ML sources.

What do sample rows look like?

The archive publishes releases and workbooks, not a tidy extract - which is precisely the problem Datadory removes. Every delivery lands as one flat table, one row per kind of business per period, carrying the three documented columns below. Each value shown is a recorded example from the field dictionary, attached to its own column rather than stitched into a made-up observation:

# one row per kind_of_business x period, normalized
kind_of_business : text   | "Clothing stores"
period           : date   | "Oct. 1953"
est_sales_musd   : number | 6734

The second half of the proof is the archive map - what each layer covers, at what interval, and how much of it exists:

archive                       | series_interval | opens    | files
advance_monthly_retail_trade  | monthly         | Oct 1953 | ~1010
quarterly_ecommerce           | quarterly       | Q4 1999  | 107
annual_revision_benchmark     | annual          | 1996     | 109
sic_state_area_workbooks      | era 1986-1996   | 1986     | 22

Request a sample and you receive real rows from whichever layers your project touches - an advance report page, an e-commerce quarter, a SIC-era state sheet - rendered in that same normalized schema, with the full field dictionary riding along.

What fields does the dataset include?

Three columns carry every archived table. Kind of Business is the row label and the classification story in miniature: current-era files speak NAICS - the 448 clothing and clothing accessories branch - while SIC-era files speak the older coding, 56 men's and boys' clothing stores and 57 clothing accessories. Period is the column header, monthly on the advance side and quarterly on the e-commerce side. Estimated sales is the figure itself, in millions of dollars, not adjusted unless a given release labels it otherwise.

One honesty note: this archive predates modern data dictionaries, so our field definitions carry an inferred confidence flag - reconstructed from release layouts rather than verified cell-by-cell. The complete dictionary with worked examples, including the columns listed in the folded note below, ships with every sample.

What does coverage look like across geography, time and granularity?

Geography: United States national totals throughout. The SIC-era workbooks are the exception that proves the rule - separate state and area breakdowns for 1986-1996, the only place this archive descends below the national line.

Temporal: four openings instead of one. Advance monthly reports begin October 1953; the quarterly e-commerce series begins Q4 1999; annual revisions begin 1996; the SIC-based files span 1986-1996. Combined with the live NAICS panel in the companion MRTS product, the clothing record becomes effectively continuous from the mid-century mark forward.

Granularity: monthly estimates in the advance layer, quarterly in e-commerce, annual and benchmark in the revision layer - all by kind of business, first under SIC and then under NAICS. That classification handoff is a feature for researchers: it lets you test whether apparent structural breaks in clothing sales are real behavior or merely re-labeling.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Who uses this data, and for what?

  • Multi-cycle backtesting - validate demand and inventory models against retail cycles back to the 1950s instead of a single post-1992 regime.
  • Structural-break analysis - the SIC-to-NAICS handoff plus the annual revision files separate genuine shifts in clothing demand from mere re-classification.
  • E-commerce penetration history - the Q4 1999 origin of the online series is the baseline every omnichannel growth curve quietly leans on.
  • Forecast calibration - benchmark revisions show how first estimates matured into final ones, which is exactly the error structure an advance-print forecaster needs to model.
  • Long-run category mix - men's versus women's versus family clothing stores and shoe stores, tracked across decades of changing wardrobes and retail formats.

Which personas get the most value?

Quant investors and economists get the deep past their backtests assume but rarely have - federal retail estimates across seven decades. Data scientists and ML engineers get a stable, kind-of-business-keyed panel long enough to train through multiple downturns. Market researchers and consultants get citable federal figures for any apparel-market history slide. Journalists and academics get primary releases for the exact month they are writing about, not someone's summary of them. E-commerce strategists get the 1999 baseline that sizes how young their channel really is. Persona-specific framings live in apparel retail data for data scientists, developers and builders, journalists, academics and students and e-commerce operators.

What should I know before requesting a sample?

Three things worth knowing upfront. First, field definitions are inferred - the archive's layouts were reconstructed rather than verified cell-by-cell, which is why this dataset scores 7 against the 9s of its modern siblings; treat the sample as the place to confirm the schema. Second, the classification seam is real: SIC 56 and 57 do not map one-to-one onto NAICS 448's subcodes, so any long-run clothing splice needs an explicit crosswalk assumption - we surface ours rather than hiding it. Third, older releases contain withheld and unavailable cells, and the benchmark files mean early prints were later revised; pin the vintage you train on, because the revised record is the one your backtest should trust.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - three documented columns, one row per kind of business per period
fieldtypedefinitionexample
kind_of_businesstextRow label identifying the retail line; SIC-era files use SIC codes (56 men's and boys' clothing stores, 57 clothing accessories) rather than the NAICS 448 branch."Clothing stores"
perioddateColumn header of the release tables: monthly periods in the advance reports, quarterly periods in the e-commerce reports."Oct. 1953"
est_sales_musdnumberEstimated sales in millions of dollars, not adjusted unless the release labels the series otherwise.6734

Questions buyers ask

How far back does US clothing-store data go?

Total retail estimates open in October 1953 with the advance reports. Clothing-specific detail arrives in two coding systems: SIC-coded lines such as 56 men's and boys' clothing stores in the 1986-1996 workbooks, then the NAICS 448 branch from January 1992 in the companion MRTS feed. Spliced, that is a continuous seven-decade clothing record.

Which apparel lines appear under SIC instead of NAICS?

In the SIC-era files the apparel rows carry Standard Industrial Classification codes - 56 for men's and boys' clothing stores and 57 for clothing accessories - rather than the familiar NAICS 448 family. Boundaries differ slightly between the two systems, which is why a serious long-run splice needs a documented crosswalk rather than a naive join.

How does this archive relate to the live MRTS feed?

They are complementary halves of the same program. MRTS carries the current monthly NAICS panel from January 1992 onward; this archive holds everything that panel supersedes - prior advance releases, retired quarterly e-commerce reports, the benchmark revisions and the entire pre-NAICS era back to 1953. Pair them and no window is missing.

Does the archive include state-level detail?

National totals are the default across all four layers. The exception sits in the SIC-era workbooks, which add separate state and area breakdowns for 1986-1996. For recent months, modeled state detail comes from companion Census products rather than this archive, so sub-national apparel questions need both pieces.

Why does this dataset score 7 instead of 9?

Documentation, not accuracy. The archive predates machine-readable dictionaries, so field definitions carry an inferred confidence flag and some older cells are withheld or unavailable - hence the 7 against the 9s scored by its modern siblings. Statistical authority is federal-grade either way; the discount reflects schema friction, not data quality.

Datasets that pair with this one

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing