Cargo Ground Transportation · U.S. Census Bureau / Bureau of Transportation Statistics
Commodity Flow Survey (CFS) 2022
Datadory delivers cargo ground transportation data covering the Commodity Flow Survey (CFS) 2022: the joint Census Bureau and BTS measure of US domestic freight, 37,576,546 shipment-level records with origin, destination, SCTG commodity, mode, value in dollars, weight in pounds and great-circle miles - delivered daily, weekly, or hourly.
API, files, or your warehouse. Daily, weekly, or hourly.
- Where it covers
- United States - FIPS states, metro areas and the survey's own CFS areas, as origin-destination pairs
- How far back
- 2022 survey year, released across 2025-2026 (final estimates May 2025, summary tables June 2025, publication January 2026, PUMS January 2026); conducted every five years with 2017 and 2012 cycles also available
- How fine
- Shipment-level microdata - one row per shipment - plus roughly 30 pre-tabulated summary tables cut by commodity, mode, distance band and geography
What is the Commodity Flow Survey (CFS) 2022?
One survey stands underneath nearly everything America knows about its own freight: the Commodity Flow Survey, run jointly by the U.S. Census Bureau and the Bureau of Transportation Statistics, measuring domestic shipments by US establishments every five years. It is the primary microdata source for ground cargo flows and the foundation of the Freight Analysis Framework.
The 2022 cycle arrived in stages. Final estimates landed 20 May 2025, the summary tables followed on 26 June 2025, the survey publication came 20 January 2026, and the shipment-level Public Use Microdata Sample closed it out on 28 January 2026. That last piece is the prize: identifiers running 00000001 through 37576546 - 37,576,546 possible shipments, each carrying where a load started, where it ended, what it was, how it traveled, what it was worth and how far it went. Datadory scores it 10 out of 10 on its quality rubric - a mark held by just 145 of the 1,744 cataloged datasets against a 7.81 average. Get a sample of this dataset and see the rows before anything else.
What does a sample row look like?
Flat, wide, keyed - one row per sampled shipment. Straight off the verified 2022 dictionary:
SHIPMT_ID : 00000001
ORIG_STATE : 06 DEST_STATE : 48 # California -> Texas
SCTG : 35 MODE : 02 # truck move
SHIPMT_VALUE : $125,000
SHIPMT_WGHT : 42,000 lbs SHIPMT_DIST_GC : 1,745 miRead it as one truckload priced three ways: dollars, pounds and mile-weighted effort. The second row shape shows the classifier half of the record:
SHIPMT_ID : 01234567
ORIG_CFS_AREA : 31080 DEST_MA : 19100 # metro-to-metro lane
SECTOR : 31-33 TEMP_CNTRL_YN : N HAZMAT : 0
EXPORT_YN : N WGT_FACTOR : 12.4That last field is the quiet one doing the heavy lifting. WGT_FACTOR is the tabulation weight that inflates the sample into population estimates - skip it and your national totals describe only the surveyed fraction of American freight. With it, summed weighted SHIPMT_VALUE reproduces the headline figures the summary tables publish, which is exactly the cross-check that keeps a model honest.
What fields does the dataset include?
Eighteen documented variables, every one defined and worked into an example during verification. The table below carries the ten that do most of the analytical work: three nested origin-destination pairs (state, metro area and the survey's own CFS area), the shipper's industry SECTOR, the SCTG commodity code, the transportation MODE - which distinguishes for-hire truck from private truck, a distinction most freight datasets flatten away - plus dollar value, pound weight and great-circle miles.
The remaining eight ride under additional fields on request: the temperature-control flag, the export pair and country, the hazmat class and the tabulation weight. Field definitions are marked verified, meeting the standard held by 1,495 of the 1,744 datasets (85.7%) in the Datadory catalog.
What does coverage look like across geography, time and granularity?
Geography - three nested cuts, twice over. Every record holds origin and destination as a FIPS state pair, a metropolitan-area pair and a CFS-area pair, so a corridor reads at whatever resolution the question needs. Suppression is coded rather than dropped - 00 for a withheld origin state, 00000 for a withheld CFS area - which is the correct way to handle confidentiality but the wrong thing to ignore in a row count.
Temporal - the 2022 survey year, delivered across a staged release train: final estimates May 2025, summary tables June 2025, publication and PUMS microdata in January 2026. The survey runs every five years, and the 2017 and 2012 cycles are available alongside, which makes two-cycle comparisons straightforward and continuous annual series impossible - pair it with a monthly indicator when between-cycle freshness matters.
Granularity - shipment-level microdata, one row per shipment, plus roughly 30 pre-tabulated summary tables cut by commodity, mode, distance band and geography for when the question stops needing row-level detail. The microdata file runs large - hundreds of megabytes compressed - while a single summary table fits in a spreadsheet; Datadory ships either shape, cut to your commodities and lanes.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
The cadence follows your pipeline, not the other way around. Rows arrive flattened and typed, keyed on the shipment identifier, with the mode, SCTG and sector code sheets attached so nothing lands as an unexplained integer - joining a delivery against carrier telematics, facility locations or a rate table is a join statement rather than a documentation archaeology project. A sample cut to your commodities, lanes and geographies comes first either way.
Who uses this data, and for what?
- Market sizing - weight SHIPMT_VALUE to population by commodity, mode or distance band and the US domestic freight market gets a denominator with a federal pedigree. See market researchers x cargo ground transportation.
- Network planning - rank origin-destination pairs by value, weight or ton-miles before siting distribution capacity; metro and CFS-area endpoints make the ranking lane-specific rather than state-vague.
- Demand modeling - millions of shipment records with value, weight, distance and commodity feed lane-choice and freight-demand models directly; part of the shelf serving supply-chain mapping.
- Investment backtests - the for-hire versus private-truck split in MODE lets quants test tonnage proxies against carrier fundamentals across cycles; see investors & quants x cargo ground transportation.
- Cold-chain and hazmat studies - the temperature-control and hazmat flags isolate refrigerated and dangerous-goods flows by commodity and corridor without a second source.
- Reporting and coursework - the official citation-grade measure of what moves across America, suppression rules included; see journalists & academics x cargo ground transportation.
Which personas get the most value?
Market Researchers & Consultants get the federal benchmark under every freight TAM model (market researchers view). Data Scientists & ML Engineers get a clean keyed microdata set that trains demand models without scraping work (data scientists view). Investors & Quant Researchers get cycle-over-cycle trucking-demand proxies they can test against carrier revenues (investors view). Journalists, Academics & Students get quotable national freight figures with the methodology already written (journalists view). Developers & Data-Product Builders get a stable schema to seed freight-lane lookup features from (developers view).
What are the limitations?
Stated plainly, because they shape the analysis:
- Five-year means five-year. This is a benchmark, not a monthly indicator; intra-cycle movement needs a companion series, and treating a quinquennial snapshot as a leading signal misuses it.
- Suppression is coded, not absent. Values of 00 and 00000 mark withheld geography; summing past them silently biases fine-grained regional work toward the areas the survey could publish.
- Weights are not optional. WGT_FACTOR turns a sample into the economy; skipping it produces numbers that look precise and describe only the surveyed fraction.
- Distance is great-circle. SHIPMT_DIST_GC measures straight lines, not road miles, so route-length calculations need a routing layer on top.
- Code labels live in the companion sheets. Mode and sector arrive as codes; read them against the full dictionary - shipped with every delivery - before assuming any particular value.
Why request this through Datadory
Because the raw release measures in the hundreds of megabytes, splits its dictionary into a separate workbook, and expects you to reconcile staged releases that arrived eight months apart. Datadory normalizes the microdata and the summary tables into the single field dictionary above, resolves code sheets onto every row, flags suppression codes explicitly instead of leaving them to be discovered downstream, and schedules recurring deliveries keyed on the shipment identifier. Browse the shelf on the cargo & ground transportation data hub or the best cargo ground transportation datasets, then get a sample cut to your lanes.
Which datasets sit next to this one?
Neighbors that bracket the questions this survey leaves open. Freight Analysis Framework Version 5 (FAF5) republishes these flows as an origin-destination matrix across 132 domestic regions with forecasts to 2050 - the head-to-head runs in vs Freight Analysis Framework Version 5 (FAF5). FHWA's Freight Performance Measurement Program adds highway travel-time reliability between the endpoints. ATA Economics' truck tonnage measures track the month-to-month pulse between survey years. Eurostat and UK road-freight statistics extend the same questions across the Atlantic, and Data.gov's catalog reaches everything else published on US freight. Together: the microdata, the modeled matrix, the highway reality, the monthly pulse and the international comparison.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
SHIPMT_ID | string | Shipment identifier, running 00000001 through 37576546 - one per sampled shipment. | 00000001 |
ORIG_STATE / DEST_STATE | string | FIPS state codes for the shipment's origin and destination; origin 00 means completely suppressed. | 06 -> 48 |
ORIG_MA / DEST_MA | string | Metropolitan area codes for the shipment's origin and destination. | 31080 -> 19100 |
ORIG_CFS_AREA / DEST_CFS_AREA | string | CFS-area codes - the survey's own finer origin-destination geography; 00000 marks suppression. | 00000 |
SECTOR | string | Industry sector of the shipper establishment, on NAICS sector lines. | 31-33 |
SCTG | string | Two-digit Standard Classification of Transported Goods code for the shipment's primary commodity. | 35 |
MODE | string | Mode of transportation for the shipment, distinguishing for-hire truck from private truck. | 02 |
SHIPMT_VALUE | number | Value of the shipment in dollars - the column freight-market sizings hang their totals on. | 125000 |
SHIPMT_WGHT | number | Weight of the shipment in pounds; paired with distance it yields ton-miles per record. | 42000 |
SHIPMT_DIST_GC | number | Great-circle distance between origin and destination in miles. | 1745 |
What teams do with it
- US freight-market sizing Sum SHIPMT_VALUE weighted to population by SCTG commodity, mode or distance band and you hold a defensible total addressable market for ground freight - built on the federal benchmark rather than a trade-press estimate.
- Lane-level network planning Origin-destination pairs down to metro and CFS-area level turn 'where does our freight actually flow' into a query: rank corridors by value, weight or ton-miles before siting a distribution center.
- Trucking-demand backtests Mode codes separating for-hire from private truck let investors test tonnage proxies against carrier revenue histories across the 2017 and 2022 cycles.
- Freight-demand ML features 37.6M shipment records with value, weight, distance and commodity make the training set for lane-choice and demand models - the same substrate the Freight Analysis Framework itself is built on.
- Modal-share and cold-chain studies The temperature-control flag isolates refrigerated flows by commodity and corridor; the hazmat flag does the same for dangerous goods.
- Citation-grade freight context Every figure traces to the joint Census Bureau/BTS publication, which is why it holds up in a policy brief, a prospectus or a news chart without a sourcing argument.
Questions buyers ask
What exactly is the Commodity Flow Survey?
The joint U.S. Census Bureau and Bureau of Transportation Statistics measure of domestic freight shipments by US establishments, conducted every five years. It is the primary microdata source on what America ships, in what quantity, by which mode and at what value - and the foundation underneath the Freight Analysis Framework.
How many shipment records does the 2022 microdata hold?
Shipment identifiers run from 00000001 to 37576546, so 37,576,546 possible shipments in the 2022 Public Use Microdata Sample published in January 2026. Each record carries origin, destination, commodity, mode, value, weight and distance, with a tabulation weight for producing national estimates.
What time period does the 2022 edition cover?
The 2022 survey year, released in stages: final estimates in May 2025, summary tables in June 2025, and both the survey publication and the shipment-level microdata in January 2026. Because the survey runs every five years, the 2017 and 2012 editions sit alongside for two-cycle comparisons.
Can the data be cut to specific commodities, modes or lanes?
Yes - that is the native shape of the microdata. Every record already carries a two-digit SCTG commodity, a mode code separating for-hire from private truck, and nested state, metro and CFS-area origin-destination pairs, so a pull for one commodity on one corridor is a filter rather than a rebuild.
How should the tabulation weight be used?
Multiply before summing. WGT_FACTOR inflates each sampled shipment into its population share; raw sums describe only the sample. Weighted sums of SHIPMT_VALUE reproduce the national figures the summary tables publish, which makes them the quickest consistency check on any derived analysis.
Does it cover shipments moving only inside the United States?
Domestic movements form the core of the record. Export movements are flagged rather than excluded - EXPORT_YN marks them and EXPORT_CNTRY names the destination country - so international legs can be isolated or filtered out depending on the question.
How does CFS 2022 differ from FAF5?
Granularity versus reach. The CFS is the survey - shipment-level records for the 2022 year - and the Freight Analysis Framework is partly built from it, republishing flows as an aggregate zone-pair matrix across 132 regions with forecasts out to 2050. Take CFS for row-level detail, FAF5 for complete matrices and forward projections.
Notes on this record
- The survey behind the framework The Commodity Flow Survey is the microdata foundation of the Freight Analysis Framework - model the flows yourself and you are re-deriving, with full visibility, what FAF5 publishes as a finished matrix.
- 37,576,546 shipments, one key Identifiers run 00000001 through 37576546 in the January 2026 microdata release - a stable primary key that joins cleanly against carrier, rate and facility tables.
- For-hire versus private truck, kept distinct Most freight sources collapse the two; the CFS mode codes separate them, which is the difference between measuring the market and measuring only the part that invoices.
- Weights make it national WGT_FACTOR inflates each sampled shipment to population scale; weighted sums reproduce the published headline totals, unweighted ones do not.
- Verified, not assumed All eighteen dictionary variables were checked against the published data dictionary during the cataloging pass - the same verified standard met by 85.7% of the catalog.
- Sample policy Samples ship in the exact schema shown above, cut to your commodities, lanes and geographies, with code sheets confirmed and populated before anything recurring starts.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.