Datadory notebook
County-to-county commuting flows: every home-work pair, delivered as typed rows
Datadory delivers passenger ground transportation data covering America's county-to-county commuting flows: worker pairs linking each residence county to its workplace county, with mode share across drove-alone, carpool, transit, walk and bike, travel-time bands and vehicles-available cross-tabs attached, from block group to nation across two decades of annual releases - delivered daily, weekly, or hourly.
1,744 datasets. Pick your catch.
What county-to-county commuting flows actually measure
A commuting flow is a count of workers aged 16 and over who live in one county and work in another: an origin-and-destination matrix, not a ridership total, a fare ledger or a timetable. It answers the question every ground-transportation model eventually asks - how many people does this corridor have to move, and in which direction - before a single boarding is counted.
The grain runs deeper than the name implies. Five-year estimates cover every geography down to the block group - roughly 85,000 tract geographies and about 840 county equivalents in the current 2020-2024 release - while one-year estimates stop at populations of 65,000 and over. Corridor work touching rural or micropolitan counties lives in the five-year file; metro trend work takes the 2024 one-year release alongside it.
What does a delivered flow row look like?
Rows arrive typed and keyed to named geographies, every estimate travelling with its margin-of-error twin. The verified frame:
# Table B08301 - Means of Transportation to Work (verified variable set)
B08301_001E Estimate!!Total: -> int
Total workers 16 years and over
B08301_003E Estimate!!Total:!!Car, truck, or van:!!Drove alone -> int
B08301_004E Estimate!!Total:!!Car, truck, or van:!!Carpooled -> int
B08301_010E Estimate!!Total:!!Public transportation (excluding taxicab) -> int
B08301_011E ...bus specifically, one tier beneath the transit subtotal
B08301_019E Estimate!!Total:!!Walked -> int
# Reliability ships beside every count
B08301_001E = n <-> B08301_001M = +/-m
# The flow layer puts two geographies on one row
residence county x workplace county -> workers making that trip, margin attachedThree properties decide whether these rows survive contact with a real model. Direction is recorded: the A-to-B count and the B-to-A count are different rows, because morning and evening demand are different questions. Reliability is recorded: a thin county-pair cell declares its uncertainty instead of posing as fact. And the vocabulary never changes with scale: the same columns answer nationally, by metropolitan area, by county and by tract, so a panel built once re-levels without remapping anything.
Get a sample of commuting flow data cut to your corridors and counties - the field dictionary travels with it.
Which fields carry the analytical weight?
The load-bearing columns stack into three jobs: measurement, identification and scope.
| Field or table | What it holds | Why it matters |
|---|---|---|
| B08301_001E | Total workers 16 years and over | The denominator every share divides against - counted, not modelled |
| B08301_003E / B08301_004E | Drove alone; carpooled | The substitutable pool any carpooling or commuter-benefits case has to beat |
| B08301_010E / B08301_011E | Public transportation excluding taxicab; the bus tier beneath it | Transit demand decomposed into modes an operator can act on |
| B08301_019E | Walked | The walk-access catchment hiding inside the mode split |
| Table B08303 | Travel time to work banded by minutes | Corridor and labour-shed sizing without a routing engine |
| Table B08134 | Vehicles available crossed against mode | Separates carless riders from optional ones - the distinction equity arguments turn on |
| B07404 / B07410 series | Residence-year-to-residence-year flows | Movement between vintages, so migration is separable from commuting |
Beyond the dictionary sit the remaining B08301 cells - bicycle, motorcycle, taxicab, worked from home, plus the subway, long-distance train, light rail and ferryboat tiers - and person-level microdata for cross-tabulations nobody has published yet. Name them when requesting a sample and they arrive populated, documented against the delivered extract.
How do you turn flow counts into catchment and market sizing?
Labour-shed and trade-area sizing. Travel-time bands describe who can reach a site and how long it takes them; vehicles available says whether they own the car that would. Site-selection teams weigh candidate locations on distributions instead of map guesses.
Corridor cases. Origin-destination flows paired with mode share turn "this train seems busy" into a defensible service case: how many workers make the trip, which modes absorb them today, and what substitution rate a new option implies.
Market sizing for mobility products. Drove-alone counts at fine geography size the pool a carpooling app, commuter-benefits platform or micromobility launch can realistically convert - computed before a dollar is spent, off the same rows the business case will later be audited against.
Can daily ridership validate a flow-based forecast?
Yes, and the cleanest test case is commuter rail around New York, where MTA Daily Ridership Data: Beginning 2020 tracks nine counters day by day - Subway, Bus, LIRR, Metro-North, Staten Island Railway, Access-A-Ride, bridges and tunnels traffic, plus the congestion-zone entries added in January 2025 - in roughly 17,700 rows covering March 1, 2020 through August 19, 2026, one row per mode per date. Because LIRR and Metro-North serve exactly the long-distance county pairs that dominate New York's inflow matrix, a corridor-level flow estimate can be checked against actual boardings on the same geography.
Where the two disagree, the suspects are knowable in advance. Vintage mismatch first: five-year files blend five survey years, so treat them as rolling averages rather than point-in-time counts when quoting a single year. Stale assumptions second: a share built on an old mode split will miss a corridor whose behaviour moved.
What calibrates the behaviour behind the flows?
National Household Travel Survey (NHTS) supplies the behavioural half of the conversion. Nine waves reach back to 1969, and the 2022 public-use core holds 7,893 households, 16,997 persons and 31,074 weighted travel-day trips, each carrying mode codes, home-based-work purpose flags, trip distances, durations and the full-file survey weights. It explains why people travel the way the flow matrix says they do.
Its boundary matters as much as its content: household geography stops at region, division, CBSA and state FIPS, so NHTS calibrates the commuting shares rather than replacing them as a second flow matrix. Run together - flows for volume and direction, survey microdata for behaviour and purpose - and a demand model inherits both the population and the motives. The fuller workflow sits on the data scientists use cases page.
Is there a non-US equivalent for benchmarking?
Not at county grain, no. The nearest structural analogue is Eurostat Passenger Transport Statistics (tran_hv_psmod / rail_pa), whose flagship table reports the share of inland passenger-kilometres taken by trains, buses and coaches, and cars annually from 1990 through 2024 across 37 geographies including the EU27 aggregates and EFTA countries. Harmonized methodology, national level only - no European record in the slice pairs one home region with one workplace region.
That difference belongs in any deliverable. US county-to-county flows answer how many people travel from A to B for work; Eurostat answers what fraction of all travel happens by each mode. Pairing the two lets a study state both how large a corridor is and how its modal mix compares internationally - but neither substitutes for the other, and mixing their units is the most common presentation error in cross-market decks.
Who builds on county-to-county commuting flows?
Transit planners and network developers read mode share at tract and block-group resolution to find where dependence already concentrates, then argue corridors from origin-destination flows instead of anecdote.
Site-selection and location-strategy teams describe labour catchments around candidate sites - who can reach them, how long it takes, whether they own a car - before committing capital; see the market researchers use cases page.
Investors and quant analysts read commuting structure as a slow indicator behind real estate, logistics and workforce exposure; the workflow lives at investors quants use cases.
Data scientists and ML engineers get typed panels whose margins arrive attached - defensible features for catchment, demand and siting models; see data scientists use cases.
Policy and equity researchers cross vehicles available against mode to separate households that ride transit by choice from those that ride it because nothing else exists.
Commute behaviour moves over years, not weeks, so a panel built on these rows stays defensible without constant rebuilding - a rarity in any data catalog.
Why get commuting flows through Datadory?
Because the hard part was never the first number - it is the tenth. Choosing the wrong release for the question's geography. Margins stripped by a naive pipeline until thin cells pose as facts. Flow tables that refuse to line up with the marginal mode tables without a mapping nobody documents. Vintages stitched mid-series with silent breaks. Each is survivable once; none is fun to re-solve in every new project.
Datadory handles them upstream of delivery: releases matched to the question, margins travelling beside their estimates, join keys stable across vintages, flow pairs and mode splits landed in one consistent frame - as API responses, scheduled files, or warehouse-native tables.
Name your counties and corridors when you request a sample and it arrives already cut to that scope. The production feed follows the same shape, so whatever you prototype on the sample survives delivery intact.
Where should you go next?
Start with American Community Survey - Journey to Work / Commuting Data sampled to your corridors; product-level detail - sample rows, the full field dictionary, coverage chips - lives on its dataset page.
This page is one thread of a wider map. The passenger ground transportation data guide walks all fourteen cataloged records in the slice, the passenger-ground-transportation data hub browses them as products, and the best passenger ground transportation datasets ranking sorts the leaders side by side. For neighboring threads: the commuting mode share by census tract walkthrough covers the tract-level frame, and transit ridership by agency monthly explains the boardings record those flows feed. When descriptions stop and real rows need to start, get a sample scoped to your counties - the field dictionary travels with it.
| Dataset | Unit of observation | Geography in output | Temporal depth | What it adds |
|---|---|---|---|---|
| American Community Survey - Journey to Work / Commuting Data | Worker counts per geography, plus county-to-county flow rows pairing origin and destination; person-level microdata alongside | Every U.S. county equivalent and tract, down to block groups in the 5-year release | Annually since 2005; current releases 2024 1-year and 2020-2024 5-year | The flow matrix itself, with the mode split, travel-time bands and vehicles-available cross-tabs that convert counts into markets |
| National Household Travel Survey (NHTS) | One row per household, person, vehicle and trip, weighted to national totals | Household geography stops at region, division, CBSA and state FIPS | Nine waves, 1969-2022; 31,074 weighted travel-day trips in the 2022 core | Behavioural calibration - mode codes, trip purpose and weights explaining why the flows look the way they do |
| National Transit Database (NTD) Monthly & Annual Ridership Datasets | One row per agency x mode x type-of-service x month | Agency-reported service across the United States, 834 NTD IDs | January 2002 into June 2026, roughly 370,000 rows | The supply-side denominator any modal-share or substitution claim divides against |
| MTA Daily Ridership Data: Beginning 2020 | One row per mode per date across nine counters including LIRR and Metro-North | New York systemwide by mode, including the congestion-zone entries added January 2025 | March 1, 2020 through August 19, 2026, roughly 17,700 rows | Day-by-day validation for the long-distance county pairs dominating New York's inflow matrix |
| CTA Ridership - Daily Boarding Totals | One row per service date, split bus versus rail | Chicago systemwide | January 1, 2001 through June 30, 2026, 9,312 daily observations | Two decades of daily boardings for validating Midwest corridor assumptions |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
American Community Survey - Journey to Work / Commuting Data
NAME
MTA Daily Ridership Data: Beginning 2020
CTA Ridership - Daily Boarding Totals
service_date · day_type · bus …+2 more
National Transit Database (NTD) Monthly & Annual Ridership Datasets
ntd_id · agency · reporter_type …+7 more
National Household Travel Survey (NHTS)
Eurostat Passenger Transport Statistics (tran_hv_psmod / rail_pa)
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
What is a county-to-county commuting flow?
A count of workers aged 16 and over who live in one county and work in another - an origin-and-destination matrix rather than a ridership total, fare ledger or timetable. The American Community Survey has published the pairing annually since 2005, with five-year releases reaching block-group geography and one-year releases covering populations of 65,000 and over.
Which release should a corridor analysis use?
The 2020-2024 five-year release when the question touches rural or micropolitan counties, because it covers every geography regardless of population; the 2024 one-year release when currency matters more than resolution. Teams needing both run them side by side and treat the five-year figures as rolling averages rather than point-in-time counts.
Do the flows include mode split and reliability figures?
Yes. Mode cells decompose workers into drove alone, carpool, transit excluding taxicab with the bus tier beneath, walk, bicycle and work from home, and every estimate ships with its margin-of-error twin, so a thin county-pair cell admits its uncertainty instead of posing as precision.
Can commuting flows validate a ridership forecast?
Yes - pair a corridor's flow estimate with observed boardings on the same geography: the MTA's nine daily counters since March 2020 for New York rail corridors, Chicago's daily bus-versus-rail series since 2001, or the National Transit Database's roughly 370,000 agency-by-mode-by-month records as the systemwide denominator. Where the two disagree, suspect vintage mismatch or a stale mode-split assumption first.
What does a Datadory sample include?
Real rows cut to the counties, corridors and tables you name, with the field dictionary attached and margins intact; additional tables - travel time, vehicles available, county flows, person-level microdata - arrive populated on request. Samples precede any commitment, and the production feed follows the same shape, delivered daily, weekly, or hourly.