Datadory notebook

Transit Ridership by Agency, Monthly: The National Panel, Delivered

1,744 datasets. Pick your catch. Every guide here is built on what the catalog can actually prove.

1,744 datasets. Pick your catch.

What is transit ridership by agency, monthly?

The query sounds modest - boardings, by operator, by month - and exactly one record answers it at national scale: National Transit Database (NTD) Monthly & Annual Ridership Datasets, the measurement layer of American public transit. Every operator that reports to the Federal Transit Administration files into it, because Congress requires them to. That makes it a census, not a survey: no sampling frame, no opt-in panel, no recall bias - every reporting agency appears.

The grain is the query's grain, verbatim. One row per agency x mode x type-of-service x month, carrying four measures: unlinked passenger trips, vehicle revenue miles, vehicle revenue hours and vehicles operated in maximum service. The complete monthly module verifies at roughly 370,000 rows across 834 distinct agency identifiers, running January 2002 through the most recent reported month (June 2026 at research time).

Quality checks out too. It scores a perfect 10 on our field-documentation rubric - a ceiling shared by just 145 of the 1,744 datasets we catalog - and it anchors a passenger ground transportation slice of 14 primary records where five hit the same mark. Get a sample of transit ridership data cut to your agencies, modes and months before anything else; everything below shows what arrives.

What does a monthly ridership row look like when delivered?

Illustrative rows in the delivered shape - one California operator, two modes, one recent month:

# City of Fairfield, California - ntd_id 90092 - month 2026-06
mode=MB   tos=PT   _3_mode=Bus      # fixed-route bus, purchased transportation
upt=9,284     vrm=24,788    vrh=1,990    voms=12

# same agency, same month, demand response
mode=DR   tos=PT
upt=1,781     vrm=19,406    vrh=1,900    voms=10

# Redding Area Bus Authority - ntd_id 90093 - Small Systems Reporter
mode=DR   tos=PT   month 2026-06
upt=7,351     vrm=43,122    vrh=2,540    voms=22

Read the anatomy rather than the digits - each value above is the documented example for its column. Three habits separate careful users from casual ones. First, upt counts unlinked trips: a rider who transfers boards again, so summing across modes overcounts unique people. Any total worth quoting states whether it is boardings or passengers. Second, the efficiency ratios come free of the rows themselves - trips per revenue mile, trips per revenue hour - because vrm and vrh ride beside the count on every line. Third, voms is the fleet signal: vehicles operated in maximum service, the number that sizes an operator's network before anyone reads the ridership at all.

One operator contributing several rows per month is the design, not noise. Fixed-route bus and demand response behave like different businesses wearing the same logo, and the mode ladder keeps them separable down to nineteen codes.

Which fields carry the analytical weight?

All thirteen documented monthly fields verify against the official per-column dictionary - the reason this record sits at the top of our documentation rubric against a catalog average of 7.81 across 1,744 datasets.

FieldWhat it holdsWhy it matters
ntd_idFTA-assigned identifier every agency must hold before filingThe join key across months, years and companion tables
agencyLegal name of the transit propertyDisplay, deduplication, account mapping
mode / _3_modeNineteen operational codes, grouped into rail, bus and otherMode-share curves and national roll-ups off one column pair
tosDirectly operated versus purchased transportationSeparates in-house fleets from contracted service
dateMonth of observationThe spine of any time-series model
uptUnlinked passenger tripsThe ridership measure itself
vomsVehicles operated in maximum serviceFleet size and utilization signal
reporter_typeFull Reporter, Small Systems, Rural and kinDecides what your totals actually cover
uza_name / stateUrbanized area and headquarters stateMetro-scoping and regional aggregation without geocoding

How far does the panel reach, and what does it cover?

  • Geographic: the United States end to end - 834 distinct agencies in the monthly module, keyed to urbanized area and state, drawn from the roughly 900-plus operators that report annually overall.
  • Temporal: January 2002 through the most recent reported month, June 2026 at research time - a quarter-century panel that walks through the pre-pandemic peak, the 2020 collapse and the uneven recovery month by month. Annual metric tables cover report years 2022-2024, with older report years held in the program archive.
  • Granularity: one row per agency x mode x type-of-service x month. Long format by design, which rewards forecasting and punishes spreadsheet eyes - more on that below.

One boundary deserves stating up front rather than discovered mid-analysis: the monthly module carries Full Reporters. Agencies running fewer than 30 peak vehicles and rural reporters sit outside it, their estimated figures living in the companion annual release instead. Roughly 900-plus agencies file overall against 834 in the monthly table, and the gap is precisely the small systems that dominate agency counts in rural maps. Ask for total-industry figures and the sample merges both releases; the caveat travels documented either way.

Which ridership ledger fits which job?

The federal census answers the national question, but five sibling records in the same slice answer adjacent ones better, and they divide cleanly by grain and geography. Lined up as of August 2026:

Can you get below monthly?

Yes, from two city systems, and each changes what a trend means.

CTA Ridership - Daily Boarding Totals compresses Chicago into 9,312 rows - one per service date since January 2001, split bus versus rail, flagged weekday, Saturday or Sunday-holiday. Twenty-five years of seasonality open in a spreadsheet, observed through June 30, 2026, which makes it the fastest way to prototype a day-type model before scaling to the national panel.

MTA Daily Ridership Data: Beginning 2020 goes finer still: nine modes - subway, buses, LIRR, Metro-North, Staten Island Railway, Access-A-Ride, bridges and tunnels, plus the two congestion-pricing counters - about 17,700 long-format rows covering March 1, 2020 onward. It is the only cataloged series where post-2020 recovery and congestion-zone effects resolve mode by mode, day by day.

Both join to the federal census on agency identity plus calendar date, so a monthly national claim can be stress-tested against daily city evidence without reshaping anything.

How does the American panel read internationally?

Direct agency-to-agency comparison stops at the border, because no other country publishes the agency-x-month grain at national scale. Two cataloged records give the international frame:

A practical pattern: build the US panel from the NTD, normalize per urbanized-area population using American Community Survey Journey to Work tables to size each agency's catchment, then set the result against Eurostat modal shares so the American trend reads in a European context. Three records, one argument.

Who builds on agency-month ridership data?

Data scientists and ML engineers inherit a quarter-century panel with built-in seasonality and a documented structural break - training substrate for demand models that need real shock absorption; see data scientists use cases.

Investors and quant researchers read monthly ridership curves as urban-activity factors - footfall proxies for mobility, real-estate and retail theses, keyed to metro rather than sentiment; the workflow lives at investors quants use cases.

Market researchers and consultants size US public-transit markets by mode and metro off a congressionally mandated census rather than a vendor's opt-in panel; see market researchers use cases.

Competitive intelligence teams rank operators by unlinked trips to prioritize fare-collection, fleet-tech and service-planning outreach, with reporter_type and voms sizing each account before the first call; see competitive intel product teams use cases.

Journalists and academics cite filed figures whose provenance is a federal reporting mandate - the strongest kind of attribution a mobility claim can carry.

How is transit ridership data delivered?

Shape is part of delivery, not your problem afterward. The native form is long and loves time-series work; samples ship pivoted wide per agency on request - one row per agency and month, a column per mode, keys preserved so both shapes still join. Mode codes arrive decoded beside their spelled-out names. Annual metric families arrive joined onto the monthly spine where the question needs money attached to boardings.

Because deliveries repeat on a schedule, each month's report becomes another point on a panel rather than another file to reconcile - captures accumulate under stable keys, and restatements stay traceable to the extract they came from.

Why get transit ridership data through Datadory?

Because the hard part was never the first extract - it is the tenth. Long-format rows that punish anyone expecting one column per mode. Mode codes (MB, DR, VP) that need a decoder ring. Annual metrics living in separate tables that need joining onto the monthly spine. Reporter classifications that quietly decide whether your totals cover the whole industry or only its largest players. Small-system exclusions that surface three notebooks deep, right when a rural map stops adding up.

Each is survivable once; none is fun to re-solve in every new project. Datadory handles them upstream of delivery: pivots cut to the shape your model wants, dictionaries traveling with every extract, annual families joined where useful, coverage boundaries surfaced explicitly rather than discovered mid-analysis.

Name your agencies, modes and months when you request a sample and it arrives already cut to that scope - the production feed follows the same shape, so anything prototyped on the sample survives delivery intact.

Where to go next

Start with the passenger ground transportation data guide, which maps all fourteen cataloged records in the industry and shows where the ridership census sits beside trip-level ledgers and feed catalogs. The passenger-ground-transportation data hub browses the same catalog as product pages with field dictionaries and sample rows, and the best passenger ground transportation datasets ranking scores the top ten side by side.

For depth on the threads named here: Mobility Database vs National Transit Database compared weighs the world's GTFS feed discovery layer against the federal measurement layer, the Mobility Database - Global GTFS & GTFS-Realtime Feed Catalog covers the supply side of the same networks, and the National Transit Database (NTD) Monthly & Annual Ridership Datasets product page documents every field row by row.

Ridership ledgers compared (Datadory catalog, as of August 2026)
DatasetRow grainCoverageSettles questions like
National Transit Database (NTD) Monthly & Annual Ridership DatasetsAgency x mode x type-of-service x month (~370,000 rows, 834 NTD IDs)January 2002 through the most recent reported month; annual metrics for report years 2022-2024National agency benchmarking, demand forecasting, modal share
CTA Ridership - Daily Boarding TotalsSystemwide per service date, bus versus rail (9,312 rows)January 2001 through June 30, 2026 observed, day-type flaggedWeekday/weekend seasonality and city-scale prototyping
MTA Daily Ridership Data: Beginning 2020Per mode per date, nine modes (~17,700 rows)March 1, 2020 to present, including the two congestion-pricing countersPost-2020 recovery and congestion-zone tracking mode by mode
BTS Open Data Portal - Passenger Travel CollectionAgency-day transit ridership; port-month border entriesDaily transit ridership from 2022 onward; border crossings back to 1996 through July 2026Recent daily counts with cross-border travel context
Eurostat Passenger Transport Statistics (tran_hv_psmod / rail_pa)Country x mode x year (37 geographies)Annual modal split 1990-2024; rail passenger volumes to 2025European cross-country benchmarking against the US panel
UK Bus Statistics & Bus Open Data Service (BODS)Country/region/local authority x year (BUS01-BUS09)Annual from 2004-05 with historical journeys reaching toward the 1950s; plus 7,207 timetables and 949 vehicle-location feedsGreat Britain bus patronage, fares and concessionary travel

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Passenger Ground Transportation United States - all reporting transit agencies, keyed to…

National Transit Database (NTD) Monthly & Annual Ridership Datasets

ntd_id · agency · reporter_type …+7 more

Passenger Ground Transportation City of Chicago and the CTA service area across Cook County…

CTA Ridership - Daily Boarding Totals

service_date · day_type · bus …+2 more

Passenger Ground Transportation Metropolitan Transportation Authority service region: New York…

MTA Daily Ridership Data: Beginning 2020

Passenger Ground Transportation All U.S. states, counties, places, tracts and block groups

American Community Survey - Journey to Work / Commuting Data

NAME

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

What is the best dataset for monthly transit ridership by agency?

The National Transit Database's machine-readable monthly module. It verifies at roughly 370,000 rows across 834 distinct agency identifiers - one row per agency, mode, type of service and month, from January 2002 into the most recent reported month, with unlinked passenger trips, vehicle revenue miles, revenue hours and peak vehicles on every row. Datadory delivers it typed, decoded and joined to its annual companions.

Does the monthly panel include small and rural transit systems?

Not all of them. Agencies reporting fewer than 30 vehicles in maximum service and rural reporters sit outside the monthly module; their estimated figures live in the companion annual release. Roughly 900-plus agencies file to the program overall against 834 in the monthly table. Ask for total-industry figures and the sample merges both.

Can I get daily ridership instead of monthly?

Two city systems supply it. Chicago's daily boarding totals hold 9,312 observations since January 2001 split bus versus rail with a day-type flag, and New York's daily series tracks nine modes - subway, bus, commuter rail, Staten Island Railway, Access-A-Ride, bridges and tunnels and the two congestion-pricing counters - one row per mode per date from March 2020 to present.

Can the monthly table be shaped wide per agency?

That is the standard cut. The native shape is long - one row per agency, mode, type of service and month - which suits time-series models and frustrates spreadsheet eyes. Samples ship pivoted wide on request: one row per agency and month with a column per mode, keys preserved so both shapes still join.