For Data Scientists & ML Engineers · Passenger Ground Transportation

Passenger Ground Transportation Data for Data Scientists

Passenger ground transportation data for data scientists starts with the NYC TLC trip records, the National Transit Database and the Mobility Database's GTFS catalog.

financial time series api for backtesting · alternative data for quantitative research · where to get training data for passenger ground transportation models

14datasets cleared the bar for this shelf
12rated top-tier for this persona
9.1mean quality, our 10-point scoring

API, files, or your warehouse. Daily, weekly, or hourly.

What passenger ground transportation data can data scientists actually train on?

Datadory qualifies 14 passenger ground transportation datasets for data scientists and ML engineers: 12 score maximum relevance for modeling work, two more clear the bar as discovery layers. Quality runs deep for a public-data domain: a 9.07-of-10 slice average beats the 7.81 catalog-wide mean, with 13 of 14 records scoring 8 or higher versus 62.8% catalog-wide.

The 14 best passenger ground transportation datasets for machine learning

Ordered by relevance to modeling work first, then Datadory's 0-10 quality score. The one-line verdict under each name is written for pipeline builders.

2. National Household Travel Survey (NHTS) — Train travel-behavior models on weighted household, person, vehicle and trip microdata in multi-format files. Access runs over bulk delivery with static refresh (csv, sas, spss, xpt); quality 10/10.

3. National Transit Database (NTD) Monthly & Annual Ridership Datasets — Join monthly unlinked passenger trips, VRM and peak vehicles by agency and mode back to 2002 into demand forecasts; commercial delivery terms clears training use.

4. NYC TLC Trip Record Data (Yellow/Green Taxi & FHV/HVFHV Uber/Lyft) — Train trip-duration, fare and demand models on billions of trip-level Parquet rows from 2009 onward — the canonical NYC mobility benchmark. Access runs over bulk delivery with monthly refresh (parquet, pdf, shapefile, csv); quality 10/10.

7. Chicago Taxi Trips — Model fares, demand and spatial patterns from trip-level rows (2013-2023 plus 2024-present) with census tracts.

8. Chicago Transportation Network Providers Trips (Uber/Lyft/Via) — Train fare and demand models on per-trip Uber/Lyft/Via rows across three era datasets (2018-present).

9. CTA Ridership - Daily Boarding Totals — Fit daily seasonality models on bus/rail boardings from January 2001 with day-type flags.

11. MTA Daily Ridership Data: Beginning 2020 — Model daily recovery curves across subway, bus, commuter rail and congestion-zone entries from March 2020 onward.

13. UK Bus Statistics & Bus Open Data Service (BODS) — Model UK bus patronage and mileage from annual BUS01-BUS09 tables with history back to 1950; OGL v3.0 permits training use. Access runs over bulk delivery with quarterly refresh (ods, csv, xml, siri-vm); quality 9/10.

14. Data.gov GTFS Catalog — Discover ~90 GTFS-related federal catalog records before ingesting individual agency feeds.

How do you combine these sources into one modeling stack?

Compose around grain, then join keys. For demand targets, join National Transit Database (NTD) Monthly & Annual Ridership Datasets monthly unlinked trips by agency and mode to CTA Ridership - Daily Boarding Totals and MTA Daily Ridership Data: Beginning 2020 dailies, and let the American Community Survey - Journey to Work / Commuting Data PUMS microdata contribute commute mode share as covariates. European breadth comes from Eurostat Passenger Transport Statistics (tran_hv_psmod / rail_pa) modal-split tables and UK Bus Statistics & Bus Open Data Service (BODS) BUS01-BUS09 patronage history back to 1950; BTS Open Data Portal - Passenger Travel Collection adds agency-day ridership and port-month border crossings over Socrata when cross-border context helps. The known gap: nothing here links individual travelers across modes, so identification stays ecological unless you bring device-level data.

Which sources should your pipeline start with?

Pick by task. Fare and duration benchmarks: TLC Parquet, Chicago taxi and TNP tables. Demand forecasting: NTD monthly joined to CTA and MTA dailies. Behavior modeling: NHTS weighted microdata with ACS Journey-to-Work covariates.

Straight answers

Where can I get training data for ridership forecasting models?

Join the National Transit Database's monthly unlinked trips, VRM and peak vehicles by agency and mode back to 2002 with CTA daily boardings since January 2001 and MTA daily ridership from March 2020 onward.

Is there NYC taxi trip data with bulk Parquet download?

Yes — NYC TLC Trip Record Data distributes yellow, green and FHV/HVFHV trip records as monthly Parquet files from 2009 onward: billions of rows covering pickup and drop-off times, locations, distances, fares and tips.

Do Uber and Lyft trip-level datasets exist for modeling?

Chicago publishes them. The Transportation Network Providers tables cover per-trip Uber, Lyft and Via rows across three era datasets from 2018 to present, including fares, tips, shared-ride flags and census tracts, served through Socrata in CSV, JSON, GeoJSON and XLSX with daily refresh.

Rows before rollout

Sample rows from any shelf entry — the field dictionary and coverage notes ride along. If the shelf misses what you need, say so; sourcing requests are half our job.

Talk to us