Datadory notebook

Is rail transportation data free to download? What exists, and how it arrives

Datadory delivers rail transportation data covering the industry's five strata through one pipeline: 302,771 North American track segments carrying owner, trackage rights and corridor designation, roughly 550 US federal rail records whose accident series reaches back to 1975, Eurostat freight tonne-kilometres split by commodity for every EU economy, nine themes of Great Britain official statistics led by 447 million quarterly journeys, and live departures from four European systems - typed rows, delivered daily, weekly, or hourly.

1,744 datasets. Pick your catch.

What is the question actually asking?

It reads like a question about terms and decodes into a question about coverage: someone wants the quantitative core of the rail industry and wants it without a sourcing project attached. The honest answer has two halves.

Yes, the record exists, and it runs deep. Fifteen rail transportation datasets sit in the Datadory catalog - drawn from a universe of 1,744 spanning every industry we index, which averages 7.81 on our quality rubric - and together they cover the five strata the industry actually runs on: physical network geometry, government statistics, national performance shelves, live movement feeds, and the model-training corpora stacked on top. No two overlap on grain, which makes the shelf composable rather than redundant.

What the question usually underestimates is the second-order problem. These records were shaped by four separate publishing cultures - a mapping agency, a federal safety regulator, a European statistics office and train operators running live networks - so polylines arrive in one vocabulary, station calls in another, statistic rows in a third. Getting hold of any single file was never the hard part. Assembling five strata published at four grains into one table that stays correct is. That is the part Datadory sells: the landscape below is what exists; the delivery section is how it lands as rows.

Which named products carry rail transportation data?

Six records do the counting.

NTAD North American Rail Network Lines (quality score 10, one of just 145 perfect scores in the catalog) is the continental base map: 302,771 polyline segments covering all 50 US states, DC, Canada and Mexico in one coordinate frame, each carrying 34 attributes - the reporting mark of the railroad holding operational control, up to nine further marks with trackage rights across it, FRA district, subdivision and branch names, main-track count, measured length, STRACNET strategic-corridor designation and a passenger flag separating Amtrak from commuter, high-speed and tourist operations. Out-of-service and abandoned lines ride in the same record, which is exactly what right-of-way and abandonment audits need.

Data.gov US Rail Catalog (score 9) is the American card index: roughly 547 federal, state and city rail records led by more than 100 Federal Railroad Administration safety programs, including a highway-rail accident family whose incident-level detail reaches back to 1975, beside about 53 Bureau of Transportation Statistics network and terminal layers.

Eurostat Railway Freight Transport Statistics (score 9) owns the European volume picture: 375 billion tonne-kilometres moved in 2024, down 0.8 percent on the year and still chasing the 410 billion peak of 2018. Germany accounts for 126 billion tkm - roughly a third - ahead of Poland at 57 billion and France at 32 billion. Splits run by NST 2007 commodity group with metal ores leading, by national, international or transit haul, and across fourteen dangerous-goods classes; annual series run roughly 2014-2024 with quarterly commodity detail from Q1 2023 onward. The passenger companion answers on the same footing: 429,575 million passenger-kilometres across the EU27 in 2023.

ORR Data Portal - UK Rail Statistics (score 9) is the regulator's own count: nine themes and roughly thirty table families headlined by 447 million journeys, 16.3 billion passenger kilometres and GBP 2.9 billion revenue in January-March 2026, broken out one column per train operating company, with the flagship operator table running back to April 2011. Every table travels with a signed cover sheet - source statement, revision policy, a named responsible statistician.

European Union Agency for Railways (ERA) holds the regulatory registers: the RINF inventory of operational points and line sections, certification databases, and 347 Common Safety Indicator codes tabulating accidents, serious injuries, fatalities and precursors from 2006 through 2024. EU Open Data Portal - Railway Datasets adds breadth - about 6,020 harvested records, including 217 RINF-keyword and 448 ERTMS deployment files - and the World Bank Railways Goods Transported Indicator frames the macro comparison: one verified national figure per country-year in million tonne-kilometres for 100-plus economies, 1995 through 2021, with China topping the panel at 3,018,200.

The operational layer moves with the traffic. National Rail Data Portal (UK) serves live departures resolved to one record per train per station call, alongside daily timetable extracts, station and operator reference and historical service performance across England, Scotland and Wales. SNCF Open Data Portal contributes France: timetables rolling 151 days ahead and annual ridership for about 3,000 stations covering 2015 to 2024. Open Transport Data Switzerland archives a full timetable year across 74 products, preserved in 117 versioned snapshots of the Timetable 2026 release, while DB Open Data and DB Developer Portal documents roughly 5,400 German stations down to facility status. OpenRailwayMap - Global Railway Infrastructure maps the world's railway plant feature by feature across five attribute styles, Transitland indexes thousands of registered transit feeds under permanent identifiers with hash-versioned histories, and the community shelves - Kaggle Railway Datasets Collection and HuggingFace Railway Datasets - hold gigabyte-scale track-fault imagery (the largest verified archive runs about 3.2 GB), the WHU Railway 3D point-cloud benchmark and railway-domain text. Those are fixed snapshots built for computer-vision and language work, not trend panels.

What do delivered rail rows actually look like?

Rows exactly as they land - one unit of record per line, identifier first:

# NTAD North American Rail Network Lines -- one segment between topology nodes
fraarcid : 717196       stateab : BC          country : CA
rrowner1 : KPR          subdiv  : KELOWNA     net : X       miles : 0.2974

# Eurostat Railway Freight Transport Statistics -- country-year observation
geo   : DE              time  : 2024          unit : MIO_TKM
value : ~126000                               # roughly a third of the EU27 total

# ORR Data Portal - UK Rail Statistics -- GB usage, latest reported quarter
table : 1223            period: 2026Q1
journeys : 447000000    passenger_km : 16300000000    revenue_gbp : 2900000000

# SNCF Open Data Portal -- one station-year ridership record
nom_gare : Abbeville    code_uic : 87317362   region : DRG HDF-NORMANDIE
total_voyageurs_2024 : 1024704                non_voyageurs_share : 0.18

# OpenRailwayMap - Global Railway Infrastructure -- one mapped station element
name : London Bridge    railway : station     crs : LBG     platforms : 15    fare_zone : 1

Three properties reward attention before anything gets built on top. First, the identifier spine: an FRA-assigned arc id keys every North American segment, the operator column keys British usage, and CRS and UIC codes key European stations - stable tokens that let geometry meet traffic without fuzzy-matching railroad or station names. Second, the pairing: value and unit travel together, so a tonne-kilometre figure never lands beside a tonnage lift undistinguished. Third, the grain: segments dissolve cleanly by owner, subdivision or corridor, station calls stack into train-level journeys, and the same rows serve a map, a model and a monthly report without re-shaping.

Which cut answers which question?

This shelf refuses to flatten to one unit of record, so start there.

Questions about infrastructure resolve to polylines - 302,771 of them in North America, millions of mapped elements worldwide. Ask who owns the rails under this shipment, which corridors carry STRACNET designation, how much out-of-service track sits inside this right-of-way.

Questions about journeys resolve to station calls and their realised counterparts - per-train records in Britain, stop-level events in Switzerland and France. Ask did punctuality hold after the timetable change or what did this platform actually serve yesterday.

Questions about volume resolve to statistic rows: country-period tonnage and tonne-kilometres by commodity in Europe, operator-columns in Britain, one national figure per economy-year at the macro level. Ask where freight is growing or which market is a third of the EU total.

Questions about material condition resolve to photographs, point clouds and text. Ask can a detector tell a cracked fastener from a shadow. When two candidates tie on all three axes - grain, geography, clock - the scorecard in best rail-transportation datasets settles it.

How deep does each series run?

Five clocks run across the shelf, and knowing which one governs your question saves a week of false starts.

  • Safety history compounds across five decades: the FRA highway-rail accident series inside the federal catalog reaches back to 1975, while ERA's Common Safety Indicators tabulate country-year outcomes from 2006 through 2024.
  • Passenger usage runs operator-deep: ORR's flagship table spans April 2011 to March 2026, and SNCF ridership covers about 3,000 stations from 2015 to 2024 - long enough for structural stories, recent enough for last quarter.
  • Freight volume is young at high resolution: Eurostat annual series run roughly 2014-2024, but the quarterly NST 2007 commodity detail begins in Q1 2023 - so a panel joining quarterly commodity splits to the longer annual histories runs three years shorter than its annual leg. Build schemas knowing which legs are deep and which are young.
  • Geometry is a maintained snapshot, not a time series: the North American layer was first compiled in 2016 and answers where the track is and who controls it now, not how ownership evolved. Treat it as the spine other facts hang from.
  • Timetables are archaeological: Switzerland's 117 versioned snapshots preserve a full timetable year, France rolls 151 days ahead, and Britain serves today's forecasts alongside historical performance periods.

Geography sets the last boundary honestly. Only the geometry layer spans borders cleanly - United States, Canada and Mexico in one file - and only the World Bank indicator reaches every economy at once. The national shelves go deeper inside their fences: Britain for performance and finance, France for ridership, Germany for facilities, Switzerland for timetable depth, Brussels for cross-country freight comparisons.

Who builds on rail transportation data?

  • Transportation economists and policy analysts ground infrastructure-spending arguments in measured track mileage by owner, state and network class, with STRACNET tagging adding the defense-mobilization angle to resilience studies.
  • Logistics and intermodal planners read ownership and trackage rights off a lane before negotiating - including the tenant railroads whose consent a deal quietly requires.
  • Insurance, real estate and right-of-way teams tie segments to parcels, easements and abandonment exposure along a corridor, using the out-of-service flags most maps silently omit.
  • Telecom and utility route designers rank rail rights-of-way as preferred conduit paths before spending on field survey.
  • Investors and analysts test an operator's claims against the regulator's own count: 447 million journeys, GBP 2.9 billion revenue, one column per train operating company.
  • GIS and mapping teams drop survey-grade linework straight into basemaps, dashboards and routing prototypes across three countries.

The personas who lean hardest: data scientists turning a continent-wide graph with categorical richness into centrality and routing models - see our data scientists take on rail transportation - market researchers anchoring sizing work on defensible mileage and ridership figures, and developers and builders wiring mapping products against identifier-stable rows.

Why get rail transportation data through Datadory?

Because the hard part was never locating a fact about rails - it is holding five strata together in one table that stays correct. Polylines arrive in one vocabulary, station calls in another, statistic rows in a third; identifiers get renamed between editions; units drift; and a corridor audit dies quietly when segment totals are summed over the wrong filter. Datadory handles the reconciliation upstream of you: normalized encodings, units kept beside values, the FRA arc id preserved as the join spine, operator columns and CRS/UIC codes carried through untouched, and every delivery shipping the complete field dictionary, sample rows and a coverage statement naming geography, vintage and granularity. When a new snapshot arrives, it lands as new rows under the same fields - no integration project repeats.

Start with a sample: name the states, subdivisions, corridors, operators or countries you need, and the extract comes cut to them - request a sample and compare real rows before anything is committed.

The rail transportation shelf compared - what each named product contributes (catalog position as of August 2026)
DatasetUnit of observationCoverage depthBest used for
NTAD North American Rail Network LinesOne polyline segment between topology nodes: owning railroad mark, co-owners, up to nine trackage-rights holders, FRA district, subdivision and branch, main-track count, measured length, STRACNET flag, passenger code302,771 segments across all 50 US states, DC, Canada and Mexico; compiled 2016, maintained as a current snapshotNetwork geometry, capacity and market-entry work, right-of-way and abandonment audits
Data.gov US Rail CatalogOne normalized record per federal, state or city rail dataset: program title, publishing agency, coverage description~547 records; more than 100 FRA safety programs with accident series from 1975, ~53 BTS network and terminal layersDiscovery across US rail programs and five decades of safety history
Eurostat Railway Freight Transport StatisticsCountry-period observation: tonnes lifted and tonne-kilometres by NST 2007 commodity, haul type, dangerous-goods classEU27 plus EFTA and candidate countries; annual ~2014-2024, quarterly Q1 2023-Q4 2024Cross-country freight benchmarking and modal-shift analysis
ORR Data Portal - UK Rail StatisticsOperator-columned usage, performance, finance and safety cells with provisional markers and signed cover sheetsNine themes, ~30 table families; flagship operator table April 2011-March 2026Regulator-grade GB performance and accountability reporting
European Union Agency for Railways (ERA)Operational points and line sections from the infrastructure register, certification records, country-year safety indicators347 Common Safety Indicator codes, 2006-2024, across member states plus CH, NO and UKInteroperability compliance and European safety comparison
World Bank Railways Goods Transported IndicatorOne verified national figure per country-year: rail freight in million tonne-kilometres100-plus economies, 1995-2021, compiled from UIC Railisa and OECD StatisticsMacro context and exogenous features for country models
National Rail Data Portal (UK)One record per train per station call: due, expected and actual times, platform, plus timetable, reference and incident recordsTens of thousands of daily GB movements across England, Scotland and WalesLive departure products and historical punctuality studies
SNCF Open Data PortalStops, routes and trips in scheduled formats plus station-year ridership with UIC codes and non-traveller shareTimetables rolling 151 days ahead; ridership for ~3,000 French stations 2015-2024French journey planning and station catchment sizing
OpenRailwayMap - Global Railway InfrastructureIndividual mapped features: tracks, stations, yards, signals, electrification, speed sections with CRS/UIC-coded stationsWorldwide wherever rail is mapped; five attribute stylesGlobal base-map context and station geocoding

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Rail Transportation All 50 US states plus District of Columbia

NTAD North American Rail Network Lines

Rail Transportation United States, with several BTS and FRA series extending into…

Data.gov US Rail Catalog

Rail Transportation EU27 aggregate plus EU

Eurostat Railway Freight Transport Statistics

Rail Transportation Great Britain national rail network, with operator-level…

ORR Data Portal - UK Rail Statistics

Rail Transportation EU member states plus Switzerland

European Union Agency for Railways (ERA)

Rail Transportation All EU Member States plus EFTA and candidate countries…

EU Open Data Portal - Railway Datasets

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

Which dataset shows who owns a stretch of track?

NTAD North American Rail Network Lines. Each of its 302,771 segments carries the reporting mark of the railroad holding operational control, co-owner columns, up to nine further marks with trackage rights, the FRA district, subdivision and branch names, main-track count, measured length, STRACNET strategic designation and a passenger flag. Reading only the first owner understates how shared a busy corridor really is.

Where does European rail freight stand right now?

375 billion tonne-kilometres moved in 2024, down 0.8 percent on the year and chasing the 410 billion peak of 2018. Germany accounts for roughly a third - 126 billion tkm - ahead of Poland at 57 billion and France at 32 billion. Splits run by NST 2007 commodity group led by metal ores, by haul type, and across fourteen dangerous-goods classes, with quarterly series tracking 2023 and 2024.

How far back do rail statistics reach?

Deepest on safety: the US highway-rail accident series runs from 1975 onward, and ERA's Common Safety Indicators tabulate 2006 through 2024. Passenger usage spans April 2011 to March 2026 in Britain. Eurostat annual freight series run roughly 2014-2024, with quarterly commodity detail beginning Q1 2023, and the World Bank indicator covers 100-plus economies from 1995 through 2021.

Can rail geometry be joined to traffic and incident records?

Yes - the rows are designed for it. Every North American segment keys on an FRA-assigned identifier with topology nodes preserved, so ownership, trackage rights and corridor designation resolve through one token and segments dissolve cleanly by owner, subdivision or corridor. British usage keys on the operator column, European stations on CRS and UIC codes, which lets geometry, ridership and incident references meet without name-matching.

How is rail transportation data delivered?

As typed rows - API, files, or straight into your warehouse, daily, weekly, or hourly, your call. Every delivery carries the complete field dictionary, sample rows for the records you name, and a coverage statement fixing geography, vintage and granularity. New snapshots accrue as new rows under the same fields, so nothing re-integrates.