Rail Transportation
SNCF Open Data Portal
Datadory delivers SNCF Open Data Portal data covering national TGV, Intercités and TER timetables in GTFS and NeTEx rolling 151 days ahead, real-time trip updates and service alerts, annual ridership for roughly 3,000 French stations from 2015 to 2024, and the network-infrastructure layers beneath them - 166 catalogued datasets normalised into one production feed.
API, files, or your warehouse. Daily, weekly, or hourly.
- Where it covers
- France nationwide - TGV, Intercités and TER scope for the national timetables, Ile-de-France for the Transilien feed, roughly 3,000 stations for ridership
- How far back
- Timetables roll 151 days ahead; a real-time layer covers services in the immediate horizon; station ridership runs annually from 2015 to 2024; infrastructure snapshots reach back to 2003
- How fine
- Per-station annual counts, per-train scheduled stops, per-line infrastructure attributes
What is the SNCF Open Data Portal?
SNCF Open Data Portal is the Rail Transportation catalog's operator-layer record for France's national railway group, and it is unusually complete for one publisher. Its 166 catalogued datasets cover the whole stack of a national railway: HORAIRES SNCF publishes theoretical timetables for TGV, Intercités and TER services in both GTFS and NeTEx, rolling 151 days into the future; sncf-transilien-gtfs carries the Ile-de-France suburban network in the same formats; a real-time layer of trip updates and service alerts tracks services in the immediate horizon; frequentation-gares logs annual ridership for every one of the roughly 3,000 French stations, 2015 through 2024, straight from ticketing; and the infrastructure layers from SNCF Reseau map the physical network itself - track consistency since 2003, nominal maximum line speeds, electrified lines, safety equipment, level crossings and dozens of companion tables. Corporate series covering finance, headcount, safety incidents and CO2 emissions sit alongside.
That combination - schedules, live operations, footfall and physical plant from a single publisher under consistent identifiers - is why this record anchors the France slice of the rail catalog. Get a sample of this dataset to see the ridership panel and timetable structure mapped against your routes.
What do sample rows look like?
Two stations off the 2024 ridership panel:
nom_gare : Abbeville
code_uic_complet : 87317362 code_postal : 80100
direction_regionale: DRG HDF-NORMANDIE
segmentation : DRG class B / marketing Proximite
non_voyageurs : 0.18
total_voyageurs_2024 : 1,024,704
total_voyageurs_non_voyageurs_2024: 1,249,639
nom_gare : Ablon-sur-Seine
code_uic_complet : 87545269 code_postal : 94480
direction_regionale: DEX GIF
segmentation : marketing IDF
non_voyageurs : 0.00
total_voyageurs_2024 : 1,412,566Read the anatomy rather than the towns. Each row is one station-year: a ticketing-derived passenger total, a second total that folds in non-travelling companions and visitors, the share of those non-travellers broken out on its own, and administrative context - full UIC code, postal code, the SNCF regional stations directorate responsible, the marketing segmentation class - attached to the same row. Abbeville and Ablon-sur-Seine differ by barely 400,000 passengers, but one is classified Proximite in the Hauts-de-France directorate and the other IDF under DEX GIF: the segmentation columns carry precisely the kind of context a footfall or catchment model needs without a lookup table.
The same discipline holds across the rest of the catalogue: identifiers first, measures second, so a ridership figure joins to a timetable stop and a network layer through the UIC code alone.
What fields does the dataset include?
Eight verified fields anchor the ridership panel - the surface where every value is exemplified in the sample rows above. The timetable, real-time and infrastructure surfaces carry their own column sets, which fold under additional fields on request rather than appearing here half-invented.
What does coverage look like across geography, time and granularity?
Geography - France, at national scale. The HORAIRES SNCF timetables span TGV, Intercités and TER services across the country; the Transilien feed adds the Ile-de-France suburban network in depth; and the ridership panel covers all three thousandish stations SNCF manages, from Paris terminals down to single-platform halts like Abbeville.
Temporal - the schedule layer rolls 151 days forward at any moment, so planning horizons stay populated; the real-time layer covers services inside the immediate operating window; and the historical spine runs the other way - station ridership tabulated annually from 2015 through 2024, infrastructure snapshots maintained since 2003. Ten years of station-level footfall plus two decades of network history give trend work a base that does not need stitching together.
Granularity - three grains, cleanly separated. Ridership lands per station per year. Timetables land per train per scheduled stop. Infrastructure attributes land per line or per asset. Because each grain keys to stable identifiers - UIC codes for stations, train identifiers for services - you can aggregate up or drill down without re-keying anything.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Your cadence is decoupled from the publisher's own publication rhythm - take the full catalogue once as files, or keep a warehouse table current so new timetable windows and each year's ridership figures diff cleanly into yesterday's rows. Deliveries arrive normalised to the field dictionary above, with UIC codes preserved as the join spine: a station's ridership, its scheduled stops and its position in the network layers all resolve through the same identifier, so joins against your own route, site or territory references hold without fuzzy-matching station names.
Who uses this data, and for what?
- Demand forecasting and revenue management - ten years of ticketing-derived station totals, split between travellers and non-travelling visitors, calibrate demand models per station and per season before any survey spend.
- Network and timetable analysis - 151-day-forward schedules in GTFS and NeTEx support service-planning studies, connectivity scoring and accessibility research across TGV, Intercités and TER services alike.
- Location intelligence and retail siting - station-level footfall plus the marketing segmentation classes turn stations into scored micro-locations for catchment and site-selection work.
- Operational monitoring - the real-time trip-update and service-alert layer supports delay dashboards and disruption studies keyed to the same train identifiers as the schedule.
- Infrastructure and policy studies - electrification state, line speeds, safety equipment and level-crossing inventories pair with ridership to ground investment and policy arguments in measured network facts.
- Journalism and academic research - citable, station-attributed numbers on French mobility rather than press-release aggregates.
Which personas get the most value?
Data scientists get a decade of station-level footfall with clean categorical context and a schedule corpus in standard exchange formats - the raw material for demand, disruption and accessibility models; see the data scientists use cases page. Developers and builders get identifier-stable rows across schedules, real-time messages and network layers, ready to drop into journey-planning and mapping products; see the developers and builders use cases page. Market researchers and consultants get a measured read on French mobility behaviour by station and region, ready to cross with consumer and economic panels; see the market researchers use cases page.
How does it compare to alternatives in its slice?
Within rail transportation data, this record owns the French operator layer: one publisher covering schedules, live operations, station footfall and physical network in consistent identifiers. Its neighbours own different geographies and different layers. Deutsche Bahn's portals play the same role for Germany, the UK National Rail data portal for Britain, and the Swiss open transport platform for Switzerland - each strong at home, none covering French stations. Eurostat's railway freight statistics supply cross-country volume comparability but stop at national aggregates, while the NTAD North American rail network lines layer maps North American trackage rather than European operations. For multi-country journey-planning work the Transitland feed registry aggregates many operators' schedules including this one - shallower per-operator, broader in reach. Pair the SNCF record with Eurostat and you get French depth plus European context; pair it with its German and UK counterparts and you get the three largest Western European networks on comparable terms.
What should I know before requesting a sample?
Notes worth having in hand:
- One operator, whole-of-network depth - Schedules, real-time operations, station footfall and physical infrastructure all come from the single SNCF publisher, so identifiers reconcile across layers instead of across vendors.
- Join keys travel with the rows - Full UIC codes ride on station records, and common train identifiers connect scheduled services to real-time alerts, so joins need no name-matching.
- History where it counts - Ridership runs 2015 through 2024 annually and infrastructure snapshots reach back to 2003, giving a decade-plus of trend depth.
- Beyond the trains - Finance, headcount, safety and CO2 series make the record useful for corporate analysis, not just operations.
Where this set stops, neighbours pick up: Deutsche Bahn's portal for Germany, the UK National Rail portal for Britain, the Swiss open transport platform for Switzerland, and Eurostat for cross-country freight aggregates. Name the regions, stations or services you need, and the sample returns shaped to them with the complete field dictionary attached.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
nom_gare | string | Station name on the Frequentation en gares ridership panel. | Abbeville |
code_uic_complet | string | Full UIC station code - the join key to timetable stops and network layers. | 87317362 |
total_voyageurs_2024 | integer | Total passengers at the station for the stated year, derived from ticketing. | 1024704 |
total_voyageurs_non_voyageurs_2024 | integer | Passengers plus non-travelling visitors and companions for the stated year. | 1249639 |
non_voyageurs | number | Share of non-travellers recorded at the station. | 0.18 |
direction_regionale_gares | string | SNCF regional stations directorate responsible for the station. | DRG HDF-NORMANDIE |
segmentation_marketing | string | Marketing segmentation category such as Proximite or IDF. | Proximite |
trip_id / VehicleJourneyRef | string | Common train identifier linking real-time service alerts to scheduled services. | on request |
Questions buyers ask
How many datasets sit behind the SNCF Open Data Portal?
The catalogue reported 166 datasets at research time in August 2026. Operations dominate - timetables, real-time messages, ridership and infrastructure layers - with corporate series on finance, headcount, safety and CO2 emissions alongside, all delivered by Datadory through one normalised feed.
How far ahead do the timetables run?
HORAIRES SNCF publishes theoretical schedules 151 days into the future for TGV, Intercités and TER services, in both GTFS and NeTEx formats. The Transilien feed covers the Ile-de-France suburban network in the same formats, so planning horizons stay populated well past a typical quarter.
How far back does station ridership reach?
The frequentation-gares panel tabulates annual totals for every one of the roughly 3,000 French stations from 2015 through 2024, derived from ticketing. Each year separates travelling passengers from non-travelling visitors, which makes the series usable for both demand modelling and catchment analysis.
Does the record include real-time train information?
Yes. Trip updates and service alerts cover services inside the immediate operating window, and each message references the same train identifier as the scheduled service it concerns, so a disruption resolves against its timetable entry without guesswork. Real-time fields are documented in full with your sample.
What network infrastructure detail is available?
SNCF Reseau layers map track consistency since 2003, nominal maximum line speeds, electrified lines, safety equipment, station lists and level-crossing inventories, among dozens of companion tables. Per-line attributes such as electrification state and speed class fold under additional fields on request.
Can a sample be cut to my regions or stations?
Yes. Name the regions, stations, lines or services you care about and the sample arrives shaped to them, with the complete field dictionary attached. Samples precede any commitment, and the schema you see in the sample is the schema you ship against.
Notes on this record
- One operator, whole-of-network depth Schedules, real-time operations, station footfall and physical infrastructure come from a single publisher, so identifiers reconcile across layers instead of across vendors.
- Join keys travel with the rows Full UIC codes ride on station records and common train identifiers tie real-time alerts to scheduled services - joins need no name-matching.
- History where it counts Ridership runs annually 2015 through 2024 and infrastructure snapshots reach back to 2003, a decade-plus of trend depth in one record.
- Beyond the trains Finance, headcount, safety and CO2 series make the record useful for corporate analysis, not just operations.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.