Datadory notebook
How Competitive Intel Product Teams Use Passenger Airlines Data
1,744 datasets. Pick your catch. Every guide here is built on what the catalog can actually prove.
1,744 datasets. Pick your catch.
What can you actually learn about an airline competitor without buying a feed?
Datadory catalogs 14 datasets for Passenger Airlines - 6 primary records plus 8 related ones - and the economics favour a watching brief: 5 of the 6 primary datasets are free, and only Amadeus charges, and only beyond its test quota. Against that, 82.4% of the full 1,744-dataset catalog is free, so this slice is not even an outlier; it is simply a vertical where mandatory government filing does most of the collection work for you.
The honest caveat up front, because it shapes every workflow below: nothing here archives rivals' sold fares, and no record returns a fare ladder. What you get is capacity, schedule reliability, service attribute benchmarks, network structure and demand context - the inputs from which competitive moves are inferred rather than announced.
Which dataset reveals a rival carrier's monthly capacity and share shifts?
BTS TranStats - Air Carrier Statistics (T-100) & Airline On-Time Performance is the backbone. It indexes 21 aviation databases, ships as public-domain bulk CSV with no registration step, and accumulates roughly 5-7 million on-time flight rows per year back to October 1987. The T-100 segment table logs one row per carrier, aircraft type, service class, origin-destination segment and month, capturing passengers, seats, available seat-miles and load factors from 1990; DB1B ticket samples run from 1993.
Because filing is mandatory and field-uniform, competitor share becomes arithmetic instead of estimation. Divide one carrier's seats or ASKs on an airport pair by all carriers' totals and you own a monthly share series nobody's investor-relations team wrote for you. The service-class code keeps the read clean: first class, coach and charter are separable, so a charter surge never masquerades as scheduled-capacity growth.
Two limits belong in your runbook. Monthly filing lags mean you see a network move weeks after frequencies changed, not when it was decided. And the T-100 All Carriers variant is what adds foreign carriers serving the United States - domestic tables alone will understate competition on international segments.
Can you catch network and seatmap changes faster than the filings cycle?
For competitor work the tells are structural: a carrier's routes table thinning out of specific city pairs, seatmaps showing densified cabins on a route where legroom was a selling point, status queries surfacing retimed frequencies. Treat sustained availability changes as network-strategy signals rather than one-off inventory noise.
Where do you benchmark customer experience against the competition?
Airline Passenger Satisfaction (103K Survey Responses) is the slice's only labelled customer-experience microdata: 129,880 survey responses rating 23 service attributes - inflight wifi, seat comfort, online boarding and cleanliness among them - against a satisfied versus neutral/dissatisfied label, pre-split into 103,904 training and 25,976 test rows.
For a product team the competitive value is calibration. Ranking attribute importance gives you a defensible answer to which experience levers separate delighted flyers from indifferent ones - the same weighting question behind your own NPS programme - and it lets you sanity-check whether a rival's wifi-first marketing targets a variable passengers actually weight.
How reliable is a rival's operation, and what baseline should you compare against?
Reliability comparisons need history, and history is where the free layer shines. For calendar year 2015 exactly, the 2015 Flight Delays and Cancellations (5.8M U.S. Flights) mirror packages 5,819,079 per-flight rows - scheduled versus actual times, delay minutes, cancellation and diversion flags - under commercial delivery terms with a 14-row airline file and a 322-row airports.csv carrying coordinates. commercial delivery terms means a derived reliability score can ship inside a commercial product without clearing legal.
Frozen at December 2015, though, it is a baseline, not a monitoring feed; our tagging scores it relevance 1 for competitor tracking for precisely that reason. Anything current comes from the BTS on-time tables, which keep accruing month after month back to October 1987 - long enough to establish whether a rival's publicised 'operational turnaround' is trend or seasonality.
How do you map a competitor's network structure and alliance footprint?
Network reach is documentary. OpenFlights Airline, Airport and Route Database ships three plain CSVs - 7,698 airports, 6,162 airlines and 67,663 directional routes with IATA/ICAO codes, coordinates and codeshare flags - under the commercial delivery terms. Codeshare flags are the competitive detail: they let you see where two marketing names ride one physical flight, which is how alliance depth reads in data rather than press releases.
Age it correctly. Airport and airline files still receive periodic refreshes, but the route table comes from a third-party feed frozen at June 2014 - treat it as structural reference for geography and equipment types, never as current schedule evidence.
Which competitor-signal sources rank strongest for this workflow?
Ranked by how much competitor behaviour each record exposes per dollar of effort:
- BTS TranStats (T-100 & On-Time Performance) - mandatory monthly filing converts to computable capacity and load-factor share from 1990; commercial delivery terms.
- World Bank IS.AIR.PSGR - annual demand context for about 265 economies since 1970 under commercial delivery terms-4.0.
- OpenFlights database - 67,663 directional routes plus codeshare flags, frozen June 2014; commercial delivery terms.
- 2015 Flight Delays (5.8M) - commercial delivery terms reliability baseline for calendar 2015 only; static.
What does a six-step competitor-tracking loop look like?
Sequenced by how soon each step pays off, using only sources named above:
- Backfill shares. Load T-100 segment extracts from 1990 onward and compute monthly seats and ASK share per carrier on your watched airport pairs; split by service class so charter stays out of scheduled comparisons.
- Baseline reliability. Compute on-time and cancellation rates from the commercial delivery terms 2015 corpus (5,819,079 rows) joined to airports.csv coordinates, then extend with BTS monthly extracts to confirm or retire the pattern.
- Benchmark experience. Rank attribute importance on the satisfaction survey's built-in 103,904/25,976 split and reuse the attribute list in your own voice-of-customer instrument.
- Frame the market. Overlay OpenFlights' 67,663 routes on World Bank IS.AIR.PSGR growth rates to separate a rival's bold entry into a growing market from one propping up a shrinking one.
- Re-run monthly. Diff the newest BTS vintage against your stored baseline; alert on three consecutive months of share movement rather than any single print.
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
BTS TranStats - Air Carrier Statistics (T-100) & Airline On-Time Performance
CRSDepTime · DepTime · CRSArrTime …+9 more
Amadeus for Developers - Airline Code, Flight Offers & Routes APIs
type · iataCode · icaoCode …+12 more
Airline Passenger Satisfaction (103K Survey Responses)
2015 Flight Delays and Cancellations (5.8M U.S. Flights)
AIRLINE · FLIGHT_NUMBER · TAIL_NUMBER …+4 more
EUROCONTROL Aviation Data & Dashboard
UK CAA Aviation Data & Analysis
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
Is there a free dataset for tracking competitor airline capacity?
Yes - BTS TranStats publishes census-level T-100 segment traffic (passengers, seats, ASKs and load factors by carrier, aircraft type, service class and airport pair) monthly from 1990 as public-domain bulk CSV with no registration, and monthly on-time records per flight back to October 1987.
Can I compute competitor market share from public airline data?
On US segments, yes: dividing one carrier's T-100 seats or available seat-miles on an airport pair by all carriers' totals yields a monthly share series from mandatory uniform filings. Use the service-class field so charter traffic never contaminates a scheduled-capacity comparison.