Airline Passenger Satisfaction (103K Survey Responses)

Datadory delivers airline passenger satisfaction 103k survey responses data covering 129,880 surveyed passengers, each rating 14 onboard and digital service attributes from inflight wifi to cleanliness on a 0-5 scale, joined to traveler profile, cabin class, flight distance, delay minutes and a satisfied versus neutral-or-dissatisfied label - the largest labeled customer-experience panel in our passenger airlines slice. API, files, or your warehouse; daily, weekly, or hourly.

What is Airline Passenger Satisfaction (103K Survey Responses)?

A census-scale opinion poll of flying, packaged for modeling. Airline Passenger Satisfaction (103K Survey Responses) is a labeled customer-experience panel: 129,880 surveyed passengers, each described by who they are (loyal or disloyal customer, business or personal traveler, Business, Eco or Eco Plus cabin), how far they flew, how late their flight ran, and how they scored fourteen separate touchpoints of the journey - the wifi, the booking flow, the gate, the food, the seat, the crew, the baggage belt, the cleanliness of the cabin. Every row closes with the binary verdict the whole file exists for: satisfied or neutral or dissatisfied.

That verdict column is the asset. Most customer-experience sources stop at averages; this record keeps the individual judgment attached to its causes, 129,880 times over, which is why it has become a standard classification benchmark - more than 138,000 downloads and 511 public notebooks trace back to it. It ships as a ready-made modeling split: 103,904 training rows and 25,976 test rows.

For the Passenger Airlines slice this is the only labeled experience microdata among its primary datasets - the neighbours count flights, seats and traffic, not opinions. In Datadory's catalog of 1,744 datasets across 159 viable industries the record scores 8/10 for quality with a verified field dictionary behind it. [Get a sample of this dataset](#request) and we cut it to the fields and segments you name.

What do sample rows look like?

Three of the shipped sample rows, exactly as the columns present themselves:

id 19556 | Female | Loyal Customer | age 52 | Business travel | Eco    | distance   160 mi
wifi 5 | schedule 4 | booking 3 | gate 4 | food 3 | online_boarding 4 | seat 3
entertainment 5 | on-board 5 | leg_room 5 | baggage 5 | checkin 2 | inflight_service 5 | cleanliness 5
dep_delay 50 min | arr_delay 44 min -> LABEL: satisfied

id 90035 | Female | Loyal Customer | age 36 | Business travel | Business | distance 2,863 mi
wifi 1 | schedule 1 | booking 3 | gate 1 | food 5 | online_boarding 4 | seat 5
entertainment 4 | on-board 4 | leg_room 4 | baggage 4 | checkin 3 | inflight_service 4 | cleanliness 5
dep_delay 0 min | arr_delay 0 min -> LABEL: satisfied

id 12360 | Male | disloyal Customer | age 20 | Business travel | Eco    | distance   192 mi
wifi 2 | schedule 0 | booking 2 | gate 4 | food 2 | online_boarding 2 | seat 2
entertainment 2 | on-board 4 | leg_room 1 | baggage 3 | checkin 2 | inflight_service 2 | cleanliness 2
dep_delay 0 min | arr_delay 0 min -> LABEL: neutral or dissatisfied

The three rows already stage the tension the full 129,880 resolve. Respondent 90035 rated the hard product lavishly - seat 5, food 5, entertainment 4 on a 2,863-mile business run - yet scored wifi 1 and schedule convenience 1, and still landed on the satisfied side of the line. Respondent 12360 scored almost nothing above 2 anywhere and dissatisfied accordingly. One row proves satisfaction survives a bad airport hour; the other shows the label tracking the average of fourteen judgments rather than any single grievance. Request a sample and the same columns come back filtered to the segment you name.

What fields does the dataset include?

Twenty-three fields define every response, all marked verified against the published schema:

  • id - respondent identifier within the extract.
  • Gender - Male or Female.
  • Customer Type - loyalty status: Loyal Customer or disloyal Customer.
  • Age - respondent age in years.
  • Type of Travel - Business travel or Personal Travel.
  • Class - cabin flown: Business, Eco, or Eco Plus.
  • Flight Distance - miles flown on the trip being rated.
  • Fourteen service ratings on an integer 0-5 scale, where 0 means not applicable: Inflight wifi service, Departure/Arrival time convenient, Ease of Online booking, Gate location, Food and drink, Online boarding, Seat comfort, Inflight entertainment, On-board service, Leg room service, Baggage handling, Checkin service, Inflight service, Cleanliness.
  • Departure Delay in Minutes and Arrival Delay in Minutes - operational context for the trip; arrival delay carries some missing values.
  • satisfaction - the target label: 'satisfied' or 'neutral or dissatisfied'.

What is not there matters as much: no airline code, route, airport or flight-date column exists, which is precisely why the record works as a general experience model and cannot be used to rank a named carrier. Any carrier-level enrichment sits under additional fields on request and is confirmed against your use case before anything ships.

What geography, time range, and granularity does it cover?

  • Geography: not stated by the publisher - and unrecoverable from inside the files, because airline, route and airport identifiers were never included. Treat the panel as geography-blind by construction; that is the trade for a 129,880-row labeled sample.
  • Temporal: a single static compilation. No per-response date is recorded, so the survey window cannot be reconstructed from the data; the record has stayed at one version since publication.
  • Granularity: one row per surveyed passenger per flight experience, with all twenty-three fields populated on that row.

The consequence cuts both ways. Nothing here decays into a stale time series because there is no series to decay - the file answers "what makes a passenger satisfied" identically next year. But any question needing a when or a where needs a companion: the T-100 traffic tables or the World Bank's ICAO-sourced passengers-carried series carry the operational and geographic axes this record deliberately omits.

How is this dataset delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

You pick the channel and the cadence; the field dictionary above travels unchanged across all three. Teams benchmarking classifiers tend toward the full file load once and versioned thereafter; CX analytics groups usually scope to the attribute grid plus the segments they track; product teams embedding a satisfaction model take a scoped feed. Name the fields and filters when you [get a sample](#request) - the sample ships first either way, and changing cadence afterward is a settings conversation, not a re-integration project.

Who uses this data, and for what?

A labeled verdict on 129,880 journeys earns its keep in five specific jobs:

  • Satisfaction-driver ranking - order the fourteen attributes by how strongly they move the satisfied label, so a product roadmap fixes the wifi before it repaints the lounge. The 0-5 grid makes the ranking computable rather than anecdotal.
  • Classification benchmarks - train and test splits arrive pre-cut at 103,904 and 25,976 rows, so model comparisons stay like-for-like across teams and papers.
  • Segment profiling - slice by Customer Type, Type of Travel and Class to watch how the same cabin scores differently for a loyal business flyer versus a disloyal personal one.
  • Delay-impact analysis - join the two delay-minute columns against the label to price operational failure in satisfaction terms rather than only in compensation claims.
  • Teaching and methods work - a clean, fully labeled, mid-sized tabular problem with a genuine class balance question, citable without footnote gymnastics.

Worked examples continue on the passenger airlines pages for data scientists, market researchers and competitive intelligence and product teams.

Which personas get the most value?

Data scientists and ML engineers get the canonical use: a large, verified, fully labeled tabular panel with the train/test split already drawn, so the pipeline starts at feature engineering instead of data assembly. Market researchers and consultants get driver analysis without a fieldwork budget - 129,880 completed interviews nobody had to recruit. Competitive intelligence and product teams get the attribute-level yardstick for experience investment cases, with the caveat that no carrier identifier exists to point it at a named rival. Journalists, academics and students get a teaching set whose provenance is documented and whose limitations are stated plainly on this page. Sales and growth teams selling CX tooling into airlines get the reference structure of a satisfaction model without licensing a vendor's black box. All delivered daily, weekly, or hourly.

Provenance note - the record circulates under the Kaggle name, uploaded in February 2020, and several near-identical copies exist on the platform under different account names. Our verification pass flagged that upstream collector and sampling frame as undocumented; cite it as a Kaggle-hosted survey unless you identify the original instrument first. The source profile covers the platform's wider catalog at Kaggle.

Quality note - Datadory scores this record 8/10 against a catalog average of 7.81 across all 1,744 datasets. Field definitions are confidence-rated verified - a bar met by 85.7% of the catalog. What holds it at 8: undocumented geography and survey period, one column with missing values, and the unresolved upstream attribution above.

Label note - the target takes exactly two values, so treat it as a classification problem, not a regression on happiness. 'Neutral or dissatisfied' is one bucket; if your use case needs the middle separated, that recoding happens on request with the schema confirmed first.

Where to go next - for the operational axis this record omits, 2015 Flight Delays and Cancellations (5.8M U.S. flights) supplies per-flight identifiers and on-time outcomes; for network-level reference data, OpenFlights airline, airport and route database maps the carriers and airports; for the macro view, the World Bank air transport passengers carried comparison sets this microdata against aggregates. The full shortlist lives at best passenger airlines datasets.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary — Airline Passenger Satisfaction (103K Survey Responses)
FieldTypeDefinitionExample
idintegerRespondent identifier assigned within the extract.19556
GenderenumRespondent gender.Female
Customer TypeenumLoyalty status of the passenger.Loyal Customer / disloyal Customer
AgeintegerRespondent age in years.52
Type of TravelenumPurpose of the trip being rated.Business travel / Personal Travel
ClassenumCabin class flown.Business / Eco / Eco Plus
Flight DistanceintegerDistance of the flight in miles.160
Inflight wifi serviceintegerRating for inflight wifi, 0 (not applicable) to 5.5
Departure/Arrival time convenientintegerRating for schedule convenience, 0-5.4
Ease of Online bookingintegerRating for the online booking experience, 0-5.3
Gate locationintegerRating for gate location convenience, 0-5.4
Food and drinkintegerRating for food and beverage, 0-5.3
Online boardingintegerRating for the online boarding process, 0-5.4
Seat comfortintegerRating for seat comfort, 0-5.3
Inflight entertainmentintegerRating for inflight entertainment, 0-5.5
On-board serviceintegerRating for cabin crew service, 0-5.5
Leg room serviceintegerRating for leg room, 0-5.5
Baggage handlingintegerRating for baggage handling, 0-5.5
Checkin serviceintegerRating for check-in service, 0-5.2
Inflight serviceintegerRating for inflight service, 0-5.5
CleanlinessintegerRating for cabin cleanliness, 0-5.5
Departure Delay in MinutesintegerDeparture delay experienced on the trip.50
Arrival Delay in MinutesnumberArrival delay experienced on the trip; contains some missing values.44.0
satisfactionenumTarget label for the response.satisfied / neutral or dissatisfied
Additional fields-Folded under 'additional fields on request': any carrier-, route- or period-level enrichment layered onto these rows. Confirmed with you before delivery - none ship by default.on request

Coverage — geography, temporal range, granularity

DimensionCoverage
GeographyNot stated by the publisher; no airline, route, airport or geographic identifier exists in the files, so geography cannot be recovered
TemporalSingle static compilation of survey responses; no per-response date recorded, unchanged since publication
GranularityOne row per surveyed passenger per flight experience, 23 fields per row

Record shape — volume and split

MeasureValue
Total labeled responses129,880
Training rows103,904
Test rows25,976
Columns per row25 (index + id + 23 documented fields)
Uncompressed size~15 MB across the two CSV files
Datadory quality score8 / 10 (catalog average 7.81)

Questions buyers ask

What does Airline Passenger Satisfaction (103K Survey Responses) include?

129,880 passenger-level survey records across 23 fields: six profile and itinerary columns (gender, loyalty status, age, travel purpose, cabin class, flight distance), fourteen 0-5 service ratings from inflight wifi through cleanliness, two delay measures in minutes, and a binary satisfied versus neutral-or-dissatisfied label.

How big is the airline passenger satisfaction dataset?

129,880 labeled responses in total, delivered as a pre-built modeling split: 103,904 training rows plus 25,976 test rows, about 15 MB uncompressed. The split arrives intact, so benchmark runs are reproducible out of the box instead of re-cut after delivery.

Which airlines or countries does the satisfaction survey cover?

Neither can be recovered. Geographic coverage is not stated by the publisher, and the files contain no airline code, route, airport or flight-date column, so no response can be tied to a carrier or country. Each row records only the passenger profile, cabin class, flight distance and delay minutes.

Does the dataset have missing values?

One column only. Arrival Delay in Minutes holds some missing values; every one of the fourteen service ratings arrives as a complete integer from 0 (not applicable) to 5, and the satisfaction label is fully populated. Plan imputation or row filtering before scoring classifiers on the 103,904-row training side.

What predicts satisfaction in this survey?

The structure points at the digital and physical touchpoints: business-class loyal travelers flying longer distances dominate the satisfied side, while the sample's low scorers cluster on inflight wifi, schedule convenience and check-in service. The 0-5 attribute grid is what makes driver ranking possible rather than a single overall score.

Can a Datadory sample be scoped before I commit?

Yes. Name the fields you need from the dictionary above - the attribute grid, the profile columns, the delay pair, the label - and any filter such as business-travel loyal customers only, and the extract returns cut to that shape with the schema intact, delivered by API, files, or your warehouse.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing