Datadory notebook

Hotel booking cancellation prediction datasets: what to train on

The hotel booking cancellation prediction dataset most teams start with is Hotel Booking Demand (Kaggle - 119k Portuguese Hotel Bookings): 119,390 labeled bookings from one resort and one city hotel in Portugal, 32 fields including lead_time, ADR and the is_canceled flag, downloaded as a roughly 17 MB CSV under commercial delivery terms 4.0.

1,744 datasets. Pick your catch.

Which hotel booking cancellation prediction datasets can you actually train on?

Labeled, row-level booking data is scarce in this vertical. Of the 23 datasets Datadory catalogs for Hotel & Resort REITs (16 primary plus 7 related), twelve of the primaries are free, but nearly all of them are financial or macro: SEC filings, Nareit sector aggregates, Yahoo Finance price bars, Census and BEA demand series. Only one record in the slice carries per-booking rows - Hugging Face Datasets - Hotel Search (quality 6), a live query over the Hub returning roughly 20-40 hotel-related repositories at any time, among them booking-demand corpora of about 119,000 rows with ADR and cancellation flags covering two European hotels, arrivals 2015 through 2017.

The canonical upload of that same corpus sits one industry over, in the Hotels, Resorts & Cruise Lines slice: Hotel Booking Demand (Kaggle - 119k Portuguese Hotel Bookings), which Datadory scores 9 of 10. A third corpus completes the kit: 515K Hotel Reviews in Europe (Kaggle) (quality 8, commercial delivery terms), whose 515,738 reviews carry no cancellation label but supply the sentiment and review-history features most published models bolt on. Expect to cross the industry boundary once to get the labels.

What does the 119,390-row Portuguese bookings file contain?

One row per booking, 32 columns per row, about 17 MB of CSV. The two properties are coded in the hotel field - H1 is the Resort Hotel, H2 the City Hotel, both in Portugal - with arrivals spanning July 2015 through August 2017 at daily resolution. The label lives in is_canceled, backed by the outcome pair reservation_status (Canceled, Check-Out or No-Show) and reservation_status_date.

The remaining columns describe why a booking might cancel. Timing: lead_time, four arrival-date components, stays_in_weekend_nights, stays_in_week_nights, days_in_waiting_list. Party and product: adults, children, babies, the meal package (Undefined/SC, BB, HB, FB), reserved_room_type versus assigned_room_type, required_car_parking_spaces. Commercial history: market_segment, distribution_channel (TA for travel agents, TO for tour operators), agent and company IDs, is_repeated_guest, previous_cancellations, previous_bookings_not_canceled, booking_changes, customer_type (Group, Transient, Transient-Party, Contract), deposit_type (No Deposit, Non Refund, Refundable), adr, total_of_special_requests, and country of origin in ISO 3155-3:2013 form. Verified sample rows show how wide behavior runs: one Resort Hotel booking made 342 days ahead of a July 2015 week-27 arrival and not canceled, another with lead_time of 737 days.

How do you train a cancellation classifier without leaking the label?

Six steps take the raw CSV to a defensible model:

Because nothing moves underneath you, the file works as a benchmark dataset: two teams downloading today get byte-identical inputs.

Where do demand features come from when the labels stop in 2017?

The corpus freezes at August 2017, so anything you deploy needs exogenous signal, and the free government layer provides it. The Census Bureau Economic Indicators Briefing Room publishes the Quarterly Services Survey's quarterly traveler-accommodation revenue (NAICS 721) from roughly 2004 to present - the highest-frequency official read on lodging receipts between REIT earnings cycles - next to MRTS retail and food-services series reaching back to 1992. The BEA Travel and Tourism Satellite Account contributes 27 annual Excel workbooks spanning 1998 through 2023 with explicit traveler-accommodations output, employment and price tables; regular production ended in February 2026, so treat the archive as a fixed deflator.

Will a model trained on two Portuguese hotels transfer to a lodging REIT portfolio?

Treat the corpus as a methods asset, not a portfolio forecast. Three limits decide transfer. Scope: two properties, one country - a classifier tuned on City Hotel dynamics may not generalize to a 14-name REIT cohort such as HST, RHP, PK or SHO. Vintage: arrivals end in August 2017 and Kaggle last touched the file on 2020-02-13, a textbook case of dataset snapshot staleness; 181 of the 1,744 datasets Datadory catalogs share the static profile. Format: the corpus ships CSV, and Parquet remains rare catalog-wide at 4.9% of records - the Hugging Face mirrors are the practical route to columnar loading.

The bridge to REIT work runs through fundamentals. The SEC XBRL Company Facts API returns every tagged concept per filer in one keyless JSON call - Host Hotels & Resorts (CIK 1070750) yields about 3.0 MB across 459 us-gaap concepts, at 10 requests per second with a declared User-Agent - while an active lodging REIT files roughly 20-60 documents per year on EDGAR carrying comparable-hotel RevPAR and occupancy tables. Joining modeled cancellation pressure to those disclosed operating statistics is the honest way to connect a classroom classifier to investable lodging economics.

Where to go next

Start with the Hotel & Resort REITs data guide, the pillar that maps all 23 cataloged datasets in the industry and places the booking corpora beside the filings stack, Nareit aggregates and government demand series. Two sibling clusters go deeper on adjacent layers: the SEC EDGAR hotel REIT filings guide covers the disclosure side a deployed model gets benchmarked against, and the traveler accommodation revenue census post explains the Census QSS and BEA series behind the macro features above.

For the data itself, the dataset pages document every field and access path: Hotel Booking Demand (Kaggle - 119k Portuguese Hotel Bookings), Hugging Face Datasets - Hotel Search and 515K Hotel Reviews in Europe (Kaggle). To see how the row-level corpus differs from official arrival counts, read Kaggle Hotel Bookings vs NTTO I-94 Arrivals, browse everything in the slice from the Hotel & Resort REITs data hub, or follow the ML-specific shortlist on the data scientists page.

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Hotels, Resorts & Cruise Lines Two properties in Portugal - one resort hotel (H1) and one…

Hotel Booking Demand (Kaggle – 119k Portuguese Hotel Bookings)

hotel · lead_time · adults …+14 more

Hotel & Resort REITs Booking-demand corpus covers two European hotels (resort and…

Hugging Face Datasets - Hotel Search

Hotel & Resort REITs Six European countries - United Kingdom

515K Hotel Reviews in Europe

Hotel & Resort REITs United States national totals, with some state-level detail…

Census Bureau Economic Indicators Briefing Room

Hotel & Resort REITs United States, national totals only - no state, metro or…

BEA Travel and Tourism Satellite Account

Commodity

Hotel & Resort REITs Global connected-property inventory keyed by IATA city code…

Amadeus Hotel Search & Booking APIs

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

Can I use hotel booking data commercially?

Yes for the core corpora. Kaggle's metadata lists Hotel Booking Demand as Attribution 4.0 International (commercial delivery terms 4.0), so commercial reuse is permitted with credit to Antonio, Almeida and Nunes. The 515,738-review European corpus is commercial delivery terms commercial delivery terms, no attribution required. On Hugging Face, check every card: argilla/tripadvisor-hotel-reviews is commercial delivery terms-NC 4.0, which bars commercial use.

Is there a US hotel cancellation dataset?

No row-level US equivalent appears in Datadory's catalog of 1,744 datasets. US coverage stays aggregate: the Quarterly Services Survey publishes quarterly traveler-accommodation revenue under NAICS 721 from roughly 2004 onward, and AHLA's four impact CSVs cover rooms, jobs and guest spending for 52 state rows and 436 congressional districts - neither records individual bookings or cancellations.