Restaurants · Yelp
Yelp structured datasets
Datadory delivers yelp structured datasets data covering 6,990,280 reviews of 150,346 businesses across 11 U.S. metropolitan areas, plus 200,100 photos, check-ins and tips - one record family per business, review, user, check-in, tip and photo, with structured attributes such as ambience, parking, alcohol service and hours by weekday.
API, files, or your warehouse. Daily, weekly, or hourly.
- Where it covers
- United States - 11 metropolitan areas with restaurants the dominant business category throughout; the named metro list travels with your sample report rather than being guessed at
- How far back
- One coherent multi-year snapshot - review and check-in timestamps span several years of dining activity, and the exact vintage of the cut is stated in your sample report
- How fine
- Six record families - one object per business, review, user, check-in, tip and photo - joined on stable identifiers so text, reviewers and venues aggregate without remapping
What is the Yelp structured datasets?
Yelp structured datasets data is the platform's own published subset of real review data: 6,990,280 reviews, 150,346 businesses across 11 U.S. metropolitan areas and 200,100 photos, completed by user, check-in and tip record families. Restaurants are the dominant business category, and each establishment carries category aliases, hours by weekday, an aggregated star rating and structured attributes - ambience, parking availability, good-for-kids, alcohol service, outdoor seating.
Two things separate it from ordinary review scrapes. First, coherence: one cut, one schema, every review joined to both its author and its venue, so the whole thing reads as a connected graph rather than a pile of text. Second, the attribute grid beside the text - "family-friendly patios serving beer in walkable neighborhoods" is a filter pass over structured columns, not a reading exercise.
Within Datadory's catalog of 1,744 datasets averaging 7.81 on quality, this record scores 8/10 - credited as the largest review corpus in the slice, docked because a snapshot cannot answer questions that need this week's signal. Get a sample of this dataset and rows come back shaped exactly like the dictionary below, scoped to your metro areas and categories.
What does a sample of rows look like?
One establishment, one review, and the join keys that hold them together:
business_key : <stable establishment identifier>
name : <named when your sample is cut>
city : <one of the 11 metro areas> state : <two-letter code>
categories : <category alias list - restaurants dominate>
stars : 4.0 review_count : <integer>
attributes : ambience, parking, good for kids, alcohol, outdoor seating
hours : <open/close per weekday, Monday through Sunday>The anatomy matters more than any single value. Identity, geography, classification, commerce signals and the attribute grid arrive as separate typed columns in one row - which is what lets a Philadelphia cut join cleanly against a Phoenix cut without remapping anything. Every one of the 150,346 businesses repeats the same grid, and every one of the 6,990,280 reviews points back into it by identifier.
Which fields does the record dictionary cover?
The dictionary runs at the level the corpus itself guarantees: six record families plus the attribute groups every workflow touches. Three entries earn their own sentence. Attributes arrive structured - ambience, parking, good-for-kids, alcohol, outdoor seating - so venue-character segmentation is a group-by, not text mining. Check-ins give foot-traffic events that sit underneath the ratings, rare in a review-led corpus. Hours break out per weekday, which makes late-night and weekend analyses possible without inference.
Where does coverage run, and at what grain?
Geography - the United States, scoped to 11 metropolitan areas, with restaurants dominating the business mix throughout. The named metro list ships with your sample report rather than being inferred, because metro membership is the first thing every downstream aggregate depends on.
Temporal - one coherent multi-year snapshot. Review and check-in timestamps span several years of dining activity inside the cut, which suits training, benchmarking and behavioral studies; the exact vintage of the cut is stated in your sample report so nobody trains on an assumption. Workflows that need this week's signal belong with the daily-refreshed inspection records on the same industry shelf.
Granularity - six record families, one object per business, review, user, check-in, tip and photo, joined on stable identifiers. That grain is the point: sentiment by cuisine, reviewer lifetime value and venue-level visit patterns all fall out of group-bys once the families are joined, and Datadory ships them already joined.
How is the Yelp review corpus delivered through Datadory?
API, files, or your warehouse. Daily, weekly, or hourly.
Pick the channel your stack already speaks and set the cadence to match the decision being fed. For a snapshot corpus the honest engineering note is that nothing inside changes overnight - hourly plumbing buys little - but if your pipeline prefers one scheduled load into Snowflake, BigQuery or Redshift, that is a settings conversation, not a renegotiation.
Normalization happens before anything reaches you. The six record families arrive flattened and pre-joined on stable identifiers, attribute flags cast as booleans, hours typed, timestamps parsed, and category aliases resolved consistently across the whole cut. Field definitions travel unchanged through every channel, and every shipment includes validation rows, a coverage profile mapped to the metro areas and categories you named, and the vintage statement for the cut.
Who builds on the Yelp review corpus, and for what?
Five jobs it does better than anything else in the restaurants slice.
Train recommenders and language models. Nearly seven million reviews joined to venues and authors is enough to fit embedding spaces, star-prediction models and cold-start baselines without augmenting from anywhere else.
Benchmark NLP honestly. A single coherent cut beats an assembly of fragments when two teams need to compare sentiment classifiers on identical inputs.
Read dining behavior. Rating distributions, review-length patterns and tip language per metro turn "how do people talk about food here" from anecdote into measurement.
Analyze listing media. With 200,100 photos mapped to businesses, presentation standards and photo-review interplay become countable.
Teach with it. Methods classes get a real, sized, documented corpus - the difference between teaching SQL on a toy and teaching it on something worth querying.
Which personas get the most value?
Data Scientists & ML Engineers hold this dataset at relevance 3 in their restaurants pack - see data scientists use cases: volume plus joinable structure is exactly what model work needs. Journalists, Academics & Students also score it at relevance 3 - see journalists academics use cases - because a citable, single-cut corpus simplifies both teaching and replication.
The middle tier gets real but narrower leverage: market researchers frame dining-sentiment studies with it, e-commerce operators mine the review-photo relationship for listing media, builders prototype before committing to production feeds. Across the pack the pattern holds - everyone here works with words and scores about food, and nobody should mistake this corpus for a live operational feed.
How does it compare within restaurants data?
Within restaurants data, this record owns review text at scale - nothing else in the slice pairs nearly seven million reviews with structured venue attributes in one coherent cut. The neighbours own different jobs. NYC DOHMH Restaurant Inspection Results delivers regulatory citations on a daily cadence at 9/10; UK Food Hygiene Rating Scheme (FHRS) anchors the industry at 10/10 with nationwide hygiene verdicts; Overture Maps Places and Foursquare Places API supply the global place frames this corpus deliberately does not attempt.
On the catalog's rubric the split is explicit: this record scores 8/10 against the 7.81 average across 1,744 datasets. You trade freshness for coherence - one vintage, one schema, complete joins - and for training and benchmarking work that trade lands on the right side. Hugging Face Datasets - Restaurant Search is the natural complement, supplying pre-labelled subsets where this corpus supplies unlabeled volume. The head-to-head sits in Yelp structured datasets vs Hugging Face Datasets - Restaurant Search.
What should you know before requesting a sample?
Four things, stated plainly.
First, it is a snapshot. Review and check-in timestamps span multiple years inside one coherent cut, so trend claims should stay inside the span the vintage statement defines rather than reaching toward the present.
Second, column names get pinned, not presumed. The publisher documents the corpus at the level shown above, so this page asserts only what the corpus guarantees - counts, families and structured attribute groups - and every remaining column name, type and population rate is verified against real rows before your sample ships.
Third, metro membership comes explicit. The 11 metropolitan areas are named in your sample report, because every geographic aggregate downstream depends on that list being right rather than assumed.
Fourth, plan for text. Six joined families with multi-million-row review text is a real storage and join exercise; Datadory ships the families pre-joined and flattened so your first pipeline run consumes typed columns instead of doing the integration itself.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
business record | object | One row per establishment: identity, location within one of the 11 metropolitan areas, category aliases, opening hours by weekday, aggregated star rating and structured attributes. Restaurants dominate the corpus. | <one row per establishment> |
review record | object | One row per written review: the full review body joined to its business and author, carrying the star rating awarded. 6,990,280 of them. | <review body> | stars: 4 |
user record | object | One row per reviewer account, joining every review and tip back to its author so per-user behavior aggregates without remapping. | <one row per reviewer> |
check-in record | object | Timestamped visit events recording when people arrived at a venue - the foot-traffic signal sitting underneath the ratings. | <visit timestamp per venue> |
tip record | object | Short user notes left on a business, distinct from full reviews and useful for menu and service-language mining. | <short note> |
photo record | object | Photo references with captions mapped to their business - 200,100 images for listing-media and presentation analysis. | <photo mapped to business> |
attributes | flags | Structured business attributes: ambience, parking availability, good-for-kids, alcohol service and outdoor seating. | outdoor seating, good for kids |
categories | alias array | Category aliases placing each establishment in the taxonomy; restaurant placements dominate the corpus. | <restaurant category alias> |
hours | object | Opening hours broken out per weekday rather than collapsed into one string. | <Monday-Sunday open/close> |
stars | number | The review score carried per review and aggregated onto the business record. | 4 |
Additional fields on request | - | Exact column names, types and per-field population rates for all six record families, enumerated and defined with examples pinned against real rows for your named metro areas before the sample ships. | <on request> |
Questions buyers ask
What is included in the Yelp structured datasets?
6,990,280 reviews, 150,346 businesses across 11 U.S. metropolitan areas, 200,100 photos, plus user, check-in and tip record families. Restaurants are the dominant business category, and establishments carry category aliases, hours by weekday, star ratings and structured attributes such as ambience, parking, good-for-kids, alcohol service and outdoor seating.
Which fields does each record include?
At the guaranteed level: six record families - business, review, user, check-in, tip and photo - with business rows carrying location, category aliases, weekday hours, star rating and the structured attribute group, and review rows carrying the full text joined to author and venue. Exact column names, types and population rates are pinned against real rows in your sample report.
How current is the Yelp review corpus?
It is a coherent multi-year snapshot rather than a live feed: review and check-in timestamps span several years inside one cut, and the exact vintage is stated in your sample report. For monitoring-type questions that need current signal, pair it with the daily-refreshed inspection records in the same industry shelf.
Which restaurant dataset should I train a sentiment model on?
Use the Yelp review corpus for unlabeled volume - 6,990,280 reviews joined to venues and authors - and take pre-labelled subsets from Hugging Face Datasets - Restaurant Search, led by yelp_restaurant_review_labelled in the 1M-10M row band. Labels on top of a shared vocabulary beat annotating from scratch.
Does the corpus include foot-traffic signals?
Yes - check-ins arrive as their own record family, one timestamped visit event per row joined to its venue. That gives a behavioral layer underneath the ratings, useful for popularity and visit-pattern studies, alongside tips as short-form text distinct from full reviews.
Can a sample be scoped to my metro areas and categories?
Name the metro areas and categories and the sample returns real rows shaped exactly like the dictionary above, with the named metro list, per-field population rates and the vintage statement of the cut reported up front so you know precisely what you are evaluating before committing to a feed.
Datasets that pair with this one
- Hugging Face Datasets - Restaurant Search The labelled counterpart: 975 community corpora including yelp_restaurant_review_labelled - labels on top where this corpus supplies volume. Scores 7/10.
- NYC DOHMH Restaurant Inspection Results About 295,000 violation rows for all five boroughs - regulatory ground truth where a review snapshot offers model fuel. Scores 9/10.
- Foursquare Places API 100M+ places with chains, hours, price tiers and popularity - the place frame this text-led corpus deliberately does not attempt. Scores 8/10.
- best restaurants datasets The pooled industry ranking, scored across all eight restaurants records we track - this corpus sits fourth.
- restaurants data hub The full industry view, from hygiene ratings to review corpora and place frames.
- review corpus What a packaged review collection actually promises, field by field.
- check-in data Foot-traffic events as a data product, and how they sit beneath ratings.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.