Consumer Electronics · RTINGS Ltd.

RTINGS.com Lab-Tested Reviews Data

Datadory delivers rtings com lab tested reviews data covering the independent bench measurements behind RTINGS.com: 4,847 products bought and tested across 28 categories in a 40,000 square foot Montreal facility, each review carrying dozens to hundreds of standardized tests under a versioned protocol - 0-10 scores per usage scenario for TVs (Mixed Usage, Home Theater, Bright Room, Sports, Gaming), measured values such as a 165Hz native refresh rate, editorial verdicts and the worth engine's cheaper-alternative and upgrade picks.

API, files, or your warehouse. Daily, weekly, or hourly.

Where it covers
US-market products purchased at retail and tested at the source's Montreal facility; pricing references US retail
How far back
Reviews published since 2011, with units physically re-tested when firmware updates change measured performance; older reviews retain their historical test-bench versions
How fine
One review per product carrying dozens to hundreds of individual tests - roughly 4,847 products bought and tested across 28 categories, several hundred full reviews in the TV branch alone

What is RTINGS.com Lab-Tested Reviews?

RTINGS.com Lab-Tested Reviews is the Consumer Electronics catalog's measurement layer: 4,847 products bought at retail and tested by RTINGS Ltd. across 28 categories in a 40,000 square foot Montreal facility, running since 2011. The distinction from everything else on the shelf is who holds the meter. GSMArena catalogs what vendors claim, Icecat distributes manufacturer data-sheets, ENERGY STAR certifies power draw - RTINGS buys the unit, puts it on the bench and publishes what the instruments said.

Every review follows a documented, repeatable protocol under a versioned test bench (2.2 is current for televisions), which is what makes two reviews comparable years apart. A television is scored against five usage scenarios - Mixed Usage, Home Theater, Bright Room, Sports and Gaming - plus performance sub-scores such as Brightness in SDR and HDR. The category roster reaches well past displays: soundbars, projectors, headphones, speakers, mice, keyboards, printers, routers, laptops, vacuums, air conditioners, air purifiers, dehumidifiers, humidifiers, mattresses, cameras and running shoes all pass through the same discipline of buy-at-retail, bench-test, publish-numbers.

Get a sample of this dataset - name the categories or brands you track, and real bench rows come back before you commit to a feed.

What do sample rows look like?

One television, one test entry, read straight off the review structure:

product_name      : Samsung S95H OLED              category : tv
test_bench        : 2.2                            status   : tested
test_name         : Native Refresh Rate            score    : 9.5
rendered_value    : 165Hz                          unblurred: true
usage_name        : Mixed Usage                    insider_only: false
linked_description: The TV has excellent motion handling thanks to its 165Hz panel.
worth_comparison  : TCL S325                       upgrade_comparison: LG G5 OLED

The anatomy is the point. One row is one test observation inside one review: the unit identified by name, the methodology pinned by bench version, the instrument reading both as a number (score 9.5) and as the human-facing rendering (165Hz), and the usage context attached alongside. Then come the two columns no spec sheet has - worth_comparison, naming the cheaper alternative the source's engine would steer a buyer to, and upgrade_comparison, naming the tier above. A row therefore answers three questions at once: what was measured, how good it was, and what else your money should consider.

What fields does the dataset include?

Thirteen fields carry each test observation, definitions verified against live review pages rather than inferred from documentation. The spine is structural - product_name, category, test_bench, test_name, status - so any row identifies its unit, its methodology vintage and its place in the protocol without ambiguity. The measurement core is score (0-10), rendered_value and usage_name; the flags unblurred and insider_only record the visibility state of each value upstream; and linked_description carries the editor's verdict paragraph for the usage scenario.

The pair that earns the table its keep are the worth-engine columns. worth_comparison and upgrade_comparison encode the source's own recommendation graph - which cheaper model it steers you to and which step-up it names - so brand-vs-brand positioning questions become a join instead of a reading assignment.

What does coverage look like across geography, time and granularity?

  • Geography - US-market products bought at retail and benched in Montreal, with pricing references tied to US retail. Scores describe the specific units purchased, which matters where regional variants differ in panel or tuner.
  • Temporal - reviews since 2011, and here is the twist most review sources lack: units get physically re-tested when firmware updates change measured performance, so a score can move after collection. Older reviews keep their historical test-bench versions instead of being silently restated, letting a bench 2.2 panel compare honestly against an older measurement.
  • Granularity - one review per product carrying dozens to hundreds of individual tests; roughly 4,847 products bought and tested across 28 categories, several hundred full TV reviews among them. Each review is a dense payload because the entire test suite travels with it.

The honest caveat: population follows the source's editorial queue. The 4,847 units are the ones RTINGS chose to buy and measure - absence means untested, not disapproved, and niche SKUs outside the queue will not appear. Teams needing full-assortment coverage join this set against a catalog source rather than expecting it from here.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

You pick the channel and the cadence; the field dictionary above travels unchanged through all three. Review pages never reach your pipeline - rows arrive flattened into one observation per test with the vocabulary already normalized, so a score column compares against a score column and a firmware re-test shows up as a dated movement rather than a surprise overwrite. Because units are genuinely re-measured when software changes performance, cadence is not cosmetic here: hourly delivery catches a re-test the week it lands. Cadence changes are a settings conversation, not a re-integration project, and a sample cut to your categories and brands comes first either way.

Who uses this data?

  • Competitive-intel and product teams benchmark their own lineups against rivals' bench numbers - display brightness, refresh behavior, input latency class - including the re-tests that move a rival's standing after a firmware push.
  • Market researchers and consultants frame category studies around standardized per-usage scores across 28 categories instead of stitching together incompatible star-rating piles.
  • Data scientists and ML engineers use the 0-10 scores and rendered measurements as labels for quality-prediction models, with test_bench marking methodology vintages so training sets do not silently mix protocols.
  • Retail buyers and merchandising teams let the five TV usage scores arbitrate assortment - a bright-room score for floor placements, a gaming score for the enthusiast end - backed by instruments rather than margins.
  • Journalists, academics and students cite a consistent methodology run continuously since 2011, with the bench version recorded on every observation.
  • Developers and data-product builders power comparison tools off the worth-engine columns, inheriting the source's cheaper-alternative and upgrade recommendations without maintaining their own ranking heuristics.

Which personas get the most value?

Competitive-intel and product teams hold this dataset at the top of their consumer-electronics pack - nothing else in the industry slice measures rival hardware with instruments, and the firmware-triggered re-tests make it the only feed where a competitor's picture quality can legitimately change overnight. Market researchers rank it nearly as high: standardized per-usage scoring is what turns 'reviewers liked it' into a comparable series. Data scientists get labeled measurement data rare in consumer electronics, where most sources publish narrative text. Retail merchandisers get scenario-resolved scores - bright room versus home theater versus gaming - that map onto how sets are actually sold. Across all of them the constant is standardization: one protocol vocabulary holding steady across 28 categories and fifteen years.

How does RTINGS.com Lab-Tested Reviews compare with other consumer electronics datasets?

Within Datadory's consumer-electronics shelf this occupies the independent-instrument corner, and the neighbors measure different things. Consumer Reports Ratings Hub also labs its subjects but compresses them into survey-backed reliability predictions rather than continuous measurement suites. GSMArena Phone Database and Icecat Open Catalog document specifications and manufacturer content at far greater breadth - 14,807 phones and 30 million-plus data-sheets - but neither observes hardware, only describes it. ENERGY STAR Certified Products certifies efficiency compliance, not performance.

Pairing beats choosing. ENERGY STAR covers the efficiency claim, GSMArena and Icecat cover the retail-ready description, Consumer Reports covers ownership reliability - and RTINGS supplies the only standardized measurement layer among them. If the question is which TV is actually brighter, this is the record; if it is which SKUs exist, start with a catalog source and join on product identifiers.

What should I know before building on it?

Three things deserve respect in the rows. First, coverage follows the editorial test queue: the 4,847 units are the ones the source chose to purchase and measure, so joins into your own catalog want product-identifier mapping - a step handled in delivery rather than left to you. Second, methodology has versions: test_bench marks which protocol produced each number, and honest cross-era comparisons respect the vintage instead of treating a decade of scores as one flat series. Third, scores can move after publication because firmware changes are re-tested - which is precisely why scheduled deliveries accumulate dated observations, turning score movement itself into a signal about how a vendor maintains its products over time.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary — RTINGS.com Lab-Tested Reviews (definitions verified against live review pages)
fieldtypedefinitionexample
product_namestringReviewed product name exactly as shown in the review header.Samsung S95H OLED
categorystringProduct category path the review sits under.tv
test_benchstringVersion of the testing methodology the review was measured under; older reviews retain their historical bench versions rather than being restated.2.2
test_namestringName of the individual test or scored metric within the protocol.Native Refresh Rate
scorenumberNumeric score from 0 to 10 for that test or usage scenario.9.5
rendered_valuestringHuman-readable measured result as the source renders it.165Hz
usage_namestringUsage scenario being scored. Televisions carry five: Mixed Usage, Home Theater, Bright Room, Sports and Gaming.Mixed Usage
unblurredbooleanWhether the score value is visible to non-members upstream.false
insider_onlybooleanWhether the individual test result is restricted to members.false
statusstringTest state for the metric.tested
linked_descriptiontextEditorial summary paragraph accompanying the usage verdict.on request
worth_comparisonstringCheaper alternative flagged by the source's worth engine for this product.TCL S325
upgrade_comparisonstringHigher-tier alternative the source's worth engine points buyers toward.LG G5 OLED

Questions buyers ask

How many products does RTINGS.com have lab test results for?

Roughly 4,847 products bought and tested across 28 categories, from televisions and monitors through headphones, appliances and even mattresses and running shoes. The TV branch alone holds several hundred full reviews, each carrying dozens to hundreds of individual bench tests under the current versioned protocol.

What usage scores do televisions receive?

Five scenarios: Mixed Usage, Home Theater, Bright Room, Sports and Gaming - plus performance sub-scores such as Brightness in SDR and HDR. Scoring a set per scenario rather than once overall is what makes the same panel comparable between a dark media room and a sunlit living room.

Do scores change after a product is reviewed?

Yes, occasionally downward or upward: units are physically re-tested when firmware updates change measured performance, so a score can move after you have collected it. Deliveries arrive dated, so a re-test lands as a visible movement in your tables rather than a silent overwrite.

How far back does the testing history go?

Reviews date to 2011, and older reviews retain the historical test-bench version they were measured under instead of being restated to the current protocol. That preserves comparability discipline - a bench 2.2 result stays a bench 2.2 result, marked as such in every row.

Can a sample be scoped to my categories or brands?

Yes. Name the categories - televisions, monitors, headphones, whatever you merchandise - and the sample returns cut to that shape with the full schema intact, one row per test observation keyed on product and test name so it joins straight onto your own catalog identifiers.

What makes this different from other product review data?

Instruments instead of opinions. Every number originates from a controlled bench procedure executed on a unit the source bought at retail, under a documented versioned protocol - not aggregated star ratings or user surveys. It is the only standardized measurement layer among Datadory's eight primary consumer-electronics sources.

Notes on this record

  • Measured, not claimed 4,847 units bought at retail and benched in Montreal since 2011 - the only instrument-based measurement layer among the catalog's eight primary consumer-electronics sources.
  • Five scores per television Mixed Usage, Home Theater, Bright Room, Sports and Gaming, plus sub-scores like Brightness in SDR/HDR - scenario resolution no spec sheet offers.
  • Re-tested when firmware moves Units go back on the bench when updates change performance, so a score reflects the product as it exists now, not as it shipped.
  • Methodology versioned Test bench 2.2 is current for TVs; older reviews keep their historical bench versions, keeping cross-era comparisons honest.
  • Recommendation graph included The worth engine's cheaper alternative and upgrade pick travel on every product, turning positioning questions into a join.
  • Joins handled in delivery Rows key on product and test name; identifier mapping into your catalog happens before the feed lands, not after.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing