Datadory notebook
RTINGS TV test results JSON: the bench measurements behind every score, delivered as rows
1,744 datasets. Pick your catch. Every guide here is built on what the catalog can actually prove.
1,744 datasets. Pick your catch.
What fields sit inside the RTINGS TV test results JSON?
That last pair is what no spec sheet anywhere offers. They encode the source's own recommendation graph - which cheaper model it steers buyers toward and which step-up it names - so brand-vs-brand positioning becomes a join instead of a reading assignment. It is also why this record is the only standardized per-usage measurement layer among the 8 primary datasets in Datadory's consumer-electronics pool; everything else there is spec sheets, certifications or prices. The full dictionary lives on the RTINGS.com Lab-Tested Reviews product page.
The matrix falls out naturally. Join on category plus test name to get one row per model per test, keep rendered_value beside the numeric score so units survive the trip into your warehouse, and carry the five usage scenarios as separate columns rather than collapsing them - Mixed Usage and Gaming measure different things by design.
What do sample rows look like?
One television, one test entry, exactly as it arrives:
The anatomy is the point. One row is one test observation inside one review: the unit identified by name, the methodology pinned by bench version, the instrument reading present both as a number and as the human-facing rendering, and the usage context attached alongside. A row therefore answers three questions at once - what was measured, how good it was, and what else a buyer's money should consider.
Televisions carry five top-level usage scores - Mixed Usage, Home Theater, Bright Room, Sports and Gaming - plus sub-scores such as Brightness in SDR and Brightness in HDR. Scored per scenario rather than once overall, the same panel stays comparable between a dark media room and a sunlit living room.
How large is the RTINGS TV corpus, and what does coverage look like?
Roughly 4,847 products have been bought at retail and benched across 28 categories since testing began in 2011, and the TV branch holds several hundred full reviews among them. The roster reaches far past displays: monitors, soundbars, projectors, headphones, speakers, mice, keyboards, printers, routers, laptops, vacuums, air conditioners, air purifiers, dehumidifiers, humidifiers, mattresses, cameras and running shoes all pass through the same discipline of buy-the-unit, run-the-bench, publish-numbers. Units are purchased at US retail rather than supplied by manufacturers, and testing happens in a 40,000 square foot Montreal facility - which is why every price reference in the corpus tracks the US market.
The honest caveat: population follows the source's editorial queue. The 4,847 units are the ones chosen to be bought and measured - absence means untested, not disapproved. Teams needing full-assortment coverage join this set against a catalog source like Icecat Open Catalog rather than expecting it from here.
Why do some RTINGS scores change after publication?
Methodology has versions too. Test bench 2.2 is current for TVs, and older reviews retain their historical bench version instead of being silently restated to today's protocol. That discipline matters when lining up a panel measured this year against one measured five years ago: the vintage travels on every observation, so cross-era comparisons stay honest instead of treating fifteen years of scores as one flat series.
How do teams actually use RTINGS TV test results?
- Competitive-intel and product teams benchmark their lineups against rival bench numbers - measured brightness, refresh behavior, input latency class - including the re-tests that move a competitor's standing after a firmware push. Nothing else in the slice observes rival hardware with instruments.
- Market researchers and consultants frame category studies around standardized per-usage scores instead of stitching together incompatible star-rating piles from a dozen publications.
- Data scientists and ML engineers use the 0-10 scores and rendered measurements as labels for quality-prediction models, with
test_benchmarking methodology vintages so training sets never silently mix protocols. - Retail buyers and merchandising teams let the five TV usage scores arbitrate assortment - a bright-room score for floor placements, a gaming score for the enthusiast end - backed by instruments rather than margins.
- Journalists, academics and students cite a consistent methodology run continuously since 2011, down to the bench version recorded on every observation.
Across the 8 primary consumer-electronics records the mean quality score is 6.5 out of 10; this one's 6 lands mid-pack on breadth, not rigor - nothing else in the pool measures picture performance at all. Where the workflow extends toward pricing, Price History Tracker (Amazon/Flipkart) adds multi-year daily price series - a sample listing spans 1,366 tracked days from November 2022 through August 2026 - so purchase timing joins onto measured performance.
Which datasets complement RTINGS TV measurements?
GSMArena Phone Database (score 7, 14,807 devices) and Icecat Open Catalog (score 8, 30,266,450+ data-sheets) document what vendors claim at enormous breadth - useful for identity mapping and retail-ready content, silent on observed performance. Consumer Reports Ratings Hub is the closest analog in spirit, expert-testing 395 currently rated TVs with about 500 data points collected per television, though it compresses findings into survey-backed reliability predictions rather than continuous measurement suites.
Retail-side depth arrives from an adjacent industry record: the Best Buy Developer API - Products, Stores & Categories scores 9 and covers 725,000+ products, attaching street price and availability to any model whose bench numbers you already hold. That gives a TV researcher three joinable layers - measurement, certification and price - none of which any single source here covers alone.
Why get the TV test results through Datadory?
Self-assembly, by contrast, buys you the raw plumbing: protocol vintages to reconcile, values whose visibility varies upstream, a corpus that reshapes as benches evolve, and identifiers that need matching before anything joins. Datadory absorbs all four on your behalf. To see it concretely, request a sample cut to your categories and brands - real bench rows come back before any commitment.
Where to go next
Start with the consumer electronics data guide, the pillar for this pool: all 14 pooled records scored and grouped by workflow, with the measurement layer placed beside the spec-sheet and registry sources. For the field-level view of the source itself - field dictionary, sample rows, coverage chips - read the RTINGS.com Lab-Tested Reviews product page, and browse the rest of the shelf in the consumer electronics data hub. The ranked best consumer electronics datasets list shows where this record lands against its neighbors. Modelers weighing these scores as training labels should check the data scientists use cases page first, and a sample cut of the TV bench feed is always one request a sample away.
| Field | Type | Definition | Example |
|---|---|---|---|
| product_name | string | Reviewed product name exactly as shown in the review header. | Samsung S95H OLED |
| category | string | Product category path the review sits under. | tv |
| test_bench | string | Version of the testing methodology the review was measured under; older reviews retain historical bench versions rather than being restated. | 2.2 |
| score | number | Numeric score from 0 to 10 for that test or usage scenario. | 9.5 |
| rendered_value | string | Human-readable measured result as the source renders it. | 165Hz |
| usage_name | string | Usage scenario being scored. Televisions carry five: Mixed Usage, Home Theater, Bright Room, Sports and Gaming. | Mixed Usage |
| unblurred | boolean | Whether the value is visible to non-members upstream. | false |
| insider_only | boolean | Whether the individual test result is restricted to members. | false |
| status | string | Test state for the metric. | tested |
| linked_description | text | Editorial summary paragraph accompanying the usage verdict. | on request |
| worth_comparison | string | Cheaper alternative flagged by the worth engine for this product. | TCL S325 |
| upgrade_comparison | string | Higher-tier alternative the worth engine points buyers toward. | LG G5 OLED |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
RTINGS.com Lab-Tested Reviews Data
product_name · category · test_bench …+10 more
ENERGY STAR Certified Products & Dataset
power_consumption_in_on_mode_watts
Consumer Reports Ratings Hub
11 documented · definitions verified …+8 more
Icecat Open Catalog
GSMArena Phone Database
Price History Tracker (Amazon/Flipkart)
event · is_partner_offer · count
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
How many TVs does the RTINGS test archive cover?
Several hundred full television reviews sit inside an archive of roughly 4,847 products bought and tested across 28 categories since 2011. Every TV is scored against five usage scenarios - Mixed Usage, Home Theater, Bright Room, Sports and Gaming - under a versioned protocol, so two reviews remain comparable years apart.
Can I get RTINGS TV scores scoped to my brands or categories?
Yes. A sample cut comes shaped to the categories or brands you name, with the full schema intact and one row per test observation keyed on product and test name, so it joins straight onto your own catalog identifiers. Request a sample and real bench rows come back before you commit to a feed.