Datadory notebook
The top 1 million domains list: a consensus ranking of the web, delivered as rows
Datadory delivers the top 1 million domains list as typed rows: the Tranco Top Sites Ranking built at KU Leuven, one million pay-level domains per daily list since 2019, each rank a Dowdall-rule average of five provider rankings over a trailing 30-day window and pinned by a permanent citable list ID - delivered daily, weekly, or hourly.
1,744 datasets. Pick your catch.
What is the top 1 million domains list?
One record owns this query and Datadory delivers it as plain typed rows: Tranco Top Sites Ranking, the web-popularity ranking built at KU Leuven's DistriNet group. Every list holds exactly 1,000,000 pay-level domains - one row per registered domain, not hostname - and a fresh ranking has landed every day since the project's 2019 launch, each day retained as its own retrievable snapshot at roughly 10-15 MB compressed.
The design choice that matters is averaging. Any single popularity panel wobbles - a domain can swing hundreds of positions between days on one signal alone - and single-source charts have been gamed by attackers parking domains near the top. This list merges five independent measurements over a rolling 30-day window, and its maintainers state plainly they are unaffiliated with every provider involved. That construction is why it became the default ranking for reproducible internet-measurement work rather than marketing screenshots, and why each daily list carries a permanent citable list ID that pins the exact snapshot.
On Datadory's rubric the record scores 9/10 against a 7.81 average across the 1,744 datasets Datadory catalogs, inside an interactive media & services slice whose 16 primary records average 9.0 - the strongest slice in the catalog.
How is a domain's rank actually computed?
Five provider rankings feed each list, each measuring popularity from a different vantage point: Cisco Umbrella's DNS-resolution panel, Majestic's backlink index, Farsight's passive-DNS records, the Chrome User Experience Report's browser-experience telemetry and Cloudflare Radar's CDN-edge view.
The merge uses the Dowdall rule: a domain ranked first by a provider contributes a score of 1, second contributes 1/2, down to 1/N - so no single provider dominates the result. Bucketed ranks are normalized to the geometric mean of their boundaries before scoring, keeping coarse providers comparable with fine-grained ones. Summed, the scores produce the daily order.
The recipe has changed twice, and both dates matter to anyone modeling rank across years. Quantcast was dropped in 2020. On 1 August 2023 the Chrome UX Report and Cloudflare Radar entered the mix when the discontinued Alexa ranking was removed - ranks either side of that date are internally consistent but not strictly comparable, so segment long backtests at the seam rather than smoothing over it.
The payoff is manipulation resistance. Meaningfully moving a domain up this list means moving it simultaneously in DNS-resolution panels, backlink indexes, passive-DNS records, browser telemetry and CDN-edge data, for weeks. Short bursts of artificial traffic that flatter a daily single-source chart wash out in the 30-day window.
What do rows from the top 1 million domains list look like?
The August 25, 2026 list opens with two columns doing all the work:
# tranco top-1M daily list - head of list
rank : 1
domain : google.com
rank : 2
domain : cloudflare.com
rank : 3
domain : gstatic.com
rank : 4
domain : facebook.com
rank : 5
domain : microsoft.comgoogle.com at rank 1 illustrates the grain: the list ranks pay-level domains, so www.google.com, mail.google.com and docs.google.com collapse into one entry and Google's whole property network carries a single position. The same shape runs down all million rows - a domain and its consensus position - which makes the file join-ready against any domain-keyed table you already keep: certificate-transparency logs, DNS logs, ad-tech inventories, threat feeds. No entity-resolution pass required before the first join.
Rank semantics are stable across the entire history, so a rank from 2019 and a rank from this morning mean the same thing: five providers, 30-day window, Dowdall scoring. Several competing rankings changed methodology mid-stream without restating old values; this one documents its breaks. A subdomains variant exists where host-level grain matters - name it when requesting your sample.
What do teams build on a daily top-one-million feed?
Five workflows account for most of what gets built, mapped to the personas Datadory serves:
What can't the top 1 million domains list tell you?
Three boundaries shape honest claims. Geography: the standard list is one global popularity order - there is no per-country column, so regional questions need an explicitly geo-resolved pairing. Grain: rows are pay-level domains; anything below the registered domain appears only in the dedicated subdomains variant. Time resolution: the same 30-day window that suppresses manipulation also erases intraday spikes and one-day viral surges - the list shows a surge's aftermath, never its morning.
Two sizing facts round out the picture. Volume is fixed, not sampled: one million rows per daily list makes warehouse load arithmetic rather than guesswork, and custom configurations shorten the prefix below one million when only the head matters. And any multi-year model carries the documented methodological seam of 1 August 2023. None of this is a defect; these are the trade-offs printed on the label - know them before you publish a chart.
Which datasets pair with a top-one-million ranking?
A rank says where a domain sits. Neighboring records in the interactive media & services slice supply everything the rank does not:
- Wayback Machine CDX Server API - hundreds of billions of capture-index rows since 1996, carrying urlkey, timestamp, mimetype, statuscode, digest and length per fetch, so a top-ranked domain's actual content history sits one query away.
- M-Lab Open Internet Performance Data - petabytes of speed-test and traceroute measurements continuous since February 2009, measuring how fast users actually reach the destinations the ranking orders.
- Domain Name Industry Brief & DNIB Data - roughly 66 quarterly reports spanning 2010 Q1 through 2026 Q2 plus dashboards over approximately 330 monitored TLDs: the registration supply side beneath the demand-side ranking.
- IANA Root Zone Database - the authoritative registry layer, about 1,595 TLD rows (1,250 generic, 316 country-code, 14 sponsored, 11 test, 3 generic-restricted, 1 infrastructure).
- Stanford SNAP Large Network Collection - 80-plus social graphs scaling to com-Friendster's 1.8 billion edges, for when the question is who links to whom among the nodes being ranked.
Stacked, they answer what is big, what it contained, whether anyone could reach it, and how the namespace itself is organized.
Where to go next
Start with the Interactive Media & Services data guide for the full 19-record landscape, then the website history archive data walkthrough for the archival half of web measurement. Settle ranking-versus-network with vs Stanford SNAP Large Network Collection, browse best interactive media services datasets, read the Tranco source profile, or jump to the interactive media services data hub. Ready for rows? Request a sample of the Tranco Top Sites Ranking and name the domains, rank bands and date range - real rows come back first.
| Dataset | What it adds | Coverage depth |
|---|---|---|
| Tranco Top Sites Ranking | Consensus popularity order: one million pay-level domains, five providers averaged over 30 days | Daily lists since 2019, every day retained and citable by list ID |
| M-Lab Open Internet Performance Data | Arrival quality: how fast users reach popular destinations | Petabytes continuous since February 2009 |
| Domain Name Industry Brief & DNIB Data | Registration supply side by TLD group | Roughly 66 quarterly reports, 2010 Q1 through 2026 Q2, over approximately 330 monitored TLDs |
| IANA Root Zone Database | Authoritative TLD registry layer | About 1,595 TLD rows, kept current as delegations change |
| Stanford SNAP Large Network Collection | Graph structure among the nodes the ranking orders | 80-plus graphs up to com-Friendster's 1.8 billion edges |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
Tranco Top Sites Ranking
Wayback Machine CDX Server API
digest · timestamp
M-Lab Open Internet Performance Data
Domain Name Industry Brief & DNIB Data
quarter / reporting period · total_domain_name_registrations · com_net_domain_name_base …+3 more
IANA Root Zone Database
Stanford SNAP Large Network Collection
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
Is there still a top 1 million domains list after Alexa was discontinued?
Yes - the Tranco Top Sites Ranking remains the reference top-1M list, rebuilt daily at KU Leuven from five averaged provider panels. When the Alexa ranking was discontinued, the composition refreshed on 1 August 2023, adding the Chrome UX Report and Cloudflare Radar alongside Cisco Umbrella, Majestic and Farsight. Datadory delivers the current list plus every retained day since 2019.
Does the top 1 million domains list include subdomains?
The standard list ranks one million pay-level domains, one row per registered domain, so names below the domain do not appear. A dedicated subdomains variant handles hostname-level questions, and custom configurations shorten the prefix below one million when only the head matters. Name the grain when requesting a sample and the rows arrive cut to it.
How do I cite a specific version of the top 1 million domains list?
Every daily list carries a permanent citable list ID, so a paper or an audit references the exact snapshot used rather than 'latest'. Academic convention also cites the methodology paper by Le Pochat et al. Datadory pins the matching list IDs to every delivery, which keeps validation runs reproducible.
Why did a domain's rank move when nothing about the site changed?
Ranks move because five independent panels keep re-measuring the web around a domain. A rival's growth, seasonal DNS behavior or backlink churn shifts the Dowdall-scored average even when the domain's own traffic holds flat, and the 30-day window deliberately smooths short spikes. Diffing retained daily lists separates real share shifts from background noise.