Datadory notebook
Systems Software Data Guide: 25 Datasets, Coverage & Uses
A systems-software data guide in 2026 starts with three source families - package registries, vulnerability feeds and platform telemetry. This guide maps all 25 systems-software datasets Datadory catalogs - 24 primary plus 1 related - with the fields, coverage and sample detail behind each, from npm's 4.3 million package documents to Backblaze's 530 million drive-days of telemetry.
1,744 datasets. Pick your catch.
The systems software data landscape in 2026
Systems software data clusters around registries, security feeds and platform telemetry, and nearly all of it originates with the infrastructure's own operators rather than with resellers. As of August 2026, Datadory catalogs 25 systems-software datasets - 24 primary records plus 1 related entry, Libraries.io Open Source Packages, filed under application software. Together they reach from single-package install events to half a billion drive-days of hardware telemetry.
A security layer pairs NIST's NVD - roughly 381,418 CVE records with CVSS scores and CPE applicability statements reaching down to product version - against OSV.dev's 848,582 advisories across 38 ecosystems, keyed to package name, version range and commit.
Platform telemetry and adoption benchmarks cover the rest of the pool. GitHub's Octoverse aggregates count 180M+ developer accounts, Stack Overflow contributes per-respondent survey microdata back to 2011, endoflife.date maintains a 464-product lifecycle catalog, and StatCounter's OS-share series reaches back to January 2009 alongside W3Techs' server-side Unix-versus-Windows split, Valve's Steam Hardware Survey, TIOBE and PYPL.
A smaller hardware-and-commerce tail rounds out the slice: Backblaze's fleet-wide SMART telemetry, TOP500 supercomputer rankings reaching back to June 1993, Icecat's 30.2-million-datasheet parts catalog, OEC's $60.2 billion trade profile for storage devices, the Census Bureau's wholesale-trade series, Amazon's computer-component best-seller ranks and the archived AnandTech benchmark database. Quality runs high: the 24 primaries average 8.4 on Datadory's 10-point rubric against a 7.81 catalog-wide mean, and eight score perfect 10s. Every record below is indexed with field dictionaries and coverage detail on the systems-software data hub.
The core datasets to know
The 24 primary records sort into five working groups: package registries and dependency graphs (5), vulnerability and lifecycle feeds (3), platform telemetry and surveys (5), language-popularity trackers (3), and hardware, trade and commercial layers (8).
Package registries and dependency graphs
deps.dev - Open Source Insights API (quality score 10) resolves dependency graphs, advisories and OpenSSF Scorecard checks across seven registries (npm, Go, PyPI, Maven, Cargo, NuGet, RubyGems), so supply-chain health becomes comparable across ecosystems.
Repology - Packaging Hub (quality score 9) tracks 324,959 projects and 4,900,782 individual packages across 173 repositories spanning Linux, BSD, macOS, Android and Solaris-derived ecosystems - the cross-distro matrix that shows exactly where versions diverge.
Vulnerability and lifecycle feeds
OSV - Open Source Vulnerability Database (quality score 10) listed 848,582 advisories across 38 package ecosystems on 21 August 2026, following a shared JSON schema with affected ranges down to package version and commit - close enough to join straight onto a lockfile.
NVD CVE & CPE Data Feeds and APIs (quality score 10) holds about 381,418 CVE records with CVSS scores and CPE applicability statements reaching down to product version, covering 2002 through 2026.
endoflife.date Product Lifecycle Catalog (quality score 10) tracks end-of-life and support-window dates for 464 products - operating systems, databases, languages, frameworks and services - each with typically 3-30 release cycles, each record linking the vendor's own policy page.
Platform telemetry and developer surveys
GitHub Octoverse and GitHub REST/GraphQL API (quality score 9) pairs repository, commit and topic activity with annual Octoverse aggregates - the 2025 edition counts 180M+ developer accounts and 36M new accounts in the past year - while the GH Archive preserves public event history back to February 2011.
Stack Overflow Annual Developer Survey (quality score 9) releases per-respondent microdata for every edition from 2011 through 2025; the 2024 file covers 65,437 respondents x 114 columns across 185 countries.
StatCounter OS Market Share (quality score 9) charts desktop, mobile, tablet and console OS share worldwide, across six regions and down to individual countries from January 2009 to present, drawn from a sample exceeding 3 billion pageviews per month on more than 1 million websites.
Steam Hardware & Software Survey (quality score 8) reports Windows 11/10, macOS and Linux shares plus CPU cores, GPU model, VRAM, RAM, resolution and VR headsets, filterable by platform; category charts retain roughly 18 months of history and video-card listings run to hundreds of models.
W3Techs OS Usage for Websites (quality score 7) covers the server side - Unix-family versus Windows across surveyed websites - complementing client-side panels with what actually serves production traffic.
Language-popularity trackers
TIOBE Index (quality score 7) has ranked roughly the top 50 languages since June 2001 - about 300 editions - with very-long-term annual averages back to 1986.
PYPL Programming Language Popularity (quality score 7) computes monthly share for 29 languages from tutorial-search data, worldwide plus United States, India, Germany, United Kingdom and France, with observations reaching back to June 2004. Because PYPL publishes normalized shares rather than absolute counts, it pairs naturally with TIOBE's ratings as a second opinion.
GitHut - GitHub Language Statistics (quality score 6) contributes quarterly pull-request, push, star and issue counts per language from Q2 2012 through Q1 2024 - about 19,000 metric rows built on GitHub event archives.
Hardware, trade and commercial layers
Backblaze Hard Drive Test Data (Drive Stats Raw Dataset) (quality score 10) is the slice's most distinctive record: one record per drive per day with SMART attributes, failure flags and datacenter placement for the whole fleet since April 2013 - about 341,000 active drives and roughly 530 million cumulative drive-days.
TOP500 Supercomputer Lists and Statistics (quality score 9) ranks 500 systems twice a year for every edition since June 1993 - roughly 33,000 system records through June 2026 - with Green500 efficiency rankings from June 2013 and HPCG results from November 2017, each record naming vendor, country, interconnect, accelerators and OS family.
Icecat Open Catalog (quality score 8) carries 30,266,450 datasheets across 29,756 brands with standardized specs, images and GTIN mapping - the parts dictionary underneath most product-data joins.
Observatory of Economic Complexity - Hard Disk Drives Trade Profile (quality score 8) pre-aggregates HS 847170 (computer storage devices including hard disk drives): $60.2 billion of global trade in 2024 across roughly 180 exporter rows per year, with the current cube covering 2022-2024.
U.S. Census Bureau Data API - CBP / Economic Census / Wholesale Trade (quality score 9) opens about 1,800 Census datasets: County Business Patterns annually 1986-2023 by NAICS industry, geography and employment-size class, the wholesale-trade Economic Census subject series (latest vintage 2022) and continuous Monthly Wholesale Trade indicators.
The remaining three are narrower. Amazon Best Sellers - Computer Components (quality score 6) tracks sales ranks for 50 products per node, fine-grained enough for day-over-day movement. AnandTech Bench (quality score 5) froze when publication ceased on 30 August 2024; its CPU-2019 suite alone spans 155 tests x 286 products. GTDC - Knowledge Hub & Member Directory (quality score 6) lists 21 member distributors with research coverage of the IT distribution channel.
Systems software data by use case
Ten recurring jobs define how teams use this pool, and each maps to named records rather than generic categories:
Depth, identifiers and granularity: what separates usable systems-software sources
The strongest systems-software records share three traits, and each decides whether a project survives contact with production:
None of these traits shows up in a screenshot or a landing page; all three are recorded in the field dictionaries and coverage detail on each dataset's page.
Who uses this data?
Four working groups drive most systems-software data searches, and each arrives with its own phrasing.
Developers and data-product builders build directly on registry and advisory records: package registries publish per-object histories, vulnerability feeds key advisories to versions and commits, and the lifecycle catalog tracks 464 products' support windows. Catalog-wide, 1,564 of 1,744 datasets (89.7%) rate as at least moderately relevant to this persona.
Market researchers and consultants type "market sizing data sources" and "where do consultants get their data." Their citable anchors here are Stack Overflow's per-respondent microdata, StatCounter's country-level OS-share series, GitHut's twelve-year quarterly panel and TOP500's 33,000-record archive - institutional provenance that survives client scrutiny, ranked on the market researchers hub.
Journalists, academics and students ask for cited data sources for academic papers. In this slice the citation-friendly properties are explicit - NIST's vendor-neutral CVE record set, survey microdata with a recommended BibTeX entry at GitHut, and TOP500's continuous 1993-2026 ranking series.
How to build a systems software data stack
Adapting the workflows that perform best across this catalog produces six steps:
- Anchor the registries. npm, PyPI and crates.io together cover JavaScript, Python and Rust adoption with day-resolution install histories for the three largest package ecosystems.
- Resolve dependencies and advisories together. Join deps.dev's dependency graphs and Scorecard checks to OSV advisories on package-version keys so every vulnerable edge carries a fix range.
- Gate on lifecycle dates. Cross endoflife.date's 464-product catalog against StatCounter's installed-base shares so upgrade work fires before vendor support lapses.
- Layer benchmarks last. Treat StatCounter, Stack Overflow microdata, TIOBE, PYPL and GitHut as a triangulation set - no single index settles a language or platform question on its own.
- Add hardware telemetry only where the question reaches silicon. Backblaze's per-drive panel covers reliability modeling; TOP500's per-edition lists cover HPC capacity.
- Decide the longitudinal job early. Registry records describe current state; the multi-year depth lives in cumulative records like the GH Archive and Backblaze's drive panel, so pick your longitudinal backbone before the first analysis.
Systems software data: frequently asked questions
Which vulnerability databases cover open source packages?
Two anchor the space. OSV.dev serves 848,582 advisories across 38 ecosystems, with affected ranges keyed to package name, version range and commit. NIST's NVD complements it with roughly 381,418 CVE records carrying CVSS scores and CPE applicability statements down to product version.
How do I check when a software version reaches end of life?
endoflife.date tracks end-of-life and support-window dates for 464 products - operating systems, databases, languages, frameworks and services - each typically split into 3-30 release cycles, so you can look up exactly when a given version lapses and which successor applies.
What is the best data source for operating system market share?
StatCounter GlobalStats charts desktop, mobile, tablet and console OS share worldwide, by region and by country from January 2009 to present. Pair it with W3Techs for the server-side Unix-versus-Windows split and Valve's Steam Hardware Survey for gaming PCs.
Where can I find hard drive failure rate data?
Backblaze Drive Stats covers one record per drive per day with SMART attributes, failure flags and datacenter placement for the whole fleet since April 2013 - about 341,000 active drives and roughly 530 million cumulative drive-days, granular enough to model failure rates by drive model.
How can I compare package versions across Linux distributions?
Repology normalizes 324,959 projects and about 4.9 million individual packages across 173 repositories spanning Linux, BSD, macOS and Solaris-derived ecosystems, aligning version strings so divergence between distros is visible at a glance.
Can I get raw developer survey data for academic research?
Yes. The Stack Overflow Annual Developer Survey releases per-respondent microdata for every edition from 2011 through 2025 - 65,437 respondents across 185 countries and 114 columns for 2024 - deep enough for crosstabs the published reports never run.
How far back do historical TOP500 supercomputer rankings reach?
Every edition since June 1993 is on record - two lists per year, 500 systems each, roughly 33,000 system records overall - covering vendors, countries, interconnects, accelerators and OS families. Green500 extends the series from June 2013 and HPCG results from November 2017.
Keep reading
This guide is the pillar of a systems-software cluster. Continue into the connected posts and hubs:
| measure | figure |
|---|---|
| Datasets cataloged for the industry | 25 pooled (24 primary + 1 related: Libraries.io Open Source Packages) |
| Mean Datadory quality score, primary slice | About 8.4 of 10 with 8 perfect scores (catalog-wide average: 7.81) |
| Largest single objects in the pool | npm's ~4.3 million package documents; PyPI's 20,805,417 files (44.8 TB); Repology's 4,900,782 packages |
| Security coverage | 848,582 OSV advisories across 38 ecosystems; ~381,418 NVD CVE records (2002-2026) |
| Hardware telemetry grain | ~341,000 Backblaze drives, ~530 million cumulative drive-days, one record per drive per day |
| Catalog-wide baseline | Average quality score 7.81 across 1,744 datasets; 62.8% score 8 or higher |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
npm Public Registry API
PyPI - Python Package Index API and Stats
crates.io - Rust Package Registry API
deps.dev - Open Source Insights API
publishedAt · isDefault · isDeprecated …+5 more
Repology - Packaging Hub
repo · srcname · visiblename …+5 more
OSV - Open Source Vulnerability Database
published · modified · summary …+4 more
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
Which vulnerability databases cover open source packages?
Two anchor the space. OSV.dev serves 848,582 advisories across 38 ecosystems, with affected ranges keyed to package name, version range and commit. NIST's NVD complements it with roughly 381,418 CVE records carrying CVSS scores and CPE applicability statements down to product version.
How do I check when a software version reaches end of life?
endoflife.date tracks end-of-life and support-window dates for 464 products - operating systems, databases, languages, frameworks and services - each typically split into 3-30 release cycles, so you can look up exactly when a given version lapses and which successor applies.
What is the best data source for operating system market share?
StatCounter GlobalStats charts desktop, mobile, tablet and console OS share worldwide, by region and by country from January 2009 to present. Pair it with W3Techs for the server-side Unix-versus-Windows split and Valve's Steam Hardware Survey for gaming PCs.
Where can I find hard drive failure rate data?
Backblaze Drive Stats covers one record per drive per day with SMART attributes, failure flags and datacenter placement for the whole fleet since April 2013 - about 341,000 active drives and roughly 530 million cumulative drive-days, granular enough to model failure rates by drive model.
How can I compare package versions across Linux distributions?
Repology normalizes 324,959 projects and about 4.9 million individual packages across 173 repositories spanning Linux, BSD, macOS and Solaris-derived ecosystems, aligning version strings so divergence between distros is visible at a glance.
Can I get raw developer survey data for academic research?
Yes. The Stack Overflow Annual Developer Survey releases per-respondent microdata for every edition from 2011 through 2025 - 65,437 respondents across 185 countries and 114 columns for 2024 - deep enough for crosstabs the published reports never run.
How far back do historical TOP500 supercomputer rankings reach?
Every edition since June 1993 is on record - two lists per year, 500 systems each, roughly 33,000 system records overall - covering vendors, countries, interconnects, accelerators and OS families. Green500 extends the series from June 2013 and HPCG results from November 2017.