Data source
Data from Python Software Foundation, delivered clean.
3 datasets pulled from Python Software Foundation's releases, checked field by field and shipped the way you want them — daily, weekly, or hourly, your call.
- 3 datasets
- 1 industry
- Real rows on request
What Datadory delivers from Python Software Foundation
3PyPI - Python Package Index API and Stats
crates.io - Rust Package Registry API
npm Public Registry API
Pick a catch, see the rows.
Name any Python Software Foundation dataset and we send real rows from it — not a screenshot of rows. 1,744 datasets. Pick your catch.
Get a sampleAPI, files, or your warehouse. Daily, weekly, or hourly.
Straight answers about Python Software Foundation data
How many Python packages does the data cover?
877,215 projects and 9,420,495 releases were registered at the August 2026 cataloging pass, spanning 20,805,417 distribution files and 44.8 TB of package content. All four counters move continuously as maintainers publish, so totals quoted at delivery reflect the pull you receive rather than a frozen brochure number.
What fields ship on every record?
Ten verified fields per release: project name, version string, maintainer-written summary, declared dependency specifiers, a Python-version floor such as >=3.9, artifact filename and type (wheel versus sdist), an ISO 8601 upload stamp, OSV-derived vulnerability records with CVE aliases and fixed-in versions, and sorted owner-maintainer roles.
How far back does the history reach?
Across the full life of the index: every release row carries its upload timestamp in ISO 8601 precision, and an immutable per-release metadata table retains rows even after a project deletes them. Cohort analysis by publication month therefore runs over the whole archive without stitching external records.
Is one dataset enough to work with?
For most questions, yes - the registry is the load-bearing table of an entire language, and one clean flat extract answers dependency screening, adoption sizing and wheel-coverage audits on its own. When a question needs a second corpus, Datadory joins it at delivery rather than handing you two schemas to reconcile.