Datadory notebook

Npm Pypi Package Dependency Dataset Data: Dataset Structure and Field Coverage

Datadory delivers npm pypi package dependency dataset data covering comprehensive field definitions, entity mappings, and historical time series — structured for direct analytics and delivered on demand.

1,744 datasets. Pick your catch.

Which dataset covers npm and PyPI package dependencies?

Access comes in two forms. A free JSON REST API exposes platforms, project metadata with version history, dependencies and dependents. A bulk open-data release on Zenodo ships the whole relational graph as CSV under commercial delivery terms-SA 4.0. Both routes are covered below, because they answer different questions: the API answers "what depends on left-pad today?" and the bulk dump answers "what did the entire ecosystem look like on a fixed date?"

How do you turn the dependency graph into supply-chain risk signals?

The seven-table archive is shaped for joins, so the analysis workflow is short:

Who uses npm and PyPI dependency data?

Investors and quant researchers are the sharpest-fit persona in Datadory's application-software slice, and the persona pack scores both Libraries.io and the GitHub REST API at relevance 2. The stated use case for Libraries.io: track dependency adoption and package growth as supply-demand evidence in open-source and dev-tool investment research. The GitHub record's: use star and commit trajectories as an adoption signal in diligence on developer-tool and open-source companies. A dependency graph is adoption telemetry — a framework whose dependents count compounds quarter over quarter is a different asset than one whose graph is flat.

If you are benchmarking which sources deserve engineering time, note that Libraries.io's quality score of 9 sits above the catalog-wide average of 7.81 across the 1,744 datasets Datadory catalogs, and its field definitions are verified rather than inferred.

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Application Software Global

Libraries.io Open Source Packages

SourceRank

Systems Software Global - the whole JavaScript and Node.js ecosystem in one…

npm Public Registry API

Systems Software Global - the entire Python package ecosystem, with no…

PyPI - Python Package Index API and Stats

Application Software Global - all public repositories hosted on GitHub, across all…

GitHub REST API (Repositories & Search)

Application Software Global free-and-open-source Android ecosystem

F-Droid Open Source Android Apps

Application Software Worldwide plus per-country website rankings

Similarweb Top Websites & App Intelligence

seven website fields · seven app fields …+11 more

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

Which dataset covers npm and PyPI package dependencies?

Libraries.io indexes 11,722,531 open-source packages across 32 package managers, including npm at 5.96 million and PyPI at 908,000, plus Go, Maven and NuGet. A free 60-requests-per-minute REST API serves package, version, dependency and dependents records, while bulk release 1.6.0 ships a 24 GB seven-table CSV archive under commercial delivery terms-SA 4.0.

How current is the Libraries.io dependency dataset?

The API tracks Libraries.io's live daily sync, but the bulk archive is frozen at 12 January 2020 with no newer release published, and the current libraries.io/data page no longer links the archives. Refresh recent edges against the live npm registry (about 4.3 million package documents) and PyPI (877,215 projects) APIs.

Can I use the Libraries.io dependency data commercially?

Yes, under commercial delivery terms-SA 4.0 terms: attribute Tidelift and apply share-alike to derivatives. The data is explicitly not validated or curated for accuracy — human-validated package data is sold via the Tidelift Subscription — and the API's 60 requests per minute cannot be circumvented, with higher limits available for a fee.