SourceForge Software Directory Data

Datadory delivers sourceforge software directory data covering 500,000+ open-source projects hosted on SourceForge since 1999 - each row carrying the project name, description, weekly download count, review tally and star rating plus license, category, operating system, programming language and development status. Delivered daily, weekly, or hourly; get a sample cut to your segments first.

API, files, or your warehouse. Daily, weekly, or hourly.

Where it covers
Global open-source projects; operator based in San Diego, CA
How far back
Registrations from 1999 to the present; current-state listings, history accumulated via scheduled deliveries
How fine
Per-project records (~500,000+) plus facet-level aggregate counts per OS/category/license/language/status

What is the SourceForge software directory dataset?

One row per open-source project in the oldest large-scale software catalog still running. SourceForge launched in 1999 under VA Linux, passed to Geeknet, and has been operated by Slashdot Media since 2012; the platform reports roughly 20 million monthly users and over 2.6 million downloads per day. The directory that fronts those half-million projects sorts them across five facet families with live counters: operating system, category, license, programming language and development status.

What separates it from every commercial software directory is what got counted along the way. A listing row shows not just a name and blurb but weekly download volume, a review count and a star rating - adoption telemetry for a population of tools that mostly never touch an app store. KeePass alone shows 606 reviews backing a 4.9/5 rating; MinGW moves 3.8 million copies in a week. The project pages add per-star breakdowns and sub-ratings for ease, features, design and support, dated reviews with helpful-vote counts, registration dates and latest file-release details. Datadory turns the whole thing into one consistent twelve-field table keyed on the project.

Sample rows

Five rows exactly as captured during the August 2026 verification pass:

rank : 1
project_name        : MinGW - Minimalist GNU for Windows
description         : A native Windows port of the GNU Compiler Collection (GCC)
reviews             : 223
downloads_this_week : 3885542
last_update         : 2021-09-05

rank : 11
project_name        : KeePass
description         : A lightweight and easy-to-use password manager
reviews             : 606
downloads_this_week : 232700
last_update         : 2026-07-25

rank : 12
project_name        : Apache OpenOffice
description         : The free and Open Source productivity suite
reviews             : 356
downloads_this_week : 237297
last_update         : 2025-10-29

rank : 13
project_name        : WinSCP
description         : Free SFTP, SCP, S3, WebDAV, and FTP client for Windows
reviews             : 209
downloads_this_week : 219183
last_update         : 2026-06-17

rank : 18
project_name        : TortoiseSVN
description         : An Apache SVN client, right where you need it most
reviews             : 155
downloads_this_week : 118263
last_update         : 2026-08-12

Read the spread, not the names. MinGW's 3.86 million weekly downloads sit next to a last-update date of September 2021 - a finished tool still being installed by the million. TortoiseSVN's row was updated twelve days before the capture. That combination of adoption velocity and maintenance recency in one row is the reason this dataset behaves differently from any review-site export: it measures installation, not opinion.

Get a sample of this dataset filtered to your categories, languages or development-status bands, and it arrives in this exact shape.

What fields does the dataset include?

Twelve fields verify against the directory itself, every definition checked during the August 2026 pass and every example captured from a live row:

What do the facet counts reveal?

Beyond the per-project rows, the directory publishes live aggregate counts per facet value, and those aggregates are themselves citable structure:

FacetLeading values
Operating systemWindows 202,909 / Linux 195,959 / Mac 151,071 / BSD 88,579 / ChromeOS 56,859
CategorySoftware Development 48,432 / System 32,192 / Internet 26,170 / Games 23,902 / Multimedia 21,774 / Business 19,715 / Scientific 19,589
LicenseOSI-approved 176,996, plus general public catalog and Creative Commons Attribution bands
LanguageJava 43,736 / C++ 36,171 / C 27,226 / PHP 23,527 / Python 19,175 / C# 14,507 / JavaScript 14,333
StatusProduction/Stable 48,464 / Beta 45,011 / Alpha 29,543 / Pre-Alpha 23,501 / Planning 21,554 / Inactive 5,566 / Mature 4,503

Two readings fall straight out of that table. First, maturity is rare: only 4,503 projects carry the Mature tag while nearly twice as many (48,464) are merely Production/Stable, and 5,566 are openly Inactive. Second, the language distribution is a two-decade sediment layer - PHP's 23,527 projects outnumber Python's 19,175 despite everything written since about both - because projects rarely leave once registered. Datadory delivers these facet rollups alongside the row-level extract so segment sizing needs no second pass.

What does coverage look like?

  • Geography: global. Projects register from anywhere; the operator sits in San Diego, California. There is no country partition to opt into, so geographic cuts ride on project metadata rather than a coverage axis.
  • Temporal: depth without snapshots. Registration dates run back to 1999 and last-update dates run to the present week, but the directory publishes only current state - no archival releases behind it. A time series of download velocity is something you accumulate from scheduled deliveries, which is exactly how Datadory schedules them.
  • Granularity: one record per project, roughly 500,000 of them, with facet-level aggregate counts available as a companion cut. Both the per-project view and the per-facet rollup are native to the source structure.
  • Population note: the same operator lists about 122,200+ business software titles beside the open-source catalog. Datadory scopes the open-source directory by default; the business-software shelf is a separate request line.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Samples precede any commitment, and full deliveries land keyed on project name with the facet rollups attached, so each pull joins cleanly to the last and to whatever else you already hold. Weekly-download figures are the field that rewards cadence: sampled hourly they become a genuine velocity curve rather than a single number.

Who puts this data to work?

Ranked by fit recorded in Datadory's persona index:

  1. Developers & Data-Product Builders - seed dependency and component catalogs with a 27-year-deep registry of tools, licenses and platforms (developers builders x application software).
  2. Data Scientists & ML Engineers - train popularity, maintenance-risk and abandonware classifiers on download velocity paired with status and last-update labels (data scientists x application software).
  3. Investors & Quant Researchers - read download momentum across open-source infrastructure categories as a bottom-up signal on developer-tool demand (investors quants x application software).
  4. Sales & Growth Teams - build technographic account lists from the OS and language facets, then target admins of the tools their product replaces (sales growth teams x application software).
  5. Market Researchers & Consultants - size open-source segments with facet counts instead of vendor marketing claims (market researchers x application software).
  6. Journalists, Academics & Students - cite a verifiable 27-year longitudinal record of what developers actually install (journalists academics x application software).

How does it compare to the alternatives?

Capterra software directory covers commercial buyers rating paid products - 86,267 profiles across 1,025 categories with starting prices - but holds no download telemetry at all. GitHub repositories search tracks code activity: stars, forks, commits. Neither measures installation, which is SourceForge's whole edge - 3.86 million weekly downloads for MinGW is behavior, not sentiment. Libraries.io packages approaches the same population through package managers and catches modern ecosystems better, while SourceForge retains the legacy enterprise stack - FTP clients, GCC ports, ERP add-ons - that never shipped through npm. The practical pattern is adoption from here and activity from GitHub, joined on project identity. The industry hub lays out the full stack.

What are the limitations?

Stated plainly, because they shape the analysis:

  • No archive. The directory exposes current state only. Download histories, deleted projects and past ratings mean accumulating deliveries over time - a scheduling decision, not a research project.
  • Ratings run thin outside the head. Review counts concentrate heavily: KeePass's 606 reviews coexist with thousands of projects holding zero. Median-row analysis should lean on downloads and last-update, which populate far more evenly.
  • Self-reported metadata drifts. Categories, status tags and OS support are set by maintainers, some of them gone for a decade; the 5,566 Inactive rows prove the tag exists but nothing enforces its use.
  • Legacy skew is real. The catalog over-represents projects registered in the 2000s. For post-2015 ecosystems, package-manager sources complement it rather than compete.

Why request this through Datadory

Because the raw artifact is faceted HTML meant for a browser, and most questions want a table. Datadory normalizes listing rows and project pages into the twelve-field dictionary above, attaches the facet rollups so segment sizing ships with the extract, cuts samples to your categories, languages or status bands, and schedules deliveries so download velocity becomes a longitudinal asset instead of a screenshot you re-take. Browse the rest of the shelf on the application software data hub, the best application software datasets ranking, or the full catalog.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - 12 verified fields, one row per project
fieldtypedefinitionexample
project_namestringName of the open-source project as listed in the directory.KeePass
descriptiontextOne-line summary plus longer blurb describing the project.A lightweight and easy-to-use password manager
downloads_per_weekintegerWeekly download count shown on the listing row and project page header.232700
review_countintegerNumber of user reviews submitted for the project.606
ratingnumberOverall user rating out of 5 stars, with per-star breakdown and sub-ratings for ease, features, design and support.4.9
last_updatedateDate of the most recent project update or file release.2026-07-25
licensestringOpen-source license declared on the project page.GNU General Public License version 2.0 (GPLv2)
categoriesstringDirectory categories assigned to the project.Business, Database, Security
operating_systemsstringSupported operating systems listed in project details.Linux, Mac, Windows
programming_languagestringImplementation languages listed in project details.C#, C++
statusenumDevelopment status: Planning, Pre-Alpha, Alpha, Beta, Production/Stable, Mature, Inactive.Production/Stable
registereddateDate the project was registered on SourceForge.2003-11-15

Facet rollups - aggregate counts delivered alongside the row-level extract

fieldtypedefinition
os_facet_countsmapLive project counts per operating-system facet: Windows 202,909, Linux 195,959, Mac 151,071, BSD 88,579, ChromeOS 56,859.
category_facet_countsmapLive project counts per directory category, led by Software Development 48,432, System 32,192, Internet 26,170, Games 23,902, Multimedia 21,774, Business 19,715.
license_facet_countsmapLive project counts per license band, led by OSI-approved 176,996, with general public catalog and Creative Commons Attribution bands.
language_facet_countsmapLive project counts per programming language: Java 43,736, C++ 36,171, C 27,226, PHP 23,527, Python 19,175, C# 14,507, JavaScript 14,333.
status_facet_countsmapLive project counts per development status: Production/Stable 48,464, Beta 45,011, Alpha 29,543, Pre-Alpha 23,501, Planning 21,554, Inactive 5,566, Mature 4,503.

Dataset facts at a glance

attributevalue
records~500,000+ open-source projects
fields verified12 per project row, plus 5 facet rollup maps
deepest historyRegistrations back to 1999
populationOpen-source projects; ~122,200+ business titles listed separately by the same operator

Questions buyers ask

How many projects are in the SourceForge software directory?

Over 500,000 open-source projects appear in the directory, with facet counters showing 202,909 Windows-tagged, 195,959 Linux-tagged and 176,996 OSI-approved-license projects among them. The same operator lists about 122,200+ business software titles in a separate catalog alongside the open-source one.

What fields come with every SourceForge project row?

Twelve fields verify against the directory: project name, description, weekly download count, review count, overall star rating, last-update date, license, categories, operating systems, programming languages, development status and registration date. Aggregate facet-count maps for OS, category, license, language and status ship alongside the row-level extract.

Does the dataset include download history per project?

Each row carries the current weekly download figure, and the directory publishes no archival snapshots behind it. A true download-history series comes from accumulating scheduled deliveries over time, which is a scheduling decision Datadory handles rather than something recoverable retroactively.

Which geographies does the SourceForge data cover?

Projects register from anywhere in the world and the operator sits in San Diego, California, so coverage is global with no country partition built into the source. Geographic cuts ride on project metadata such as translation and audience fields rather than a dedicated coverage axis.

Can a sample be filtered to my categories or languages?

Yes. Samples can be cut to any combination of directory facets - category, operating system, programming language, license band or development status - and arrive in the same shape as the full delivery, keyed on project name so the sample joins cleanly to whatever follows it.

How fresh is the data in each delivery?

Listing rows reflect the current weekly download figure, latest update date and standing review counts at collection time, and deliveries run daily, weekly, or hourly depending on the schedule you pick. The TortoiseSVN row above, updated twelve days before the August 2026 verification pass, shows the typical lag floor.

Notes on this record

  • Provenance Source name: SourceForge, the open-source hosting platform founded in 1999 and operated since 2012 by Slashdot Media, reporting ~20 million monthly users and 2.6M+ downloads a day.
  • Adoption telemetry Weekly download counts, review tallies and star ratings with ease/features/design/support sub-scores - installation behavior no review marketplace records.
  • Facet structure Five facet families with live counters - OS, category, license, language and development status - delivered as rollup maps beside the row extract.
  • Longitudinal depth Registration dates reaching back to 1999 make this one of the longest continuous records of what developers actually install.
  • Honest limits Current state only, self-maintained metadata, and review counts concentrated in the head - stated up front because they shape the analysis.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing