Datadory notebook

Crossref DOI Bulk Export: 185.7 Million Works, Delivered as Rows

Datadory delivers publishing data covering the DOI spine of scholarly communication: 185,678,115 registered works spanning journal articles, book chapters, monographs, edited books, proceedings pieces, reports, dissertations and grants, beside roughly 169,622 journals, 33,393 member publishers and 45,739 funders - each row keyed on its DOI with contributor lists, affiliations, funding acknowledgements, deposited reference lists and citation counts riding the same record - normalized into typed rows and delivered by API, files, or straight into your warehouse, daily, weekly, or hourly.

1,744 datasets. Pick your catch.

What is a Crossref DOI bulk export?

The query shows up wherever someone needs scholarship as a table instead of a website: bibliometric studies, publisher market-share decks, training corpora, systematic-review screens. A bulk export of Crossref is a machine-readable mirror of scholarly publishing metadata exactly as member publishers deposit it, and one named product answers it outright: the Crossref REST API - Scholarly Publishing Metadata record, scored 10 out of 10 on Datadory's quality rubric and one of only 145 perfect scores across the 1,744 cataloged.

The scale behind the term: 185,678,115 registered works observed during the August 2026 research pass, spanning journal articles, book chapters, monographs, edited books, proceedings pieces, reports, dissertations and grants - with rollups across roughly 169,622 journals, 33,393 member publishers and 45,739 funders. Each work row carries its DOI, title, contributor list with affiliations and roles, funding acknowledgements, the deposited reference list, an inbound citation count and the registration timestamps that make change detection mechanical.

Datadory delivers that corpus as typed rows keyed on DOI - whole registry or cut to your named slice - through whichever channel your pipeline reads first.

How much corpus does a bulk export actually move?

A work record looks light until you multiply it. The sampled chapter below runs a compact set of core fields and still ships a 66-entry deposited reference list; add abstracts where publishers supplied them and reference arrays across 185.7 million works, and a whole-corpus cut is measured in terabytes before any derived table exists. Venue-level depth is just as concrete: one Elsevier serial in testing held 9,548 total DOIs, split 1,089 current and 8,459 backfile - the registry measures depth per title, not merely in aggregate.

That arithmetic settles the design question before any pipeline work begins: scope beats brute force. Cutting by publisher, ISSN or ISBN, work type, funder or date window turns one unbounded pull into a series of bounded, resumable jobs - per journal, per member, per cohort year - each small enough to schedule, diff and rerun. The market sizing workflow runs on precisely that shape: named slices, comparable definitions, repeatable totals.

What does one delivered row look like?

One row per DOI-registered work, exactly as a delivery lands:

# one row per DOI-registered work - the delivered grain
DOI                    : 10.1007/978-3-658-17671-6_18-1
title                  : Soziale Innovation
type                   : book-chapter     # journal-article | proceedings-article | ...
publisher              : Springer Fachmedien Wiesbaden
publisher-location     : Wiesbaden

# the signals most bibliographic feeds leave out ride on the same row
reference-count        : 66               # references deposited for this work
is-referenced-by-count : 2                # citations counted across the registry

# venue-level rollup: one Elsevier serial seen from the journal side
counts.total-dois      : 9548             # 1089 current, 8459 backfile

Read the anatomy before the digits. The DOI keys every join - outward onto citation graphs, identifier registers and trade catalogs alike; the Digital Object Identifier glossary entry covers the mechanics. reference-count and the deposited reference list describe what a work looked at; is-referenced-by-count describes what the field did with it - 66 out, 2 in on the sampled chapter, and keeping the two directions distinct is the difference between bibliometrics and folklore. created fixes first registration and never moves; indexed advances whenever a correction is processed, so consecutive deliveries diff cleanly into new registrations, corrections and citation growth.

Where do DOI registration records stop?

The registration layer answers who published what, when, and who cites it. Three questions it structurally cannot answer get records of their own.

  • Breadth beyond the registry. The OpenAlex API - Scholarly Book & Publisher Graph widens the lens to 322,044,938 indexed works - among them 5,906,748 books and 22,015,762 book chapters - beside roughly 125.7 million authors, 255,627 sources and 10,708 publishers carrying corporate-lineage fields, all joined by citation-network edges rather than name-string surgery.
  • The preprint instant. The arXiv API - e-Print Publishing Preprints holds the version before any journal said yes: a corpus commonly cited above 2.5 million papers, with submissions reaching back to 1991 - the currency check for any STEM literature scan.
  • The journal business layer. The DOAJ API - Directory of Open Access Journals curates 23,352 journals and 13,470,084 article records from publishers in more than 130 countries, each journal carrying APC, peer-review-model and preservation fields alongside the work-level volume the DOI spine counts.

Abstract coverage deserves its own honesty: abstracts are optional deposits, populated unevenly by design rather than by defect, and much of what exists arrives JATS-tagged. Screen presence against your named scope at sampling if abstract text drives the question.

How does the DOI spine join into trade-book catalogs?

The scholarly and trade hemispheres meet on identifiers, not titles. Crossref stamps ISBN and isbn-type onto book-chapter, monograph and edited-book records - the hook that hangs edition-level commerce off DOI-registered scholarship.

From there the join fans out across three records. The Wikidata publishing knowledge graph contributes 122,983,238 typed entities, linking editions, authors and publishers through ISBN-13 (P212), title (P1476), author (P50) and publisher (P123). The International ISBN Agency's global publisher registers validate any edition identifier against 287 registration groups, 1,871 registrant range rules and a global register holding over a million publisher prefixes. And the Open Library Monthly Data Dumps supply the trade baseline - roughly 12.4 GB compressed across record types of editions, works, authors and ratings.

The rule this stack enforces: match each layer to the record built for it, and keep the DOI as the spine everything else hangs from. When the choice narrows to registration versus catalog, the Crossref versus Open Library comparison settles it row by row.

Who builds on DOI-scale scholarly metadata?

  • Bibliometrics and research-evaluation teams turn is-referenced-by-count across 185.7 million works into metrics that survive peer review, with deposited reference lists exposing who cites whom - the workflow lives at citation-grade research.
  • Funding analysts and science-policy researchers join acknowledged funders against the funder directory to show who pays for which science - diligence-grade evidence, curated on the journalists & academics persona.
  • Systematic-review builders collapse weeks of manual screening passes into fielded filters across type, publisher, ISSN and funder - the same slices that make the bulk build schedulable in the first place.

Why get the bulk export through Datadory?

Because the corpus moves whether or not your extract does. DOIs have been minted continuously since 2000 and deposits keep arriving, so any static pull ages from the moment it completes - which is why a delta discipline beats repeated monolithic rebuilds. Deliveries keyed on DOI make that discipline cheap: consecutive pulls diff cleanly into new-registration and citation-growth timelines instead of piling up as undifferentiated snapshots.

Datadory normalizes before delivery. One field dictionary kept current as deposit practice evolves, funder and member directories resolved into their own tables, reference lists typed as structured rows rather than embedded blobs, and every slice sharing the same join-key spine - so the scholarly leg and the trade leg meet on one column instead of fuzzy matching.

Files, feeds, or straight into your warehouse. Daily, weekly, or hourly - your call. Extraction plumbing, pagination walks and change detection stay our problem. Start with a sample: name the publishers, prefixes, work types, ISSNs or date windows you need, and the extract arrives cut to exactly that scope with the field dictionary attached. The schema in the sample is the schema you ship against.

Where to go next

Start with the publishing data hub, which browses all nineteen pooled datasets in the industry as records, and the best publishing datasets ranking that scores ten primary records side by side - led by the DOI spine at a perfect quality score. The publishing data guide maps who publishes each record, what its fields hold and how the sources join; the Crossref source profile covers provenance.

And when you're ready to build, request a sample scoped to your publishers, years and fields - the 185,678,115th work loads from the same table as the first.

The stack behind a DOI-scale scholarly metadata build (as of August 2026)
RecordLayerGrainCoverage depth
Wikidata Publishing Knowledge Graph + SPARQLIdentifier join: editions, authors, publishers as typed entitiesEntity x property statements122,983,238 entities linking ISBN-13 (P212), title (P1476), author (P50), publisher (P123)
International ISBN Agency - Global Publisher RegistersEdition-identifier ground truthRange rules per registration group and registrant287 registration groups; 1,871 registrant range rules; over a million publisher prefixes
Open Library Monthly Data DumpsTrade edition baseline: editions, works, authors, ratingsRecord-type partitioned files~12.4 GB compressed across record types (~29.6 GB with revision history)

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Publishing Global by construction - metadata contributed by publishers…

Crossref REST API - Scholarly Publishing Metadata

DOI · title · type …+13 more

Publishing Global research output across all countries and publishers…

OpenAlex API - Scholarly Book & Publisher Graph

doi · title · type …+21 more

Publishing Global author base across institutions worldwide

arXiv API - e-Print Publishing Preprints

title · summary · published …+2 more

Publishing Global - journals from publishers in more than 130 countries…

DOAJ API - Directory of Open Access Journals

Publishing Global and multilingual, with statement density varying by…

Wikidata - Publishing Knowledge Graph + SPARQL

P212 · P957 · P1476 …+7 more

Publishing Global - roughly 150 national and regional agencies across…

International ISBN Agency - Global Publisher Registers

field confirmed dictionary · deeper attributes fold out on request …+7 more

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

How many works does a Crossref DOI bulk export cover?

185,678,115 registered works were observable during the August 2026 research pass, spanning journal articles, book chapters, monographs, edited books, proceedings pieces, reports, dissertations and grants, with DOIs minted continuously since 2000. Beside the work rows sit rollups across roughly 169,622 journals, 33,393 member publishers and 45,739 funders.

What does one delivered row contain?

The DOI itself plus title, work type, publisher name and location, contributor lists with affiliations and roles, funding acknowledgements, publication dates, venue detail, ISBNs with type markers on book records, the deposited reference list with its count, an inbound citation count, and the registration timestamps that separate first registration from later indexing.

Can references and citations both be measured from the same rows?

Yes, and they are deliberately separate facts. `reference-count` and the deposited `reference` list describe what a work looked at - 66 references on the sampled chapter - while `is-referenced-by-count` describes what the field did with it. Outbound referencing and inbound impact stay distinct instead of blurring into one number.

Does every work record include an abstract?

No - abstracts are optional deposits made by publishers, so the field is populated unevenly by design rather than by defect, and much of what is present arrives JATS-tagged. Screen abstract presence against your named scope at sampling if abstract text drives your question.

Which datasets pair with the DOI spine for books and trade catalogs?

Three complements close the gap. OpenAlex widens to a 322-million-work citation graph including 5.9 million books and 22.0 million book chapters; Wikidata links editions, authors and publishers through ISBN-13 (P212), title (P1476), author (P50) and publisher (P123) across its 122,983,238-entity graph; and the ISBN agency registers validate edition identifiers against 287 registration groups and over a million publisher prefixes.