Glossary
Digital Object Identifier (DOI)
A Digital Object Identifier (DOI) is a persistent identifier registered for a scholarly work so that citations and metadata links keep resolving over time. Crossref exposes metadata for over 185,678,115 DOI-registered works through its free REST API, updated daily as deposits arrive.
What is a Digital Object Identifier (DOI)?
A DOI is a permanent string assigned to a scholarly work - a paper, dataset or book chapter - that stays valid even when the hosting site changes, which is why citation systems key on it. In the publishing slice the reference implementation is the Crossref REST API carrying scholarly publishing metadata for over 185,678,115 DOI-registered works. Its coverage note states that 'DOIs [have been] registered continuously since 2000 with backfile content, updated daily as deposits arrive', so the corpus grows incrementally rather than in editions. A DOI is not the only identifier a bibliographic record carries: Wikidata tags include 'isbn' alongside its book and publisher linking properties, so monographs typically hold both an ISBN for the physical object and a DOI for the work.
Why does DOI coverage matter when choosing a dataset?
The DOI is usually the join key that turns scattered publication records into one corpus, so gaps become silent record loss. A source whose metadata stops at some earlier deposit date misses everything added 'daily as deposits arrive' by Crossref, and snapshots taken months apart disagree for that reason alone. Backfile reach matters in the other direction: registration 'continuously since 2000 with backfile content' means pre-2000 works appear only through retrospective deposit, so older literatures are thinner by construction. Confusing DOIs with ISBNs creates a subtler defect - matching a book chapter against a print edition fails because Wikidata treats 'isbn' as a separate linking property.
How do you evaluate DOI data in a source?
- Scale against the registry. Crossref's metadata covers over 185,678,115 DOI-registered works; a bibliographic source well below that is a subset, and you should know which subset.
- Check the update path. Crossref updates 'daily as deposits arrive'; batch-refreshed mirrors lag by design.
- Probe the backfile. Registration runs continuously since 2000 with backfile content, so test how far pre-2000 works actually reach in your extract.
- Keep identifiers distinct. Treat 'isbn' as a separate Wikidata property rather than an alias for the DOI.
- Match delivery to workload. Full-corpus analysis favors a bulk data dump over per-record API calls.
Related terms
Entries that sit next to Digital Object Identifier in this glossary:
Frequently asked questions
How many DOI records does Crossref expose?
Metadata for over 185,678,115 DOI-registered scholarly works, served through Crossref's free REST API. Coverage includes DOIs registered continuously since 2000 plus backfile content, and the corpus is updated daily as new deposits arrive.
Is a DOI the same as an ISBN?
No. A DOI persistently identifies a scholarly work such as a paper or chapter, while an ISBN identifies a specific book edition. Wikidata carries 'isbn' as its own tagging property alongside book and publisher links, so the two identifiers should be stored and matched separately.
Datasets containing this field
Datasets containing Digital Object Identifier (DOI)
6 datasets carry digital object identifier (doi) in the catalog. Open one, count the fields, judge for yourself.
DOAJ API - Directory of Open Access Journals
bibjson.title · bibjson.alternative_title · bibjson.identifier …+19 more
Hugging Face Datasets Hub - Books Search
Every listing shows the field dictionary, sample rows, and coverage before you commit. API, files, or your warehouse. Daily, weekly, or hourly.
Get sample rows