Publishing Data Provider - Scholarly and Trade-Book Datasets, Delivered · Head-to-head
arXiv API - e-Print Publishing Preprints vs DOAJ API - Directory of Open Access Journals
Which publishing data provider - scholarly and trade-book datasets, delivered data fits your job: arXiv API - e-Print Publishing Preprints, or DOAJ API - Directory of Open Access Journals. API, files, or your warehouse. Daily, weekly, or hourly.
arXiv API - e-Print Publishing Preprints
DOAJ API - Directory of Open Access Journals
Coverage, side by side
| arXiv API - e-Print Publishing Preprints | DOAJ API - Directory of Open Access Journals | |
|---|---|---|
| Geographic | Global author base across institutions worldwide | Publishers in more than 130 countries, each coded to ISO-3166 |
| Temporal | Submissions from 1991 to the present; earliest consolidated datestamp September 16, 2005 | OA start year per journal back to inception; article metadata current to today |
| Granularity | One paper version, with version history implicit in identifier suffixes | Two levels - the curated journal record and its article records |
What each contains
They tie on 1 attribute. Pick by fit, not by loyalty.
| arXiv API - e-Print Publishing Preprints | DOAJ API - Directory of Open Access Journals | |
|---|---|---|
| Publisher | arXiv | Directory of Open Access Journals (DOAJ) |
| Subject lens | Papers researchers post before peer review - physics, mathematics, computer science, quantitative biology, finance and economics | Curated journal directory: journal profiles plus the articles those journals publish |
| Unit of analysis | One paper version, with version history implicit in identifier suffixes | Two levels - the curated journal record and its article records |
| Geographic coverage | Global author base across institutions worldwide | Publishers in more than 130 countries, each coded to ISO-3166 |
| Temporal reach | Submissions from 1991 to the present; earliest consolidated datestamp September 16, 2005 | OA start year per journal back to inception; article metadata current to today |
| Scale | Millions of e-prints commonly cited above 2.5 million; one common-term probe alone matched 185,192 | 23,352 journals carrying 13,470,084 article records; the journal table spans 52 columns |
| Documented fields | 12 | 24 |
| Datadory rubric | 9/10, field definitions verified | 9/10, field definitions verified |
| Best for | Monitoring research output before peer review and tracing preprints into print | Mapping the open-access landscape by country, discipline, review model and turnaround |
What each does better
arXiv API - e-Print Publishing Preprints
Version-level depth on every paper. One record per paper version: title, the summary abstract, an author/name list with optional arxiv:affiliation, a published stamp for the version-1 submission and updated for later revisions. DOAJ's journal records know a title opened for submissions in a given year; they cannot see draft three of anybody's paper.
A preprint-to-publication trail, written by the authors themselves. arxiv:journal_ref carries the eventual journal reference, arxiv:doi the registered published version, and arxiv:comment often names the acceptance venue and page count. A worked example from the sample rows: Entropic Measures of Complexity in a New Medical Coding System, posted December 23, 2020 under primary category cs.IT with a cs.DL cross-list.
Search grammar built for literature work. Prefixes ti, au, abs, co, jr, cat and rn narrow on title, author, abstract, comment, journal reference, category or report number, combinable with boolean AND/OR/ANDNOT, parentheses, quoted phrases and date ranges. Match-count counters report the true total before you commit - one common-term probe reported 185,192 hits.
STEM breadth at preprint speed. Millions of e-prints commonly cited above 2.5 million, spanning physics, mathematics, computer science, quantitative biology, finance and economics, from authors at institutions worldwide.
DOAJ API - Directory of Open Access Journals
Journal-level business fields no preprint index carries. Publisher name and ISO-3166 country, print and electronic ISSNs, accepted manuscript languages, reuse-condition type with attribute flags, article-processing-charge presence and amount, digital-preservation participation, persistent-identifier usage and average submission-to-publication time in weeks. That is the raw material for charge benchmarks, policy audits and market maps - none of it exists in a preprint corpus.
Curation signals you can filter on. Every journal carries an admin.ticked re-application status and an admin.last_full_review date, alongside peer-review process types that separate double-anonymous review from editorial screening, plus a BOAI compliance flag. Sample record: Computational and Experimental Research in Materials and Renewable Energy - eISSN 2747-173X, publisher country ID, open since 2018, ten weeks median submission-to-publication, ticked.
Two record levels, one dictionary. The 23,352 journal records resolve down into 13,470,084 article records carrying DOI, journal linkage, authors, abstracts and page numbers, so a journal-level screen lands on citable papers without a second source.
Market sizing keyed by geography. With publisher country coded to ISO-3166 across more than 130 countries, a country-by-country view of the open-access landscape is a group-by, not a research project.
Where they're equivalent
Same shelf, same grade. Both sit in Publishing's 16-primary slice, both score 9/10 against the catalog's 7.81 average, and both cleared field-definition verification during the same research pass in August 2026 - a bar met by 85.7% of the catalog.
A shared bibliographic spine. Titled work, human names, a time stamp and a classification scheme appear in both dictionaries. That spine is why a title-year-author match bridges the two, and why neither replaces the other past it.
Stable keys everywhere. Version-suffixed entry identifiers and resolved DOIs on one side; ISSN pairs and article DOIs on the other. Either anchors a master list; joined, they chain preprint to journal to article.
Structured records with documented examples. Both dictionaries publish definitions and example values per field, so a sample cut tells you within minutes whether the grain fits your model.
The verdict
Verdict: sample both, pick by fit - the unit of analysis decides.
Take arXiv when the unit is the paper and its life story: monitoring fresh STEM literature, counting category-level submission growth as an innovation signal, watching rival organizations' research output, or following a preprint into print through its journal reference and registered DOI.
Take DOAJ when the unit is the journal and its business: charting the open-access landscape by country and discipline, benchmarking article processing charges, vetting a venue's review model and turnaround before submitting, or auditing preservation coverage across a publisher list.
Three quick tests settle most cases. Need the paper before peer review? Only arXiv has it. Need the journal's economics and vetting trail? Only DOAJ carries those fields. Modeling scholarly communication end to end, from first posting to published home? That is the both-of-them case, and it is common.
Sample both, pick by fit. See arXiv API - e-Print Publishing Preprints · See DOAJ API - Directory of Open Access Journals
Or take both in one feed
They stack along the publishing timeline, and the join is a mapping exercise rather than a miracle. Monitor on arXiv - new postings in a category, an organization's output, a topic's trajectory - resolve what landed in print using arxiv:journal_ref and arxiv:doi, then enrich each hit with the journal's DOAJ profile: review model, turnaround in weeks, publisher country, preservation status. A science-policy team can measure how fast computer-science preprints find vetted open homes; a competitive-intelligence team can pair rivals' research output with the journals willing to run it.
Two practical notes from the records. There is no shared key, so the bridge is a title-year-author match - imperfect, though arxiv:doi closes most of the gap wherever a published version registered one. And mind the grain: one side counts paper versions, the other counts journals and articles, so aggregate deliberately before comparing magnitudes.
Or take both in one feed. Datadory normalizes each record to its documented dictionary, attaches sample rows for validation, and ships them beside the rest of the publishing catalog - delivered daily, weekly, or hourly, your call.
API, files, or your warehouse. Daily, weekly, or hourly.
Fair questions
Is arXiv better than DOAJ for publishing analytics?
Different instruments, tied on craft - both score 9/10. arXiv wins whenever the question concerns papers: millions of preprints with versions, abstracts, categories, affiliations and author-supplied journal references. DOAJ wins whenever the question concerns journals: 23,352 curated titles with publisher country, review models, turnaround times and article-processing-charge data across 13,470,084 articles.
Do arXiv and DOAJ cover the same papers?
Only in passing. arXiv holds work at the pre-review stage; DOAJ's article records describe formally published papers in vetted journals. Overlap happens where a preprint later reached print - visible as a journal reference or registered DOI on the arXiv side, and as a journal-linked article record on the DOAJ side.
Which dataset is bigger, arXiv or DOAJ?
On different axes. The arXiv corpus is commonly cited above 2.5 million e-prints, and a single common-term probe matched 185,192 records. DOAJ counts 23,352 journal records and 13,470,084 article records, so it dwarfs arXiv at article level while covering far fewer titles.
Can you track a preprint into its published journal version?
Yes, wherever authors supplied the trail. Three arXiv fields exist for exactly this: `arxiv:journal_ref` for the eventual journal, `arxiv:doi` for the registered published version and `arxiv:comment`, which often names the acceptance venue. From there DOAJ completes the picture with the journal's review model, ISSN pair and turnaround statistics.
How current can arXiv and DOAJ data be?
As current as your project needs: Datadory delivers either dataset daily, weekly, or hourly - your call. Content-wise, arXiv's corpus runs from 1991 submissions to the present, while DOAJ stamps every journal record with creation and last-update dates you can filter on, observed changing as recently as January 2026.
Can Datadory deliver both datasets together?
Yes. Sample both and pick by fit, or take both in one feed - normalized to their documented dictionaries (12 fields on the preprint side, 24 on the journal side), aligned on whichever identifiers you choose, and validated against sample rows before anything ships.