Publishing · OpenAlex
OpenAlex API - Scholarly Book & Publisher Graph
Datadory delivers openalex api scholarly book publisher graph data covering 322 million indexed scholarly works - 5.9 million books and 22 million book chapters among them - plus 10,708 publishers with corporate lineage fields, 255,627 sources and full citation networks, normalized into typed rows and delivered by API, files, or your warehouse daily, weekly, or hourly.
API, files, or your warehouse. Daily, weekly, or hourly.
- Where it covers
- Global research output across all countries and publishers - a 1988 French field report and a 2013 American clinical manual share one schema, so regional analyses never stitch in a second vendor
- How far back
- Historical through current: decades of monograph backfile sit beside recent arrivals, so decade-spanning citation cohorts and new-title watchlists come off one feed
- How fine
- Entity-level records - works, authors, sources, publishers, institutions - with citation-network edges between them, plus rollups per publisher and work type
What is the OpenAlex API - Scholarly Book & Publisher Graph dataset?
It is the breadth layer of scholarly publishing data, filed under Publishing and scored 9 out of 10 on Datadory's quality rubric against a catalog mean of 7.81 across 1,744 datasets. OpenAlex, compiled by a US 501(c)(3) nonprofit as the successor to Microsoft Academic Graph, indexes the world's research; the August 2026 research pass observed 322,044,938 works, including 5,906,748 books and 22,015,762 book chapters - 27,922,510 long-form scholarly units combined.
Two things make it analytic rather than merely searchable. First, citation impact travels on the row: cited_by_count, the field-weighted fwci index, year-normalized percentiles and annual counts_by_year, so a 1988 French field report and a 2013 clinical manual compete on measured terms. Second, the publisher entity is structural: hierarchy_level, parent_publisher and lineage turn a flat list of 10,708 publishers into a corporate map. Datadory normalizes all of it into typed, documented tables delivered on your cadence.
What do sample rows look like?
One row per indexed work, wrapped in index-wide counters:
id :
title : Diagnostic and Statistical Manual of Mental Disorders
type : book
publication_year : 2013
cited_by_count : 115
language : enRead what the pinned row proves. The id anchors every join and survives downstream renames; type: book places the row in the 25-value vocabulary before any classifier guesses; and cited_by_count: 115 measures a reference work's standing directly on the row.
Two more rows show the schema's honest range. The Australian century (1999) carries 18 citations; Rapport de mission au Rwanda 28. 1 au 06. 2. 1988, a French-language government field report, carries zero - and that zero is the point. Long-tail non-English monographs persist as first-class rows rather than being dropped, which is what makes cohort and coverage studies possible. Index-wide counters - 322,044,938 works, 5,906,748 books, 22,015,762 chapters, roughly 125.7 million authors, 255,627 sources, 10,708 publishers - travel beside the rows they summarize, so headline totals stay reconcilable with the data behind them. Example rows are illustrative of the record shape; your sample ships pinned to real records cut to your filters.
What fields does the dataset include?
Sixteen field groups define the work and publisher records, definitions verified during the August 2026 research pass rather than inferred from column names. Grouped as Datadory ships them:
- Identity:
idas the stable entity URI,doiwhere registered,title, andtypeacross the 25-value work-type vocabulary. - Dates:
publication_yearbeside the finerpublication_date. - Impact:
cited_by_count, the field-weightedfwci,cited_by_percentile_yearandcitation_normalized_percentilebands, and annualcounts_by_yearseries. - People:
authorshipswith positions, ORCID-linked author ids, institutions, countries and raw affiliation strings. - Hosting: the
open_accessstatus object andprimary_location/locationswith landing page, PDF, version and hosting source per location. - Classification:
topics,keywordsandprimary_topicfor subject filtering and trend slicing. - Container:
bibliovolume, issue and page range. - Graph:
referenced_worksandreferenced_works_count- outbound citation edges as columns. - Publishers:
works_countandcited_by_countaggregates plushierarchy_level,parent_publisherandlineagecorporate-structure fields.
Fields whose worked example ships as structure rather than literal values - the nested objects among them - confirm against live records when your sample is cut rather than being padded with guessed values here. The full dictionary follows in tabular form below.
Which fields arrive only on request?
Five extensions reshape the same graph rather than adding columns to the core. Author and institution resolution joins authorships onto the author and institution entities, so country and institution become filterable dimensions across roughly 125.7 million authors. Publisher lineage tables flatten the hierarchy into parent-child pairs, turning the 10,708-publisher list into an account map of who owns which imprint. Citation-edge extracts re-serve referenced_works as an edge list ready for network analysis. Topic and year rollups re-aggregate output and citation uptake over any window you name.
The most requested extension is the crosswalk: joining this corpus onto Crossref, arXiv and DOAJ by DOI, title and author - posted, registered, indexed and open-access-screened in one pipeline instead of four stitched exports. Each extension confirms against your named scope when the sample is cut.
What does coverage look like across geography, time and granularity?
Geography - global research output across all countries and publishers. Language is recorded per work, so the French field report in the sample above is queryable as itself, not noise around an Anglophone core.
Temporal - historical through current. Decades of monograph backfile sit beside recent arrivals in one schema, which is why a single feed serves both a thirty-year citation-cohort study and a new-title watchlist without changes on your side.
Granularity - entity-level throughout: works, authors, sources, publishers and institutions as separate record sets, joined by citation-network edges rather than fuzzy name matching. Rollups arrive per publisher and work type, so record-level science and market-share arithmetic share one dictionary. One honest caveat: book coverage depends on upstream harvests, so titles without registered identifiers may be underrepresented - screen your scope at sampling if completeness against a known list matters.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
You pick the channel and the cadence; extraction plumbing, pagination walks and change detection stay our problem. Rows arrive normalized to the field dictionary above and keyed on
id, so consecutive deliveries diff cleanly into new-work and citation-growth timelines instead of piling up as undifferentiated dumps.Hourly suits teams watching new arrivals or citation movement where latency matters. Daily suits competitive dashboards tracking publisher output through the lineage tree. Weekly suits trend work, where citation curves move slowly enough that point-in-time reads lose nothing. Whichever you pick, definitions travel unchanged and schema shifts are flagged rather than discovered. A sample cut to your named publishers, fields or years comes first either way.
Who uses this data, and for what?
Six jobs the record settles outright:
- Bibliometrics and impact scoring -
fwciand percentile bands put monographs on the same measured footing as articles; see citation-grade research. - Publisher competitive benchmarking - works and citation totals per publisher, resolved through the lineage tree, expose group-level output moves; see competitor tracking.
- Research-trend monitoring - topics, keywords and the book-versus-article mix read each discipline's direction early.
- Book-market and list curation - citation uptake per title gives editors a demand signal across 5.9 million books.
- ML training corpora - titled, classified, language-tagged long-form works with citation edges as columns; see ML model training.
- Systematic review automation - typed, dated, topic-filtered screening across millions of candidates in seconds.
Which personas get the most value?
Data scientists and ML engineers rank highest fit: 322 million nodes of citation edges, affiliation chains and topic labels are labeled structure raw dumps rarely provide; see data scientists' publishing use cases. Journalists, academics and students follow - cite the exact monograph, resolve its impact standing and trace the citation trail without hand-built bibliographies. Competitive intelligence and product teams get the lineage fields that survive corporate restructuring; see competitive intel product teams' publishing workflows. Market researchers and consultants size research output by country, field and work type; see market researchers' publishing workflows. Developers and builders ship discovery and reference products on one uniform entity schema; see developers builders' publishing use cases.
How does it compare to alternatives in publishing data?
Publishing's shelf splits by lifecycle stage, and this record owns the breadth-and-citations layer. Crossref holds the registration authority upstream - the registry aggregators build on - with deeper deposit-side commerce fields per work; this one answers when the question spans books, chapters and citation networks in a single frame. arXiv sits further upstream at the preprint instant. DOAJ curates the journal-business layer. On the trade side, Internet Archive answers where a readable copy exists, while this record answers what scholarship did with the book - the contrast is worked through properly in the Internet Archive vs OpenAlex comparison.
If the question is who published what, and who cites whom, across all of research - this is the record.
Why route it through Datadory
Because the raw artifact underneath is built for application lookups checking one entity at a time, and every interesting question wants millions of works, many publishers and many years in one frame. Datadory hands over the graph itself: the sixteen-group dictionary kept verified and current, publisher lineage resolved into its own table, citation edges served as columns ready for network analysis, and deliveries keyed on id so consecutive pulls diff into clean timelines - with sample rows in hand before any commitment. Browse the rest of the slice on the publishing data hub, see how the source fits the wider catalog on the OpenAlex source profile, or start from the best publishing datasets ranking.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| Field | Type | Definition | Example |
|---|---|---|---|
id | string | Canonical entity URI for the work - the stable key every join runs through. | |
doi | string | DOI URL of the work where one has been registered; absent for books that never received one. | - |
title / display_name | string | Work title as indexed. | Diagnostic and Statistical Manual of Mental Disorders |
publication_year / publication_date | date | Publication year and, where resolvable, full publication date of the work. | 2013 |
type | enum | Work type from a 25-value vocabulary; book is defined as a complete standalone volume such as a monograph, distinct from book-chapter, book-review, reference-entry, preprint and report. | book |
cited_by_count | integer | Number of citing works known to the index - the raw influence measure per row. | 115 |
fwci | number | Field-weighted citation impact index, so a theology monograph and a genetics paper compare on fair terms. | - |
cited_by_percentile_year / citation_normalized_percentile | number | Percentile bands placing each work's citations against its cohort and field-normalized baseline. | - |
counts_by_year | text | Annual citation series per work - velocity and decay curves without recomputation. | [{year:...,cited_by_count:...}, ...] |
open_access | text | Object carrying is_oa, oa_status, oa_url and any_repository_has_fulltext for the work. | {is_oa:..., oa_status:..., oa_url:...} |
authorships | text | Author list with position, author ids and ORCID, institutions and countries, plus raw affiliation strings. | [{author_position:..., institutions:[...], countries:[...]}] |
primary_location / locations | text | Where the work is hosted: landing page, PDF, hosting source, version and publication status per location. | {landing_page_url:..., source:{...}} |
biblio | text | Container detail - volume, issue, first and last page - for works published inside serials or volumes. | {volume:..., issue:..., first_page:..., last_page:...} |
referenced_works / referenced_works_count | text | Citation-graph edges from this work to other indexed works, plus the edge count. | ["", ...]; count alongside |
topics / keywords / primary_topic | text | Topical classification of the work - subject filters and trend slicing ride on these columns. | {display_name:..., field:{...}} |
language | string | ISO language of the work; non-English scholarship resolves here rather than vanishing. | en |
works_count / cited_by_count (publishers) | integer | Publisher-entity aggregates: total works and total citations attributed to the publisher. | - |
hierarchy_level / parent_publisher / lineage (publishers) | text | Corporate-structure fields linking imprints to parent groups - an ownership map, not a name list. | {parent_publisher:..., lineage:[...]} |
Coverage - geography, temporal range, granularity
| Dimension | Coverage |
|---|---|
| Geography | Global research output across all countries and publishers - a 1988 French field report and a 2013 American clinical manual share one schema, so regional analyses never stitch in a second vendor |
| Temporal | Historical through current: decades of monograph backfile sit beside recent arrivals, so decade-spanning citation cohorts and new-title watchlists come off one feed |
| Granularity | Entity-level records - works, authors, sources, publishers, institutions - with citation-network edges between them, plus rollups per publisher and work type |
Additional fields on request - extensions confirmed at sampling
| Extension | Notes |
|---|---|
| Author & institution resolution | Authorships joined onto the 125.7-million-author and institution entities, so country and institution become filterable dimensions |
| Publisher lineage tables | Hierarchy flattened into parent-child pairs across the 10,708 publishers - imprint-to-group maps without tree walking |
| Citation-edge extracts | referenced_works reshaped as an edge list ready for network analysis over any scope you name |
| Topic & year rollups | Output and citation uptake re-aggregated by topic, work type and year window |
| Scholarly crosswalk | Joins onto Crossref, arXiv and DOAJ by DOI, title and author |
What teams do with it
- Bibliometrics and impact scoring `fwci`, year percentiles and `counts_by_year` make monograph impact comparable across fields and decades - defensible metrics instead of folklore.
- Publisher competitive benchmarking Works and citation totals per publisher, split by the lineage tree, turn 10,708 rivals' output into a table you can diff week over week.
- Research-trend monitoring Topics, keywords and work types read each discipline's direction years before surveys catch up - books lag articles, and the gap is itself a signal.
- Book-market and list curation 5.9 million books with citation uptake per title give acquisition editors and imprint planners a demand signal that sales figures alone cannot.
- ML training corpora Titled, classified, language-tagged long-form text with citation edges as columns - labeled structure few corpora offer at this scale.
- Systematic review automation Fielded filtering across type, year, language and topic collapses screening passes over millions of candidate works into queries that run in seconds.
Questions buyers ask
What does openalex api scholarly book publisher graph data contain?
Entity-level records for the world's research: 322,044,938 works observed in August 2026 - including 5,906,748 books and 22,015,762 book chapters - plus roughly 125.7 million authors, 255,627 sources and 10,708 publishers. Works carry citation counts, field-weighted impact, authorships with institutions and countries, topics, hosting locations and referenced-work edges.
How many books and book chapters does the dataset cover?
Live testing during the August 2026 research pass counted 5,906,748 works typed as complete standalone books and 22,015,762 book chapters - 27,922,510 long-form units together, inside a 322-million-work default index. The 25-value type vocabulary keeps books separate from chapters, reviews and reference entries, so monograph studies never inherit chapter noise.
Can I measure the impact of a book, not just an article?
Yes - that is what the impact fields exist for. cited_by_count gives raw citations (115 on the sampled DSM edition), fwci field-weights them against the discipline, cited_by_percentile_year places the work in its cohort, and counts_by_year traces the curve. A slow-burning monograph stops looking weak next to fast-citing article fields.
Does the dataset capture publisher ownership structures?
The publisher entity carries hierarchy_level, parent_publisher and lineage fields linking each imprint to its corporate group, beside works_count and cited_by_count aggregates. Across the 10,708 publishers indexed, market-share arithmetic resolves through the actual org chart rather than guessing from name strings that rebrand every few years.
Is non-English and older scholarship covered?
Yes. Coverage is global across all countries and runs historical through current, and each work records its ISO language. The sampled rows themselves span a 1988 French government field report with zero citations and a 1999 Australian political history with eighteen - long-tail non-Anglophone monographs persist as first-class rows rather than dropping out.
Can a sample be scoped to my fields, publishers or years?
Yes. Name the work types, topics, languages, publishers or date windows you care about and the sample returns pinned to real records in exactly the field shape shown above. Author and institution resolution, publisher lineage tables, citation-edge extracts and the crosswalk onto Crossref, arXiv and DOAJ confirm alongside the sample rather than being promised blind.
Notes on this record
- Provenance Compiled during the August 2026 research pass against live index output; the 322,044,938-work figure, the book and chapter splits and the 10,708-publisher count are point-in-time readings of that pass, not vendor claims.
- One key, every join `id` keys every entity and `doi` resolves outward onto Crossref and DOAJ records, so the scholarly hemispheres meet on real columns instead of fuzzy title matching.
- Books are not chapters `type` separates `book` - a complete standalone volume such as a monograph - from `book-chapter`, `book-review` and `reference-entry` across a 25-value vocabulary; conflating them ruins monograph impact studies.
- Impact, field-adjusted `fwci` and the year-percentile bands put a 1988 French field report and a 2013 clinical manual on one scale; raw `cited_by_count` alone would just crown whichever discipline cites most.
- Scored against the catalog Quality 9 of 10 against a catalog mean of 7.81 across Datadory's 1,744 cataloged datasets; the sixteen-record publishing slice averages 8.5 and this record sits in its top tier.
- Sample policy Samples ship in the exact schema shown above, cut to your named work types, topics, publishers and windows; lineage tables, edge extracts and crosswalk joins confirm with the sample.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.