OpenAlex — scholarly-works graph including education research
Datadory delivers openalex scholarly works graph including education research data covering 322 million works, 125.7 million authors, 134,448 institutions and 4,516 topics joined by citations, affiliations and classifications - 7.2 million of them in the Education subfield alone. Delivered daily, weekly, or hourly as files or straight into your warehouse.
What is the OpenAlex scholarly-works graph?
OpenAlex is OurResearch's index of world research, built as the successor to Microsoft Academic Graph, and the graph is the product: 322,044,938 works joined to 125,675,266 author records, 134,448 institutions, 255,627 sources, 4,516 topics, 10,708 publishers and 45,639 funders. Every work carries its DOI, title, publication date, genre and hosting source - and every one of those entities is a first-class record, not a footnote.
Three link structures turn it into an analysis surface rather than a bibliography. Authorships tie each work to its authors and their institutions, resolved to ROR identifiers and ISO country codes. Topics classify every work along a domain > field > subfield > topic hierarchy - 7,224,638 works sit in the Education subfield alone. And the citation graph is explicit: a cited_by_count on each work, referenced_works as the outgoing edge list.
Ask for the sample cut to your subfields, topics, institutions or countries; the shape you approve is the shape that ships.
What do sample rows look like?
Captured rows, exactly as they arrive - one work, plus the graph-scale counters and education rollups captured alongside it:
# one row per work - captured records, shown to fix the shape
id W3038568908
doi 10.1585/pfr.15.2402039
title Radiation Resistant Camera System for Monitoring
Deuterium Plasma Discharges in the Large Helical Device
publication_year 2020 type article
oa_status diamond country JP
first_author_orcid 0000-0003-0655-7347
# graph-scale counters captured alongside the work rows
all_works 322044938
authors 125675266
# education slices arrive pre-classified - captured rollups
Education (subfield, primary_topic.subfield 3304) 7224638 works
T12217 Higher Education Teaching and Evaluation 110439 worksNothing was tidied for presentation. The useful part is what surrounds each work row: affiliations already resolved to ROR institutions and country codes, an ORCID on the first author, an OA-status code, and topic assignments carrying IDs you can traverse. Because the education rollups ship pre-counted - 7,224,638 works in the Education subfield, 110,439 in Higher Education Teaching and Evaluation - scoping a feed to education research is a filter decision, not a classification project.
Samples ship cut to the subfields, topics and countries you name, with the same field set throughout.
What fields does each record carry?
The thirteen headline fields divide into identity, classification and graph, as tabled above. Three earn special mention.
authorships is where the institutional view comes from: each entry carries an ORCID, resolved institutions with ROR ids and country codes, the raw affiliation string and corresponding-author flags. Institution-level and country-level aggregations need no separate lookup file.
topics rides on OpenAlex's retired concepts layer - legacy concept codes still resolve, but topic filters are canonical, so anything built on the old codes should plan the migration once, on our side of the feed rather than repeatedly on yours.
referenced_works turns 322 million bibliography lines into a traversable citation graph. Paired with cited_by_count, it supports influence, co-citation and collaboration work without anyone opening a PDF.
Where does coverage run across geography, time and granularity?
- Geography: global. Affiliations resolve to ROR institution ids and ISO country codes covering essentially every country with research output, so region and country views aggregate mechanically.
- Temporal: works from the mid-1600s onward, densely covered from the 1990s - the era most education-research questions actually live in. Publication year and full ISO date both ride on each work.
- Granularity: one row per entity - works, authors, institutions, sources, topics, publishers, funders - with rollups delivered as typed aggregates rather than ad-hoc groupings.
The education depth is worth pinning down: 7,224,638 works in the Education subfield, about 41.5 million in the parent Social Sciences field, and 110,439 in the Higher Education Teaching and Evaluation topic alone, as tabled below. Set against the survey censuses on the education services hub - enrollments, completions, finance - this is the other half of the picture: published research about education, attributable to 134,448 named institutions.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
A graph arrives as tangles: nested authorships, scored topic arrays, edge lists tucked inside each work record. Teams rarely want the nest; they want the shapes underneath it - a works table keyed by ID, an authorship bridge joining works to authors to institutions, a topic dimension, and the citation edge list as its own table. Datadory does that unwinding once, keeps the keys stable across deliveries, and hands over whichever cadence the workload needs. The sample proves the shapes on your subfields before anything recurring starts.
Who uses this data, and for what?
- Research landscape mapping - who publishes what, with whom, across 134,448 institutions and essentially every country, filterable to a single topic.
- Edtech and publishing market sizing - education-research volume by topic, institution and country as demand-side denominators (market sizing playbook).
- Citation-grade evidence - literature reviews and systematic maps anchored to DOIs and explicit citation edges rather than vendor summaries (citation-grade research).
- Model features - citation graphs, topic labels and affiliation networks as training signal (ML model training).
- Trend detection - topic-level counts flag where education-research effort is moving before survey cycles pick it up.
- Institutional benchmarking - output volume and OA mix per named institution, assembled into benchmark sets on stable IDs.
Which personas get the most value?
Market researchers and consultants get demand-side denominators for anything sold into education research or academic publishing (market researchers page). Data scientists inherit a citation graph and topic taxonomy that arrive typed, keyed and already migrated (data scientists page). Competitive intelligence teams track research-output moves by institution, country and topic (competitive intelligence page). Investors and quant researchers read publishing volume and collaboration patterns as sector fundamentals (investors and quants). Journalists and academics cite DOIs and explicit counts rather than estimates (journalists and academics page). Sales and growth teams build account lists from named institutions active in their topic (sales and growth teams page).
Which notes and neighboring datasets pair with it?
Taxonomy note - the concepts layer is retired in favor of topics. Historical concept codes still resolve, but new feeds classify on topics; any scoping built on legacy codes gets migrated once at our end, not repeatedly at yours.
Identity note - raw affiliation strings ship alongside ROR-resolved institutions, because the unresolved text is sometimes the only version of a department that exists. Both stay in the feed.
Joining note - OpenAlex work IDs, DOIs, ORCIDs, ROR ids and ISSNs are the crosswalk set. They compose with the survey and statistical records below - IPEDS for output-per-campus views, Eurostat's education tables for national comparisons - instead of overlapping them. Where this graph meets the UIS education statistics comparator, the difference is unit of analysis: published works versus compiled national indicators.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
id | string | Canonical OpenAlex entity identifier - the stable work key every join starts from. | W3038568908 |
doi | string | DOI of the work where registered, the crosswalk to reference managers and publisher records. | 10.1585/pfr.15.2402039 |
display_name | text | Title of the work, or name of the entity when the row is an author, institution or source. | Radiation Resistant Camera System for Monitoring Deuterium Plasma Discharges... |
publication_year | integer | Year the work was published - the default time axis for trend work. | 2020 |
publication_date | date | Full ISO publication date where known. | 2020-06-08 |
type | enum | Work genre such as article, preprint, book-chapter or dissertation. | article |
primary_location | string | Best-known landing page, PDF location and hosting source (journal or repository) with ISSN. | landing page + host ISSN |
open_access | string | OA block: is_oa flag, oa_status (diamond, gold, green and friends) and any OA location. | oa_status: diamond |
authorships | text | Authors per work with ORCID, ROR-resolved institutions, country codes, raw affiliation strings and corresponding-author flags. | ORCID 0000-0003-0655-7347, country JP |
topics | text | Scored topic assignments carrying domain, field, subfield and topic display name with ID. | T12217 Higher Education Teaching and Evaluation |
cited_by_count | integer | Count of works in the graph that cite this work - the fast influence column. | — |
referenced_works | text | OpenAlex IDs of the works this one cites - the outgoing half of the citation graph. | edge list of cited work IDs |
indexed_in | text | Upstream indexes the record was harvested from. | crossref |
Education-relevant slices - captured rollup counts
| slice | level | works |
|---|---|---|
| Education (primary_topic.subfield 3304) | subfield | 7,224,638 |
| Social Sciences (field 33) | field | ≈41,500,000 |
| Higher Education Teaching and Evaluation (T12217) | topic | 110,439 |
Questions buyers ask
What does one row in this dataset represent?
Depends on the entity table. On the works spine, one row is one published work with its DOI, title, publication date, type, OA status and topic assignments. Authors, institutions, sources and topics each get their own row sets, joined back through authorship bridges and citation edges.
How many records does the graph hold?
Captured counters read 322,044,938 works and 125,675,266 author records, alongside 255,627 sources, 134,448 institutions, 4,516 topics, 10,708 publishers and 45,639 funders. Totals move as the index grows, so the sample reports current figures for your exact scope.
How much of it is education research?
7,224,638 works sit in the Education subfield and about 41.5 million in the parent Social Sciences field. Narrower still, the Higher Education Teaching and Evaluation topic holds 110,439 works. Feeds are scoped at whichever level of the domain-field-subfield-topic hierarchy your question lives at.
How far back does coverage run?
Works reach from the mid-1600s onward, with dense coverage from the 1990s forward. Most education-research analysis concentrates in the dense era, where citation edges and affiliation resolution are strongest, and publication year and full ISO date ride on every work.
What is a topic, and what happened to concepts?
Topics are the canonical classification: a four-level hierarchy running domain, field, subfield and topic, with scored assignments per work. The older concepts layer is retired - legacy codes still resolve, but feeds classify on topics, and that migration happens once on our side.
Can a sample be scoped to my subfields or institutions?
Yes. Name the subfields, topics, institutions or countries and the sample returns cut to that scope, with the field dictionary attached and work IDs intact. The shape you approve in the sample is the shape that ships on the recurring feed.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.