Publishing · Wikidata
Wikidata - Publishing Knowledge Graph + SPARQL
Datadory delivers wikidata publishing knowledge graph sparql data covering the 122,983,238-entity knowledge base that links books, authors and publishers as typed statements - ISBN-13 P212, publisher P123, title P1476, author P50, publication date P577 - with multilingual labels riding beside every item, delivered by API, files, or your warehouse daily, weekly, or hourly.
API, files, or your warehouse. Daily, weekly, or hourly.
- Where it covers
- Global and multilingual, with statement density varying by language community - regional reads resolve inside one graph instead of stitching country feeds, once fill rates are screened for your scope
- How far back
- Works from early printing to current titles, with editors and bots adding statements continuously - a printing-history cohort and this quarter's new editions share one dictionary
- How fine
- Entity-level throughout: editions, works, authors and publishers connected by typed statement properties, with rollups available along any edge - per publisher, per author, per class, per language
What is the Wikidata - Publishing Knowledge Graph + SPARQL?
It is the connective tissue of book metadata, filed under Publishing and scoring 9 out of 10 on Datadory's quality rubric - inside the top tier of the sixteen-record publishing slice we catalog. The knowledge base held 122,983,238 entities when the August 2026 research pass read its counter, and book publishing lives inside it as structured statements on item entities: ISBN-13 (P212) and ISBN-10 (P957) pin the edition, title (P1476) arrives language-tagged, publisher (P123) and author (P50) resolve to their own nodes, publication date (P577) grades to year-month-day precision, and P291, P31 and P921 carry place of publication, work class and subject.
Two things make it analytic rather than merely encyclopedic. First, the links are the product: publisher-book-author triples across languages turn a pile of titles into a network you can walk, which is why publishers' houses, imprint relationships and multilingual editions of one work come out as queries rather than reconciliation projects. Second, the verification was live: the research pass confirmed publisher, title, ISBN-13 and publication-date claims on a real book item, so the dictionary below describes observed statements, not documentation aspiration. Get a sample of this dataset to inspect real rows before committing pipeline time.
What do sample rows look like?
One row per book-bearing item, wrapped in the graph's own identifiers:
# one edition row - delivered grain is one row per book-bearing item entity
item : Q100708395
title (P1476) : Silver Economies, Monetisation and Society in Scandinavia, AD 800-1100
isbn13 (P212) : 978-87-7934-585-0
publication_date (P577) : 2011-01-01
# a second edition shows the same shape carrying another language market
item : Q100708549
title (P1476) : Dictionnaire historique des Academiciens de Lyon
isbn13 (P212) : 978-2-9559433-0-4
publication_date (P577) : 2017-03-01
# every row resolves outward through typed statements, not loose text
publisher (P123) : -> publishing-house item
author (P50) : -> author item
instance_of (P31) : -> book / literary work class
# graph-wide scale accompanies the rows
entities.total : 122983238 # counter read August 2026Read what the pinned row proves. The item Q-number anchors every downstream join and never gets recycled; the ISBN-13 ties the row back to the trade catalogs that speak barcode; the language-tagged title means a Portuguese edition and an English edition of one work stay distinct rows that still point at shared author and publisher nodes. The two count lines matter less than the arrow lines beneath them - P123, P50 and P31 arrive as item links, so attribution is a graph fact with its own node, ranks and references in the claims container rather than a free-text afterthought. Graph-wide scale - 122,983,238 entities at the August 2026 reading - travels beside the rows it summarizes. Example rows are illustrative of the record shape; your sample ships pinned to real records cut to your filters.
Which fields does the dataset include?
Eleven verified fields define the item record, every definition checked against live output during the August 2026 research pass rather than inferred from property names. They group cleanly: identity (P212, P957, P1476), people and firms (P123, P50), time and place (P577, P291), classification and topic (P31, P921), and the multilingual apparatus (labels / descriptions / aliases) with the claims container that keeps ranks and references attached to each assertion.
The full dictionary follows in the table below. Note the deliberate asymmetry with ordinary bibliographic feeds: half the dictionary is pointers. Publisher and author are nodes you traverse, which is precisely what makes imprint networks and multi-market edition maps computable instead of hand-built.
Which fields arrive only on request?
The extensions above reshape the same graph rather than invent new facts, and they confirm against your named scope when the sample is cut:
What does coverage look like across geography, time and granularity?
Geography - global and multilingual, with one honest structural note: coverage density varies strongly by language community, because statements accumulate where contributor communities are active. A title hot in one market may have thinner statements than its bestseller cousin elsewhere - measurable per scope, and worth screening at sampling rather than assuming.
Temporal - works from early printing to current titles, and the graph keeps moving: editors and bots add statements continuously, so the corpus is a live fabric rather than a frozen snapshot. That combination serves a printing-history study and a this-quarter new-edition watch from the same dictionary.
Granularity - entity-level throughout: book editions, works, authors and publishers, connected by typed statement properties. Rollups come naturally along any edge - per publisher, per author, per class, per language - so record-level cataloging and market-share arithmetic share one spine. Property presence is uneven by design across real-world books; fill rates for P123, P50 and P212 within your named scope are a number we measure together at sampling instead of a promise made blind here.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
You pick the channel and the cadence; endpoint etiquette, pagination walks and change detection stay our problem. Rows arrive normalized to the field dictionary above and keyed on the item Q-number with ISBN-13 alongside, so consecutive deliveries diff cleanly into new-edition arrivals and statement corrections instead of piling up as undifferentiated snapshots.
Hourly suits teams watching new statements land on titles they track. Daily suits catalog enrichment and deduplication jobs that run between warehouse loads. Weekly suits network analysis, where publisher graphs move slowly enough that snapshots lose nothing. Whichever you pick, definitions travel unchanged and schema shifts are flagged rather than discovered. A sample cut to your named publishers, languages, subjects or date windows comes first either way.
Who uses this data, and for what?
Five uses dominate, and all five lean on the same trick - treating publishing facts as a graph instead of a spreadsheet:
Which personas get the most value?
Fit tracks how much of your question lives in the relationships between books, people and firms:
How does it compare to alternatives in publishing data?
Publishing's shelf splits by job, and this record owns the linkage layer. OpenAlex is the breadth play - roughly 322 million works including 5.9 million books - when citation volume matters more than cross-market identity. Open Library Search & Works/Editions API answers which editions of a book exist inside a lending catalog, while this one answers how that book, its author and its publisher hang together across every language. The International ISBN Agency registers supply the prefix arithmetic - who may mint which ranges - but carry no per-title statements; Google Books and Goodreads hold the retail-facing descriptions and ratings. Our twin record, the Wikidata SPARQL Query Service, sees the same graph from the query-seat angle - the trade-offs are worked through in the dedicated comparison.
If the question is what connects these books, authors and publishers across markets, this is the record.
Why route it through Datadory
Because the raw artifact underneath is a living encyclopedia edited by humans and bots, and every interesting question wants thousands of editions, many publishers and several languages in one frame. Datadory hands over the graph itself: the eleven-field dictionary kept verified and current as editing practice evolves, deliveries keyed on the item Q-number with ISBN-13 alongside so consecutive pulls diff into clean timelines, label and alias tables resolved into their own columns, and the whole thing landed on the cadence your pipeline runs - with sample rows in hand before any commitment. Browse the rest of the slice on the publishing data hub, start from the best publishing datasets ranking, or see how the slice hangs together in the publishing data guide.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| Field | Type | Definition | Example |
|---|---|---|---|
P212 | string | ISBN-13 identifier of a book edition - the join key that carries the item into commercial catalogs and both ISBN registries. | 978-87-7934-585-0 |
P957 | string | ISBN-10 identifier of a book edition, surviving on older printings where the thirteen-digit form never replaced it. | - |
P123 | string | Publisher - an item link to the publishing house entity, so attribution resolves to a node in the graph rather than a spelling variant. | -> publishing-house item |
P577 | date | Publication date of the work or edition, carried as a precision-graded date value. | 2011-01-01 |
P1476 | text | Title of the work as monolingual, language-tagged text - one title string per language edition rather than one mushed column. | Silver Economies, Monetisation and Society in Scandinavia, AD 800-1100 |
P50 | string | Author - an item link to the author entity, keeping homonymous writers apart by node instead of name-matching heuristics. | -> author item |
P291 | string | Place of publication observed on book items - territory signal for rights and distribution reads. | - |
P31 | enum | Instance of - the classification edge placing each entity in a class such as book or literary work, the spine for corpus composition counts. | book |
P921 | string | Main subject - topical links from a work to subject entities, turning a catalog into a topic map. | -> subject item |
labels / descriptions / aliases | text | Multilingual labels, descriptions and aliases attached to every item, so the same work surfaces correctly across language communities. | 'Silver Economies...' @en |
claims | text | Statement container holding every property-value assertion with its rank and references - provenance travels with the fact. | [{property, value, rank, references}] |
Fields and resolutions available on request
| Extension |
|---|
| Label, description and alias tables delivered as their own multilingual columns - one row per language per item instead of nested blobs |
| Author and publisher nodes resolved into their own tables, turning attribution pointers into an addressable directory of the industry |
| External-identifier columns observed on book items - library-system and reader-community IDs sitting beside the core properties for cross-source reconciliation |
| Crosswalk joins onto Open Library editions and both ISBN registers by ISBN-13 and title, chaining catalogs into one pipeline |
What teams do with it
- Publisher network analysis Publisher-book-author triples across languages make imprint relationships and house concentration a traversable graph - rivals mapped as nodes and edges, not guessed from press releases.
- Catalog enrichment and deduplication Stable item Q-numbers joined on ISBN-13 collapse duplicate editions in your catalog and backfill publisher, subject and multilingual titles where your rows run thin.
- Knowledge-graph features for machine learning Typed statements, classes and topic links give entity embeddings and link-prediction models labeled structure at graph scale - edges as columns, not scrapes.
- Multilingual edition mapping Language-tagged titles plus shared author and publisher nodes reveal which works travel across markets and which stall at the border - sizing reads built on identity rather than title-string luck.
- Metadata quality auditing Fill rates for publisher, author and ISBN assertions measured per scope turn 'is our reference data trustworthy?' from an opinion into a dashboard.
Questions buyers ask
What does wikidata publishing knowledge graph sparql data contain?
One row per book-bearing item in a 122,983,238-entity knowledge graph, observed August 2026. Each row carries the edition's ISBN-13 and ISBN-10, language-tagged title, publisher and author as resolvable item links, publication date, place of publication, work class and subject links, plus multilingual labels, descriptions and aliases with ranks and references attached in the claims container.
How many entities does the dataset cover?
The counter read 122,983,238 data entities during the August 2026 research pass, spanning books, works, authors, publishers and everything else the community models. The book-bearing subset is not published separately, which is exactly why we screen fill rates against your named scope at sampling instead of quoting a headline edition count.
Is every book in the graph fully described?
No - and treating that as a defect would misread the source. Statement presence varies by language community and by title, so many real-world books lack publisher, author or ISBN-13 assertions entirely. We measure P123, P50 and P212 fill rates inside your named scope during sampling, so your pipeline inherits measured completeness rather than assumed completeness.
Can these rows link to my own catalog?
Yes, on two keys that behave. The item Q-number is stable and unique inside the graph, and ISBN-13 rides on the edition rows, so your catalog joins on a barcode it already stores while unresolved titles fall back to language-tagged title and author-node matching. The same keys reach Open Library editions and both ISBN registers, chaining catalogs into one pipeline.
Can a sample be scoped to my publishers, languages or years?
Yes. Name the publishers, languages, subjects, classes or date windows you care about and the sample returns in exactly the eleven-field shape above, pinned to real items from the graph. Label and alias tables, statement fill-rate readings for your scope and crosswalk joins onto the neighboring publishing records confirm alongside the sample rather than being promised blind.
Notes on this record
- Provenance Compiled during the August 2026 research pass against live graph output; the 122,983,238-entity figure is a point-in-time reading of the public counter, and the publisher, title, ISBN-13 and publication-date claims on the sampled item were verified live, not taken on faith.
- One key, every join The item Q-number keys the rows and resolves outward, while ISBN-13 (`P212`) lands on both ISBN registers and Open Library editions - so the trade catalogs and the scholarly aggregators meet this graph on one column instead of fuzzy matching.
- Statements, not strings Publisher and author arrive as item links with ranks and references inside the `claims` container, so attribution stays a graph fact you can traverse rather than a free-text afterthought that breaks on the second spelling variant.
- Fill rates are a measurement, not a promise Statement presence varies sharply across language communities and titles. We read `P123`, `P50` and `P212` coverage inside your named scope during sampling, so completeness enters your pipeline as a number.
- Scored against the catalog Quality 9 of 10 - top tier of the sixteen-record publishing slice, whose average sits at 8.5 against a mean of 7.81 across Datadory's 1,744 cataloged datasets.
- Sample policy Samples ship in the exact schema shown above, cut to your named publishers, languages, subjects and windows; label and alias tables, fill-rate readings and crosswalk joins confirm alongside the sample.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.