Publishing · Wikidata Query Service
Wikidata SPARQL Query Service Data
Datadory delivers wikidata sparql query service data covering typed result rows drawn from a 122,983,238-entity knowledge graph: each row one query binding carrying the edition entity, its ISBN-13 (P212), language-tagged title (P1476), publisher reference (P123) and publication date (P577) - shaped to your question and delivered by API, files, or your warehouse daily, weekly, or hourly.
API, files, or your warehouse. Daily, weekly, or hourly.
What is the Wikidata SPARQL Query Service dataset?
It is the window onto Wikidata's knowledge graph - the seat where questions become tables. Datadory catalogs two publishing datasets from Wikidata: the twin record holds the shelf itself, 122,983,238 entities of editions, works, authors and publishers connected by typed statements, while this one covers the query service that reads that graph. The service speaks SPARQL 1.1 over the Wikibase RDF data model: you declare the shape - which editions, which properties, which markets - and it returns result sets assembled to match.
The seat is unusually well appointed. A visual Query Builder and an interactive Query Helper assemble triple patterns without hand-written syntax, a curated examples collection supplies ready-made starts, and result views cover table, map, graph, timeline and image grid - a publisher-location pattern renders as pins before any code exists. Every one of those capabilities was confirmed live during the August 2026 research pass, alongside a probe joining ISBN-13 (P212), title (P1476) and publication date (P577) that came back with language-tagged titles and typed date literals.
Datadory takes the machinery out of the browser and hands you the result sets as flat, typed rows. Context sits on the publishing data hub; get a sample of this dataset cut to your publishers, languages or windows.
What do sample rows look like?
Two captured solutions beat a specification paragraph:
# one solution row - delivered grain is one row per query binding
book : Q100708395
title (P1476) : Silver Economies, Monetisation and Society in Scandinavia,
AD 800-1100 [language tag: en]
isbn13 (P212) : 978-87-7934-585-0
pubdate (P577) : 2011-01-01T00:00:00Z # xsd:dateTime
# a second row shows the same shape carrying another language market
book : Q100708731
title (P1476) : Ovo gigante da Mothra : Consumo do Pacifico Sul no Japao
dos anos 1960 [language tag: pt]
isbn13 (P212) : 978-85-98112-63-3
# the response envelope travels beside every result set
head.vars : book | title | isbn13 | pubdate
results.bindings : one solution object per row, each value typed uri or literal
# the graph underneath the rows
entities.total : 122983238 # counter read August 2026Read the anatomy rather than the antiquity of the first title. First, the grain is honest: one row per binding, so a publisher-year count and an edition listing differ only in the question, never in the shape. Second, identity arrives as entity references - Q-numbers that resolve to publisher and class nodes - rather than loose strings, which is what keeps a thousand-row pull joinable instead of fuzzy. Third, the typing is load-bearing: dates land as xsd:dateTime, titles carry their language tag, and a Portuguese-market edition reads natively beside an English one. Your scope ships with the sample in exactly this layout.
Which fields does the dataset include?
Seven documented fields in two layers, definitions verified against live captures in August 2026 - a bar met by 85.7% of the 1,744 datasets Datadory catalogs.
The envelope layer makes every response self-describing: head.vars names the columns, results.bindings carries one solution object per row with each value typed as URI reference or literal. Bind against that contract once and the shape survives however the query changes.
The solution layer carries the publishing payload: ?book addresses the edition entity by Q-number, ?title (P1476) arrives monolingual and language-tagged, ?isbn13 (P212) keys the join onto trade catalogs, ?publisher (P123) points at the publishing-house item, and ?pubdate (P577) preserves its recorded precision. Publisher-label resolution and rollup columns fold in on request rather than cluttering the core.
What does coverage look like across geography, time and granularity?
Geography - global and multilingual, mirroring the graph itself. Language-tagged titles mean regional reads are native rather than derived, with statement density varying by language community; screen fill rates for any scope before leaning on completeness.
Temporal - current-state coverage of a continuously edited graph, not a dated snapshot. Each row carries its own publication date (2011-01-01T00:00:00Z on the captured row), so a historical cohort study and this quarter's new-edition watch share one dictionary. Scheduled captures add the history dimension the current state lacks.
Granularity - triple-level statements aggregated into user-defined result sets per query: one row per binding, shaped by the question. Result size is bounded by per-query limits rather than by the dataset, so large analytical joins stage in pieces - handled at scoping, not discovered mid-project.
Scored honestly: 9 of 10 on Datadory's rubric against a 7.81 mean across all 1,744 cataloged datasets, sitting just under the sixteen-record publishing slice's 8.5 average - held up by the verified response contract, docked because property-level fill rates vary and need measuring per scope rather than being promised blanket.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
You pick the channel and the cadence; query construction, staging under per-query limits and publisher-name resolution stay our problem. Result sets arrive flattened to the field dictionary above - one typed row per binding, language tags intact - so consecutive deliveries diff cleanly into new editions and corrected statements instead of piling up as undifferentiated snapshots.
Hourly suits teams watching statements land on titles they track. Daily suits catalog-enrichment and audit jobs that run between warehouse loads. Weekly suits network work, where publisher graphs move slowly enough that snapshots lose nothing. A sample cut to your named publishers, languages, subjects or date windows comes first either way.
Who uses this data, and for what?
- Publisher output measurement - editions per imprint per year, grouped straight from bindings on publisher and publication date; one of 367 cataloged datasets serving market sizing.
- Rival catalog watching - competitors' edition arrivals and identifier hygiene diffed week over week; one of 195 tagged for competitor tracking.
- Missing-identifier audits - the editions lacking ISBN-13, publisher or date statements, listed outright; data-quality arguments become work queues with a length.
- Multilingual market reads - language-tagged titles separate markets natively, no translation layer in between.
- Citation-grade metadata pulls - typed, sourced edition values that survive review; the catalog holds 765 datasets clearing that bar across 133 industries, and this one resolves editions across languages inside it (citation-grade research).
Which personas get the most value?
Data scientists and ML engineers land closest fit: prototype a join interactively, then industrialize the extract - typed bindings drop into pipelines without a bespoke parser; see data scientists x publishing. Developers and builders embed lookup, deduplication and audit features behind one stable envelope via developers & builders view. Market researchers and consultants build publisher-by-year output curves and multilingual footprint maps through market researchers view. Competitive intelligence teams run rivals' edition counts as a standing feed via competitive intel view. Journalists, academics and students verify bibliographic detail across languages with typed values attached to the editions they describe - journalists & academics view.
How does it compare to alternatives within publishing data?
This record owns the question-answering layer of the Wikidata pair. The twin Wikidata - Publishing Knowledge Graph + SPARQL is the shelf: entity-level rows for enrichment, validation and network loads. The dedicated comparison walks the split - if the deliverable is a dataset, take the shelf; if it is an answer, take the window; most teams take both in one feed.
Elsewhere on publishing's shelf: Crossref REST API is the DOI registration layer with 185,678,115 works; OpenAlex API is the breadth play at roughly 322 million works including 5.9 million books; the International ISBN Agency registers cover numbering-group plumbing without per-title statements. None of those return arbitrarily shaped result sets over a multilingual graph. When the question is who published what, in which language, when, this is the record.
Why route it through Datadory
Because the raw artifact answers one question at a time in a browser tab, and every interesting business question wants thousands of bindings across publishers, years and languages in one frame. Datadory industrializes the extract: queries constructed and staged under per-query limits, publisher names resolved beside every identifier, deliveries keyed so consecutive pulls diff into clean timelines, and a sample in hand before any commitment. Browse the rest of the slice on the publishing data hub, see where the pair ranks on best publishing datasets, or get a sample of this dataset cut to your scope.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| Field | Type | Definition | Example |
|---|---|---|---|
head.vars | text | Names of the selected variables in the SPARQL result envelope - the column list every downstream parser binds to. | book | title | isbn13 | pubdate |
results.bindings | text | Array of solution objects, one per result row, mapping each variable to a typed value (URI reference or literal). | one object per row in the sample above |
?book / ?item | string | Entity variable bound to the book item - addressed by Q-number, with the URI form carried on the value itself. | Q100708395 |
?title (P1476) | text | Edition title as a monolingual literal, language-tagged per market. | Silver Economies, Monetisation and Society in Scandinavia, AD 800-1100 (en) |
?isbn13 (P212) | string | ISBN-13 string literal of the edition - the join key onto trade catalogs. | 978-87-7934-585-0 |
?publisher (P123) | string | Publisher variable bound to the publishing-house item; resolved to display names beside the reference on delivery. | publishing-house item (Q-number reference) |
?pubdate (P577) | datetime | Publication date as an xsd:dateTime literal, preserving the precision the statement records. | 2011-01-01T00:00:00Z |
What teams do with it
- Publisher output measurement Count editions per imprint per year from bindings grouped on publisher and publication date - market arithmetic without assembly work.
- Rival catalog watching Diff competitors' edition arrivals and identifier hygiene week over week; one of 195 cataloged datasets tagged for competitor tracking.
- Missing-identifier audits List the editions that lack ISBN-13, publisher or date statements, and turn data-quality hand-waving into a work queue with a length.
- Multilingual market reads Language-tagged titles separate markets natively - no translation layer, no country-code guessing.
- Citation-grade metadata pulls Typed, sourced values that survive review; the catalog holds 765 datasets clearing that bar, and this one resolves editions across languages inside it.
Questions buyers ask
What does wikidata sparql query service data contain?
Typed result rows over Wikidata's publishing graph. Each row is one query binding carrying the edition entity by Q-number, its ISBN-13 (P212), language-tagged title (P1476), publisher reference (P123) and publication date (P577), with the response envelope naming the columns and typing every value as URI reference or literal.
What does one row represent?
One solution binding, so the grain follows the question: a publisher-year count returns one row per publisher-year, an edition listing one row per edition, an identifier audit one row per edition lacking a statement. Say which grain you want and the delivery is shaped to it.
How current is the delivered data?
As current as the cadence you pick - daily, weekly or hourly. Coverage is current-state of a continuously edited graph rather than a dated snapshot, and scheduled captures add history the live view does not retain, such as per-publisher output time series built from recurring pulls.
Can I measure a publisher's output year by year?
Yes. Group editions on the publisher reference and publication date and the result is one row per publisher-year, with date precision reported as recorded rather than imputed. Rollups over any window arrive as derived columns when named at sample request.
How does this differ from the Wikidata knowledge-graph record?
Same asset at two altitudes. The knowledge-graph record ships entity-level rows suited to enrichment, ISBN validation and network loads; this one answers shaped questions - counts, audits, overlap lists - assembled per query. Prototype here, industrialize there; most teams take both in one feed.
Can a sample be scoped to my publishers, languages or years?
Yes. Name the publishers, language markets, date windows or an identifier-audit focus, and the sample returns in exactly the seven-field shape above, pinned to live records - with publisher-label resolution, rollups and crosswalk joins confirming alongside rather than being promised blind.
Notes on this record
- Provenance Compiled from the service's published documentation and live query output during the August 2026 research pass; benchmark rows were verified field by field, and the 122,983,238-entity figure is the catalog's August 2026 reading of the graph underneath.
- Window, not shelf This record answers questions; the twin knowledge-graph record ships the entities themselves. One asset, two altitudes - the comparison walks the trade-offs.
- One question, one shape Grain is declared, not discovered: one row per binding means a publisher-year count and an edition listing share a dictionary while answering different jobs.
- Fill rates are a measurement Property coverage varies across the graph's language communities, so completeness is screened per scope during sampling instead of asserted blanket.
- Scored against the catalog Quality 9 of 10 against a 7.81 mean across Datadory's 1,744 cataloged datasets - the sixteen-record publishing slice averages 8.5, and the verified response contract earns this record's mark.
- Sample policy Samples ship in the exact seven-field shape shown above, cut to named publishers, languages, windows or audit focuses; extensions confirm with the sample rather than being promised blind.
Datasets that pair with this one
- SPARQL endpoint glossary entry What a SPARQL endpoint is, how query-shaped extraction differs from fixed exports, and where this record sits in that vocabulary.
- Best publishing datasets Where this record sits among the sixteen publishing products in the catalog, scored and compared.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.