Data source

Data from Open Library, delivered clean.

3 datasets pulled from Open Library's releases, checked field by field and shipped the way you want them — daily, weekly, or hourly, your call.

  • 3 datasets
  • 1 industry
  • Real rows on request

What Datadory delivers from Open Library

3
Publishing Global bibliographic catalog spanning publish… · Monthly full snapshots of the current catalog

Open Library Monthly Data Dumps Data

Publishing Global bibliographic coverage · Works from earliest printing to current titles

Open Library Search & Works/Editions API Data

Publishing Global catalog of published books · Works from early printing history through cur…

Open Library General Catalog - Browsable Lending Library Data

Pick a catch, see the rows.

Name any Open Library dataset and we send real rows from it — not a screenshot of rows. 1,744 datasets. Pick your catch.

Get a sample

API, files, or your warehouse. Daily, weekly, or hourly.

Straight answers about Open Library data

What does Datadory deliver from Open Library?

Three scored records describing one crowd-edited bibliography: the whole-catalog snapshot files at 10 out of 10, the query layer over the identical catalog at 10 out of 10, and the browsable discovery surface at 8 out of 10. All three arrive cleaned, keyed and documented as five named deliverables.

How many works does the catalog cover?

A single keyword matched 7,723,767 works during the August 2026 verification pass, and the subject tree runs from History's 2,550,711 ebooks down to Plays' 3,037. Counts move with community edits, so every Datadory delivery pins its reconciliation total beside the rows.

What is the difference between a work and an edition here?

A work is the abstract title; editions are its concrete printings with their own publishers, dates and ISBNs. The sample work Cross carries 49 editions spanning a `first_publish_year` of 1684 and five language codes - one identity, forty-nine physical manifestations.

What do the availability flags actually mean?

`ebook_access` grades what a reader can get - borrowable, full-text or print-only - while `has_fulltext`, `public_scan_b` and the `ocaid` scanned-copy identifier say whether a digitization exists at all. Treat them as statements about Internet Archive holdings, and filter before promising digital delivery.