Google Patents Public Data on BigQuery
Datadory delivers electronic equipment instruments data covering the world's patent record through Google Patents Public Data on BigQuery: 98,176,830 publications in one queryable table - multilingual titles and abstracts, US claims, harmonized assignees and inventors, CPC/IPC classifications, citations and four dates per filing across every major patent office - delivered daily, weekly, or hourly on the cadence you choose.
What is Google Patents Public Data on BigQuery?
The whole world's patent paperwork, stacked into one addressable table. Google Patents Public Data on BigQuery is a worldwide bibliographic and US full-text patent corpus assembled by Google with IFI CLAIMS Patent Services: 98,176,830 publications in a single table in the documented snapshot - about 899 GB of them - each row carrying DOCDB-compatible publication and application numbers, multilingual titles and abstracts, US claims and descriptions, priority, filing, publication and grant dates, harmonized inventor and assignee names, examiners and art units, nested CPC, IPC, USPC, FI and F-term classification records, and citation arrays. Companion tables add derived research features, and dated snapshots such as publications_201710 and publications_201802 freeze earlier cuts.
Within Datadory's 16-dataset electronic-equipment-instruments pool this is the deep-history innovation view, and it carries a quality score of 8 out of 10 with definitions verified during cataloging - a standard met by 85.7% of the 1,744 datasets Datadory catalogs. Get a sample of this dataset - name a CPC slice like H01L for semiconductor devices, an assignee list or a country set, and matching rows come back first.
What does a sample row look like?
One captured record beats a paragraph of specification. This is the shape of a single publication row from the core table as documented:
publication_number : US-7650331-B1
application_number : US-87124404-A
country_code : US
kind_code : B1
application_kind : A (patent)
family_id : 45133502
title (en) : Semiconductor device and method...
abstract (en) : A semiconductor device includes...
priority_date : 20040315
filing_date : 20050912
grant_date : 20100119
inventor : [{"name": "John Smith"}]
assignee : [{"name": "IBM"}]
cpc.code : H01L21/02
ipc.code : H01L 21/02
citation : [{"publication_number": "US-6000000-A"}]Notice what one row settles immediately: the family ID that groups every publication sharing the same priority claims across offices, the four distinct dates that let you measure examination lag rather than guess at it, the harmonized assignee that keeps IBM-as-IBM intact across countries, and the nested classification that makes a CPC filter a one-line operation. Multiply by ninety-eight million rows and electronics questions - who files in H01L, where claims cite whom, which art units handle what - become counting problems.
What fields does the dataset include?
Twenty-one fields carry a worked example in the dictionary below, all drawn from the record shape shown above; definitions were verified against documentation during cataloging. The full schema is wider than the twenty-one: raw (non-harmonized) inventor and assignee name records sit alongside the harmonized ones, PCT application numbers, examiner names and FI/F-term classifications ride along populated but ship without fixed worked examples here, so they fold under additional fields on request - together with the companion research table's engineered columns and dated snapshot cuts if your use case wants them.
What does coverage look like across geography, time and granularity?
Geography - worldwide bibliographic coverage across the patent offices contributing to DOCDB, plus full-text claims and descriptions for US publications. Country codes make office-level splits trivial; harmonized assignees keep multi-office filers whole when you aggregate.
Temporal - historic filings through the most recent refresh, with four dates per publication (priority, filing, publication, grant) so any trendline can be drawn on the calendar that matches the question. The documented snapshot counted 98.2M publications; the live count runs higher.
Granularity - one row per patent publication, with simple-family grouping available via family_id whenever you want invention-level instead of document-level counts.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Who uses this data, and for what?
- Innovation and technology landscaping - count filings by CPC class, country and year to map where electronics R&D actually concentrates. It is one of 765 cataloged datasets serving citation-grade research, and the only one in this industry that spans every major patent office in a single table.
- Citation-network analysis - walk the cited-reference arrays between publication numbers to build forward and backward citation graphs, the raw material for patent-quality and knowledge-flow studies. Feeds the same shelf as the rest of the quant backtesting pool when innovation factors need inputs.
- Competitor roadmap evidence - track filing velocity and CPC drift for named assignees; a competitor's H04-heavy quarter says more than its press releases. One of 195 datasets tagged to competitor tracking.
- Assignee intelligence for sales - find companies investing in electronics R&D before they appear on anyone's target list, keyed off harmonized assignee names.
- Examination-lag and portfolio studies - the four-date structure turns grant latency, prosecution duration and family size into direct measurements rather than estimates.
Which personas get the most value?
Ranked by relevance in Datadory's persona tagging:
- Data Scientists & ML Engineers (relevance 3): query ~98M patent publications for CPC-level innovation and citation-network features without leaving SQL - see the data scientists x electronic equipment instruments workflows.
- Investors & Quant Researchers (relevance 3): build innovation-factor and litigation-adjacent signals from citation graphs spanning worldwide offices - investors & quants page.
- Market Researchers & Consultants (relevance 3): measure technology-field innovation intensity across countries on consistent bibliographic ground.
- Journalists, Academics & Students (relevance 3): quote filing counts and citation structures from the IFI CLAIMS/Google corpus at citation grade.
Sales & Growth Teams and Competitive Intelligence & Product Teams both tag it at relevance 2 - assignee-based prospecting and filing-velocity tracking respectively; Developers & Builders likewise at 2 for running analytics over one table instead of parsing bulk XML.
What should I know before requesting a sample?
Three honest gaps. First, vintage of the headline figure: the 98,176,830-row count comes from the documented snapshot dated 2018-11-26 - the live total is higher, but treat any exact number, including ours, as a floor. Second, text depth is uneven by office: bibliographic fields cover every contributing authority worldwide, while full claims and descriptions exist for US publications only, so claim-length analytics are a US-only read. Third, cadence upstream is not published anywhere we could verify - which matters less than it sounds, because Datadory's own delivery schedule is what sets your freshness: daily, weekly or hourly captures land on your side regardless. None of this is hidden at delivery; say which gap bites your question and the sample comes back shaped around it.
Which datasets sit next to this one?
Electronic Equipment & Instruments catalog neighbors that answer different halves of the electronics question. USPTO PatentsView API - Electronics Patents covers the US-only side with disambiguated inventors and bulk TSVs - deeper entity resolution, narrower geography. UN Comtrade Plus TradeFlow - Electrical Machinery & Instruments (HS 85/90) trades innovation signals for physical trade flows in the same HS chapters. WIPO Global Innovation Index & IP Statistics compresses IP activity into ~140-economy annual indicators when you want rankings over records. Or read the vs UN Comtrade Plus TradeFlow comparison for the head-to-head.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
publication_number | string | Patent publication number in DOCDB-compatible form. | US-7650331-B1 |
application_number | string | Patent application number, DOCDB-compatible; may be unset. | US-87124404-A |
country_code | string | Two-letter code of the publishing authority. | US |
kind_code | string | Kind code indicating application, grant, search report, correction, etc.; differs per country. | B1 |
application_kind | enum | High-level application kind: A=patent, U=utility, P=provision, W=PCT, F=design, T=translation. | A |
family_id | string | Simple patent family identifier; grouping on it returns all publications sharing the same priority claims. | 45133502 |
title_localized | text | Repeated record of publication titles with language codes. | {"text": "Semiconductor device", "language": "en"} |
abstract_localized | text | Repeated record of abstracts with language codes. | {"text": "A semiconductor device includes...", "language": "en"} |
claims_localized | text | Claims text, US publications only. | {"text": "1. A method of manufacturing...", "language": "en"} |
publication_date | integer | Publication date as YYYYMMDD integer. | 20100119 |
filing_date | integer | Filing date as YYYYMMDD integer. | 20050912 |
grant_date | integer | Grant date as YYYYMMDD integer. | 20100119 |
priority_date | integer | Earliest priority date as YYYYMMDD integer. | 20040315 |
inventor_harmonized | string | Harmonized inventor names with normalized forms across countries. | [{"name": "John Smith"}] |
assignee_harmonized | string | Harmonized assignee names grouped per country. | [{"name": "IBM"}] |
cpc | string | Cooperative Patent Classification codes assigned to the publication. | [{"code": "H01L21/02"}] |
ipc | string | International Patent Classification codes. | [{"code": "H01L 21/02"}] |
uspc | string | US Patent Classification codes. | [{"code": "438/455"}] |
citation | string | Cited references with publication numbers and categories. | [{"publication_number": "US-6000000-A"}] |
entity_status | enum | US entity size status (large/small) affecting fee schedule. | small |
art_unit | string | USPTO art unit handling the application. | 2835 |
Questions buyers ask
How many patents does the google patents public data on bigquery data cover?
The documented snapshot holds 98,176,830 publications - roughly 899 GB - in the core table, spanning every major patent office that contributes to DOCDB plus US full text. The live count runs higher than the documented figure, so treat any precise number, including ours, as a floor. Family grouping via family_id collapses documents into inventions.
Does the dataset include claims text for every country?
No. Bibliographic fields - titles, abstracts, dates, classifications, parties, citations - cover all contributing offices worldwide, while full claims and descriptions are populated for US publications only. Claim-language analytics are therefore a US read; worldwide studies work on the bibliographic spine, which is complete everywhere.
How far back does the patent history go?
Historic filings through the present refresh window, with four dates per publication - priority, filing, publication and grant - each stored as YYYYMMDD integers. Because priority dates survive across offices, multi-country trendlines stay anchored to first filing rather than drifting with later publication events.
Can I isolate semiconductor or communications patents specifically?
Yes - classification does that work. Nested CPC, IPC, USPC, FI and F-term records hang off every publication, so a filter like CPC prefix H01L isolates semiconductor devices, H03F amplifiers, H04 communication and G06 computing in one pass. Class-code slices are the standard way samples get shaped for electronics questions.
How do I tell one invention from its foreign publications?
The simple-family identifier groups every publication sharing the same priority claims across offices, so a worldwide count can be de-duplicated down to invention level in a single aggregation. Document-level counts and family-level counts answer different questions; both come from the same rows.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.