Research & Consulting Services · USPTO / PatentsView
PatentsView — USPTO Patent Data API
Datadory delivers patentsview uspto patent data api data covering the full history of United States patent grants plus pre-grant publications from 2001 - millions of documents keyed by patent identifier, each joinable to disambiguated inventors, assignees and locations, four classification systems, backward and forward citations and full claim-level text, delivered daily, weekly, or hourly.
API, files, or your warehouse. Daily, weekly, or hourly.
- Where it covers
- United States grants and pre-grant publications, with worldwide reach in the entity tables - inventor and assignee addresses resolve across countries
- How far back
- Grants from 1976 and pre-grant publications from 2001 - the full history of US patents, with each build stamped by its data-through date
- How fine
- Record-level patent, inventor, assignee, classification and citation entities with cross-entity joins - one row per document, fanning out per entity
- record families
- Granted patents, pre-grant publications, inventors, assignees, applicants, attorneys, locations, classifications, citations, related applications and full-text tables
What is the PatentsView — USPTO Patent Data API dataset?
It is the research platform built over the United States Patent and Trademark Office's own grant record, and it does the two things raw patent files never do: it disambiguates the actors and it classifies the technology. Millions of granted patents arrive as rows keyed by patent_id, carrying titles, abstracts, grant dates and patent types, joined through disambiguated inventor, assignee and location tables. Grants reach back to 1976 and pre-grant publications start in 2001, so one table spans the entire modern biotech, software and semiconductor eras.
Around the spine sit the joins that make the corpus analytic rather than merely searchable: classifications across four systems (USPC, CPC, IPC and the WIPO technology fields), backward and forward citations both US and foreign, other references, related applications, applicant and attorney rosters, and full text down to individual claims.
For research and consulting work this is the output side of the innovation economy - where NCSES counts R&D dollars going in, this record shows what came out the other end. Get a sample of this dataset scoped to your technology classes.
What do sample records look like?
A few rows from different corners of the same corpus:
# PatentsView -- USPTO Patent Data -- grant row -- research & consulting services
patent_id : 10000000
patent_title : Coherent LADAR using intra-pixel quadrature detection
patent_date : 2018-06-19
patent_year : 2018
patent_type : utility
# an older grant from the same table - the history runs deep
patent_id : 7861317
patent_date : 1981-10-06
patent_year : 1981
# classification join -- the technology address per grant
patent_id : 9954111
cpc_class_id : H01L (semiconductor devices)
# entity join -- who owns it, resolved to a stable identity
patent_id : 11872626
assignees.assignee_organization : DIVERGENT TECHNOLOGIES, INC.Read them as one document wearing three hats. The first two rows are the bibliographic spine - what was granted and when, from 1981 hardware to a 2018 optics grant, in one continuous series. The classification join adds the technology address: cpc_class_id H01L pins a patent inside the semiconductor-device literature, which beats the word "chip" in a title by any measure you care to run. The entity join carries ownership - and because the organization name travels with its disambiguated form, "Divergent Technologies" stays one company even when the printing changes.
Values arrive exactly as the office publishes them: raw codes and un-normalized names alongside their resolved identifiers. Get a sample of this dataset and read your own slice.
What fields does the dataset include?
Twelve fields carry the core extract, doing four jobs. patent_id is the join surface everything else hangs off. The bibliographic group - patent_title, patent_abstract, patent_date, patent_year, patent_type - describes the document. patent_num_times_cited_by_us_patents measures its footprint in one column. And the relational group reaches outward: inventors.inventor_name_last and assignees.assignee_organization name the actors, application.filing_date timestamps the idea rather than the grant, cpc_class_id places it technologically, and citation_sequence keeps reference lists faithful to the printed document.
Everything else rides under additional fields on request: pre-grant publications, applicant and attorney rosters, location identifiers, the remaining classification systems, foreign citations, related applications and the full-text surfaces.
<!--TABLE:field_dictionary-->
Where does coverage run, and at what grain?
- Geography: United States grants and pre-grant publications, with worldwide reach in the entity tables - a US-granted patent with a Tokyo assignee and a Munich inventor keeps all three geographies.
- Temporal: the full history of US patents. Grants run back to 1976, pre-grant publications to 2001 - five decades long enough to watch a technology class rise, peak and be absorbed into the next one. Each build is stamped with its data-through date, so trend work knows exactly where the window closes.
- Granularity: record-level entities with cross-entity joins. One row per document on the spine, then one row per classification, per inventor, per assignee, per citation in the joined tables - the fan-out that makes network analysis possible instead of approximate.
Against the wider catalogue - average quality score 7.81 across all 1,744 datasets - this slice scores 6/10, docked mainly for self-documentation that runs thinner than a statistical agency's. The entity resolution, meanwhile, is something no raw patent dump provides at any price.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
The delivery step earns its keep here, because the hard part of patent data was never finding it - it was assembling it. Reconciling inventor spellings into one person, aligning classification versions across decades, and rebuilding citation joins are exactly the projects that eat a quarter and then need defending forever. Delivered extracts arrive with that assembly done: identities resolved, classifications decoded, joins keyed on patent_id. Cadence is yours to set - daily for monitoring, weekly for landscaping refreshes, hourly if a pipeline demands it.
Who uses this data, and for what?
Technology landscaping and white-space mapping. Subgroup-level classification turns a crowded field into a map: count filings by CPC class across five decades and the thin patches announce themselves.
Competitive intelligence on R&D. Filing counts, class mix and claim depth by disambiguated assignee give the earliest public signal of where a rival's engineering budget is going - years before a product roadmap confirms it.
Citation-network analytics. Backward references and forward citing documents make influence traceable, so a foundational patent can be followed into everything built on it, with citation_sequence preserving the order the examiner saw.
Innovation benchmarking and policy studies. University-versus-corporate output splits and national innovation comparisons computed from the official record - quotable in client decks because the citation traces to the patent office itself.
Prospecting for IP-adjacent services. Firms that file in a given class are named, classified buyers for prior-art search, patent-analytics and translation vendors; assignees.assignee_organization is the account list.
Which personas get the most value?
Market Researchers & Consultants get the innovation-output layer their engagement models usually fake - technology trends measured from filings rather than sketched from a vendor deck. Competitive Intelligence & Product Teams get rivals' R&D direction as an always-current feed of classified, owned, dated documents. Data Scientists & ML Engineers get stable entity keys and panel structure clean enough to model directly, skipping the entity-resolution project that usually eats the first month. Journalists, Academics & Students get the citable official record behind almost every story about American innovation.
Related pages: market researchers using research & consulting services data · competitive intel & product teams using research & consulting services data · data scientists using research & consulting services data
What should I know before requesting a sample?
Three things worth saying plainly. First, names arrive un-normalized by design: the printed organization name travels beside its resolved identity, so analyses should group on the identifier and label with the name - never the reverse.
Second, classification depth varies by era. A recent patent may carry a dozen CPC entries while a 1976 design grant carries one legacy code, so class-count comparisons need era-aware baselines; the multi-system classification tables exist precisely to bridge that gap.
Third, the platform itself is being folded into USPTO's own data infrastructure, and the plumbing is moving with it. Datadory absorbs that churn so your extracts stay continuous - and every build is stamped with its data-through date, so pin snapshots if a report will be re-run. Name the classes, assignees or years you want and the sample arrives cut to them. Get a sample of this dataset
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| Field | Type | Definition | Example |
|---|---|---|---|
patent_id | string | Patent number identifier used as the primary key - every inventor, assignee, classification and citation row joins back to it. | 7861317 |
patent_title | text | Title of the granted patent as printed on the front page. | Coherent LADAR using intra-pixel quadrature detection |
patent_abstract | text | Abstract text of the patent, summarizing what the invention claims and how it works. | <returned with your sample> |
patent_date | date | Grant date of the patent. | 1981-10-06 |
patent_year | integer | Year of grant - the standard axis for filing-trend series. | 1981 |
patent_type | enum | Patent kind, separating utility from design and plant grants. | utility |
patent_num_times_cited_by_us_patents | integer | Count of citations received from later US patents - the headline influence measure before any network analysis starts. | <returned with your sample> |
inventors.inventor_name_last | string | Disambiguated inventor surname, carried as a nested field so one grant holds its full inventor list. | <returned with your sample> |
assignees.assignee_organization | string | Disambiguated assignee organization name - the resolved owner identity that survives renames and mergers. | DIVERGENT TECHNOLOGIES, INC. |
application.filing_date | date | Filing date of the underlying application - the moment the idea entered the record, often years before grant. | <returned with your sample> |
cpc_class_id | string | Cooperative Patent Classification class identifier - the technology address that makes the corpus sliceable. | H01L |
citation_sequence | integer | Order position of a citation within the patent's citation list, keeping reference tables faithful to the printed document. | <returned with your sample> |
Coverage chips
| Dimension | Coverage |
|---|---|
| Geography | United States patent grants and pre-grant publications, with global reach in the entity tables - inventor and assignee addresses resolve worldwide |
| Temporal | The full history of US patents: grants reach back to 1976 and pre-grant publications begin in 2001, with every build stamped by its data-through date |
| Granularity | Record-level patent, inventor, assignee, classification and citation entities with cross-entity joins - one row per document on the spine, fanning out per entity |
| Record families | Granted patents, pre-grant publications, inventors, assignees, applicants, attorneys, locations, classifications, citations, related applications and full-text tables |
What teams do with it
- Technology landscaping and white-space mapping Count filings by classification across five decades and the crowded neighborhoods and the empty ones appear on the same chart - the map that decides where an R&D partnership or a freedom-to-operate screen goes next.
- Competitive intelligence on R&D output Track rivals' filing counts, class mix and citation footprint quarter by quarter using the disambiguated assignee identity, and see a pivot into a new technology class years before it reaches a product roadmap.
- Citation-network and influence analysis Follow a foundational patent forward into everything that cites it, or walk a new filing backward to its intellectual ancestors - computable from delivered tables rather than approximated from a single count column.
- Innovation benchmarking and policy studies Split output between universities, corporations and government grantees, compare national innovation profiles, and quote figures that trace to the patent office itself - the difference between a study that survives review and one that does not.
- Prospect universes for IP-adjacent services Firms filing in a given class are named, located, classified buyers for prior-art search, patent analytics and translation work - an account list that builds itself from the very activity it sells to.
Questions buyers ask
What is the PatentsView USPTO Patent Data API dataset?
Datadory's packaged delivery of the USPTO-backed PatentsView patent corpus: millions of granted patents plus pre-grant publications, each keyed by patent identifier and joined to disambiguated inventors, assignees and locations, multiple classification systems, citations and full text.
How far back does the patent history reach?
Grants run to 1976, giving five decades of continuous classified history, while pre-grant publications begin in 2001. That span covers the entire modern semiconductor, software and biotech eras without stitching separate sources together.
What does inventor disambiguation mean in practice?
Every inventor, assignee and location resolves to a stable identifier, so differently spelled names, renames and mergers collapse into one entity. Raw printed names travel alongside for labeling and audit, which keeps a corporate R&D league table from fragmenting into spelling variants.
Which classification systems come with the data?
Four: the legacy USPC scheme, the current Cooperative Patent Classification (CPC), the International Patent Classification (IPC), and the WIPO technology fields that group patents into comparable sectors such as computer technology or semiconductors. Subgroup-level codes pinpoint the niche.
Can I trace citations between patents?
Yes. Citation tables link each patent to its backward references and forward citing documents, covering US and foreign citations with sequence positions preserved. Influence chains, foundational-patent identification and technology-diffusion networks are computable directly from the delivered tables.
Does the dataset include full patent text?
Yes - as separate text surfaces: abstracts, claims, brief summaries, detailed descriptions and drawing descriptions. Claims are the legally operative text and the usual input for prior-art screening, landscape clustering and language-model training sets.
Datasets that pair with this one
- best research & consulting services datasets The scored shortlist for this industry, ranked on documented fields, reliable delivery and freshness.
- research & consulting services data hub Where the innovation-output record sits among the industry's other data-bearing sources.
- NCSES R&D and Business Innovation Statistics The input-side companion: who spends on R&D, to sit beside the record of who files for patents.
- USPTO PatentsView API - Electronics Patents The technology-scoped slice of the same corpus, cut to the electronics classes.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.