Citation Grade Research: Datasets Built to Be Cited

Citation grade research needs primary-source data published by the record-holder itself: 765 datasets in Datadory's catalog serve it directly, led by FEC Campaign Finance Data & OpenFEC API and FRED - Federal Reserve Economic Data (St. Louis Fed). Of the 765, 660 are free and 203 state public-domain terms.

What is citation-grade research?

Citation-grade research is the practice of sourcing every figure in a story, paper or filing from the authoritative publisher of that record, so any reader can retrace the number to its origin. The pattern comes straight out of newsroom and academic workflow: instead of citing an aggregator's chart, you cite the FEC's transaction-level record of who spent on federal elections, verifiable down to each individual transaction, or FHFA's own house price index rather than a secondhand summary.

Three properties decide whether a source qualifies. The publisher must hold the primary record, the unit of observation must survive intact rather than arriving pre-aggregated, and the reuse terms must allow quotation with attribution. In Datadory's catalog, 765 datasets clear that bar across 133 industries; 203 of them state public-domain licensing and 110 ship under commercial delivery terms 4.0.

Which datasets support citation-grade research?

The scorecard below draws nine datasets from the 765 that Datadory tags for this play, chosen to show its range: political money, monetary statistics, housing, corporate registries, defense spending, genomics, chemistry and international banking.

Across the full serving pool, access divides as follows: 370 of the 765 arrive as bulk downloads, 251 through official APIs, 121 as scrapeable sites, 21 by request form and 2 over FTP. Freshness varies just as widely - 170 of the 765 refresh daily or in real time, 156 update monthly, and 226 publish on an annual cycle that matches yearbook-style references.

How do teams actually do citation-grade research?

  1. Start at the record-holder, not a copy. For US election spending that is the Federal Election Commission: the OpenFEC REST API serves Schedule E independent expenditures with 81 fields including payee_name, expenditure_amount and a pdf_url pointing at the filed document image. For UK company facts it is Companies House, whose register carries filings observed back to 1862.
  2. Match the delivery channel to the volume. Single questions suit the API route - OpenFEC allows 1,000 calls per hour on a free api.data.gov key at 100 results per page, and FRED returns up to 100,000 observations per request in JSON, XML, CSV or Excel. Whole-universe work belongs in bulk files: FHFA's master HPI file ships as CSV, JSON, XML and SQL, and the BIS locational banking CSV alone runs about 122 MB compressed with roughly 609,000 rows.
  3. Reconcile revisions before totalling anything. The FEC warns that original and amended reports both appear in its transaction files, so deduplicate on amendment indicators and file numbers before summing. SIPRI folds revisions into one workbook and republishes each spring, which is why its own guidance is to re-download the full file every cycle rather than patch old copies.
  4. Capture provenance at download time. Record the release identifier and date - GenBank Release 273.0 (August 2026), the FRED vintage date, the SIPRI DOI - because 75 of the 765 datasets here carry an explicit DOI and the rest need at least an endpoint URL and access date to make a claim auditable.
  5. Match the citation string to the license. BIS permits unrestricted use provided BIS is named as source; FHFA asks to be cited whenever index values are redistributed; SIPRI mandates the credit line "Information from the Stockholm International Peace Research Institute (SIPRI)". OpenSecrets site content is CC Attribution-Noncommercial-Share Alike 3.0 US, so commercial reuse needs prior written permission, typically via an OS Pro subscription at $3.99/month billed $40/year.

Which personas run this play?

The tagging is lopsided toward the newsroom-and-campus profile: 730 of the 765 serving datasets also carry a relevance tag for [data for journalists-academics], against 50 for [data for market-researchers] and 15 for developers and builders.

That ordering tracks what each persona needs from a citation. Market-research overlap clusters on industry statistics usable as a quoted yardstick - the IAB Insights & Industry Research Library, updated weekly, holds the IAB/PwC Internet Advertising Revenue Report behind a free account. Developer overlap clusters on documented endpoints: the World Bank Military Expenditure Indicator API returns SIPRI-sourced military spending for 200+ countries in a single call with per_page=20000, under commercial delivery terms 4.0 terms.

Free vs paid for citation-grade research

Free dominates this use case: 660 of the 765 datasets cost nothing, 88 are freemium, 14 are paid and 3 do not state pricing - 86.3 percent free, against 82.4 percent free across the entire 1,744-dataset catalog. The skew is structural, because most citation-grade records originate with governments, central banks and treaty bodies that publish under public-domain or open terms.

The paid tier buys editorial synthesis rather than raw records. The Military Balance (IISS) is a ~450+ page reference volume covering 170+ countries, delivered by request form alongside the Military Balance+ online database, and LBMA Precious Metal Benchmark Prices licenses benchmark price histories. The freemium middle trades convenience for terms: OpenSecrets bulk downloads are free after registration and approval, but commercial use requires written permission or OS Pro, and reading full IAB-branded reports requires a free account.

Frequently asked questions

Where do journalists find primary-source data for fact-checking?

Start with the regulator that owns the record. The FEC's OpenFEC API exposes every federal campaign finance filing as structured data, updated nightly, with a pdf_url linking each expenditure to the filed document image. FHFA publishes its house price indexes directly as CSV, JSON, XML and SQL, and the UK's Companies House register reaches back to 1862.

Can I republish these datasets in an academic paper?

Mostly yes, with attribution. Within the serving pool, 203 datasets state public-domain terms and 110 ship under commercial delivery terms 4.0, which permits reuse with credit. Some publishers add citation conditions: BIS requires being named as source, SIPRI mandates a specific credit line, and OpenSecrets site content is commercial delivery terms-NC-SA 3.0 US, which blocks commercial reuse without permission.

Do I need an integration key to download citation-grade data?

Sometimes, and keys are normally free. OpenFEC issues a free api.data.gov key allowing 1,000 calls per hour, FRED keys are issued free after registering an account, and Companies House passes the key as a basic-auth username at 600 requests per 5 minutes. Where no key exists, bulk files substitute: FHFA serves its master HPI CSV with no registration or terms-of-use gate.

How current is the data behind a citable claim?

It depends on the publisher's cycle. Among the 765 serving datasets, 131 update daily and 39 stream in real time, 156 refresh monthly and 226 publish annually. Annual does not mean stale: SIPRI's military expenditure series spans 1949-2025 in one workbook revised each spring, and GenBank posts daily increments between bimonthly numbered releases.

Is paid data ever necessary for citation-grade work?

Rarely, and only for synthesis rather than records. Just 14 of the 765 datasets are paid, led by editorial reference products like The Military Balance (IISS), a ~450+ page volume covering 170+ countries sold with the Military Balance+ database, and LBMA Precious Metal Benchmark Prices. Government and multilateral equivalents - OpenFEC, FHFA, BIS and World Bank indicators - stay free.