Datadory notebook

XBRL company facts bulk access, delivered whole

Datadory delivers industrial-conglomerates data covering the whole XBRL company facts population: every standardized financial figure US-listed registrants tag in their own filings - one row per company, concept, period and filing, roughly 945 us-gaap concepts for GE Aerospace alone and concept-period frames lining up 1,934 entities side by side - flattened into typed, joinable rows beside derived statement columns for 130,000-plus symbols, delivered daily, weekly, or hourly.

1,744 datasets. Pick your catch.

What does a bulk company-facts population actually contain?

Five row types carry the whole feed, shown here as they land - verified against GE Aerospace, CIK 0000040545:

# XBRL company facts -- flattened rows
row_type  registrant
cik       0000040545          name      GENERAL ELECTRIC CO
tickers   ["GE"]              exchange  NYSE          fiscal_year_end 1231

row_type  filing_index
accession 0000040545-26-000049   form 10-Q   filed 2026-07-16   fy 2026 fp Q2

row_type  xbrl_fact
tag Revenues   unit USD   period 2026-04-01 -> 2026-06-30   val 13349000000
accn 0000040545-26-000049         frame CY2026Q2

row_type  market_frame
frame CY2025Q1   taxonomy us-gaap   tag Revenues   entities 1934

row_type  ticker_reference
ticker NVDA   ->   cik 1045810   title NVIDIA CORP

Whole universe or named peer set - how should breadth decide?

The population comes in two shapes, and the analysis you intend picks between them. The all-filer population covers every registrant at once and behaves like infrastructure: load it deep, refresh it steadily, never sample it again. The named peer set carries only the companies on your list - a twenty-name conglomerate cohort - and is the right answer whenever the question is comparative rather than exhaustive.

One number decides the second question before any storage is provisioned: a filer of GE's size exposes about 945 us-gaap concepts, so an unfiltered peer pull drags in thousands of concepts you will never model. Scope the allow-list first - revenue, earnings, segment and balance-sheet concepts the model actually uses - and the feed stays proportional to the analysis instead of to the taxonomy. At Datadory the scope locks at sample request: you name the companies and the concepts, and the schema stays stable around them.

The frames view is the shortcut neither shape replaces. One concept across all filers for one calendar period is a peer panel served whole - 1,934 entities for a single quarter of Revenues - which is the fastest way to place one industrial against its median without assembling anything.

Which vintage survives a restatement?

Facts arrive keyed one row per company, concept, period and filing, so successive versions of the same metric sit side by side rather than overwriting each other. That structure forces one honest decision every build must make explicitly: as-filed or latest-revised. Point-in-time work - backtests, event studies, anything a journal reviewer can ask you to reproduce - takes each value from the filing that reported it. Current-state dashboards take the newest vintage and let the older ones stand as history. Both read off the same rows; the policy is a filter, not a rebuild.

Sizing deserves the same explicitness. A single fiscal year of a single metric can carry several vintages, and multiplied across roughly 945 concepts for a large filer, the row count outruns the headline company count fast. Concept filtering at ingest keeps the warehouse proportional - and because every value travels with the accession number that reported it, audit trail survives any amount of downstream aggregation. Nothing here needs to be reconstructed after the fact, because nothing was overwritten in the first place.

What gaps does the filings layer leave - and which records fill them?

Three limits shape planning, and none of them are secrets.

Jurisdiction. The population covers US-listed and US-registered filers, so private companies and non-filing foreign parents are absent by construction. For the UK side of a conglomerate's family tree, Companies House Register Search (UK) holds roughly 5 million live companies with officers, charges and person-of-significant-control records, and retains dissolved records for 20 years so wound-down subsidiaries stay traceable - scored 9 on the same rubric. Our head-to-head vs SEC EDGAR Company Facts & Submissions API splits the two disclosure cultures field by field.

Depth. Machine-readable tagging generally begins in 2007-2009, depending on when each filer started. The filing index reaches further - reassembled, GE Aerospace's runs back to 1994 - and for price and statement context across the longer span, Alpha Vantage Stock Data API carries 20-plus years of daily, weekly and monthly bars beside annual and quarterly statements mapped to GAAP and IFRS taxonomies.

Shape. A fact store is not a tidy one-row-per-year table, and quick screens want exactly that. Wikipedia List of Largest Companies by Revenue ranks the top 50 by consolidated revenue with industry, profit, employees and headquarters columns for a seed list, and the Fortune Global 500 extends the same one-row-per-company logic to 500 names.

Who builds on bulk company facts, and for what?

Investors and quants rebuild as-reported statement history with restatement lineage intact - the primary investors and quants on industrial conglomerates data page collects the workflows. Data scientists get a flat, keyed schema - company, concept, unit, period, filing - that joins cleanly against prices and registries without a parsing project in front of it (data scientists). Competitive-intel and product teams watch what diversified rivals disclose, segment by segment (competitive-intel teams).

Fields carried on every delivered company-facts row (as of August 2026)
FieldTypeDefinitionExample
tag / labelstringStandardized concept name and its official human-readable label.Revenues
valnumberReported value of the fact in the stated unit.13349000000
unitstringUnit of measure the observation is expressed in.USD
start / enddatePeriod start and end dates for a duration fact.2026-04-01 | 2026-06-30
accnstringAccession number of the filing that reported the fact - the provenance key.0000040545-26-000049
fy / fpinteger / stringFiscal year and fiscal period (Q1-Q4, FY) of the reporting period.2026 | Q2
form / filedstring / dateForm type that supplied the fact and the date it was filed.10-Q | 2026-07-16
framestringCalendar period identifier when the fact aligns to a standard frame.CY2026Q2
entityName / cikstringRegistrant name and zero-padded Central Index Key joining facts to identity.GENERAL ELECTRIC CO | 0000040545
The bulk-fundamentals stack: what each Datadory record contributes (Industrial Conglomerates pool, August 2026)
RecordRole in the stackCoverageDatadory score
Companies House Register Search (UK)Ownership and registry depth for the non-filing part of the family tree~5 million live UK companies; dissolved records retained 20 years9
StockAnalysis.comDerived statement columns and multiples, one row per company per metric130,000-plus tradable symbols; statements typically 10-plus fiscal years8
Fortune Global 500One row per company revenue ranking for deck-ready context500 largest corporations, annual editions6
Wikipedia List of Largest Companies by RevenueQuick revenue-ranked seed listTop 50 by consolidated revenue with industry, profit, employees and headquarters columns5

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Industrial Conglomerates United States-registered filers: domestic issuers and foreign…

SEC EDGAR Company Facts & Submissions API

Industrial Conglomerates US-listed symbols primarily (NASDAQ

StockAnalysis.com — Per-Ticker Fundamentals & Valuation Pages

market_cap · revenue_ttm · net_income_ttm …+13 more

Industrial Conglomerates United Kingdom - the England

Companies House Register Search (UK)

Industrial Conglomerates Global equities: US exchanges plus international listings

Alpha Vantage Stock Data API — Global Equity Prices & Fundamentals

Industrial Conglomerates Global

Wikipedia — List of Largest Companies by Revenue

Industrial Conglomerates Global - companies headquartered worldwide, tagged by country…

Fortune Global 500

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

What does an XBRL company facts dataset include?

Every standardized financial figure US-listed registrants tag in their own filings: revenues, segment detail, balance-sheet and cash-flow line items, each keyed to the filing that reported it. Verified against GE Aerospace, that means about 945 distinct us-gaap concepts, roughly 1,002 recent filings in the online index, and concept-period frames covering 1,934 entities for a single quarter of Revenues.

How far back do XBRL company facts reach?

Machine-readable tagging generally begins in 2007-2009, depending on when each filer started tagging under the standard taxonomy. The filing index reaches further - reassembled, GE Aerospace's runs back to 1994 - and for longer histories, delivered companions such as Alpha Vantage carry 20-plus years of prices and GAAP- and IFRS-mapped statements.

Do restated figures overwrite the originals?

No. Facts arrive one row per company, concept, period and filing, so a re-reported value lands as another dated vintage while every earlier one stands. Your build chooses an as-filed or latest-revised policy as a filter over those rows, and point-in-time reconstruction stays possible after the fact because nothing was ever overwritten.

Can I take a named peer set instead of the whole filer population?

Yes, and for comparative work it is usually the better cut. You name the companies and the concepts - a filer of GE's size exposes about 945 us-gaap concepts, so scoping the allow-list first keeps the feed proportional to the model - and the schema stays stable around them. Scope locks at sample request, with the full population available whenever breadth becomes the question.