Industrial Conglomerates · Crunchbase
Crunchbase
Datadory delivers crunchbase data covering millions of company profiles worldwide, deepest on US and European venture-backed firms: funding rounds reaching back to the late 1990s, acquisitions, investors, leadership changes, subsidiaries and portfolios, plus predictive acquisition and funding-likelihood signals - delivered daily, weekly, or hourly.
API, files, or your warehouse. Daily, weekly, or hourly.
- Where it covers
- Global, with depth concentrated in US and European venture-backed companies
- How far back
- Funding events from the late 1990s to the present - roughly three decades of private-market activity
- How fine
- One row per company, funding round, investor, person and acquisition event
What is the Crunchbase dataset?
Crunchbase is the record of who funded what, who bought whom, and who just took which job - held as profiles for millions of companies, with the operator declining to publish a verified count. Five families of records hang off each profile: funding rounds, acquisitions, investors and the people behind them, subsidiaries and portfolios, and the growth-signal stream of leadership hires and product launches that fills the gaps between deals. On top of the records sits the part its homepage leads with: acquisition-prediction, funding-likelihood and growth-insight scores computed over the corpus.
Inside industrial conglomerates research it plays a role nothing else in the slice plays. The filings layer - SEC EDGAR above all - describes listed parents. The macro layer - BEA, Eurostat, World Bank - describes industries. This is the only one of the 23 primary industrial-conglomerates datasets that covers private-market financing: the scale-up raising a Series C, the grid-equipment maker absorbed by a rival, the investor network connecting them. Coverage is global with its deepest concentration in US and European venture-backed companies. See where it sits on our industrial-conglomerates data hub.
What do sample rows look like?
Illustrative rows in the delivered shape - one record family per row:
# one row per record - company, funding round, investor, person, acquisition or signal
record_type : funding round company : <industrial-robotics scale-up>
stage : Series C amount_raised : on request
lead_investor: on request event_date : on request
record_type : acquisition event company : <grid-equipment maker>
counterparty: on request direction : acquired-by
record_type : person company : <robotics scale-up>
full_name : on request role : leadership hire
record_type : company hq_location : on request
sector_tags : [industrial automation, energy equipment]
signal : acquisition / funding-likelihood / growth scores on qualifying profilesRead the anatomy rather than the placeholders. A Series C row, an acquisition row, a person row and a bare company profile all share one spine - a company key that every family resolves back to. That is the property that turns three decades of scattered events into a panel: stack the rows and any given company's funding history, M&A exposure and leadership churn assemble themselves in order, no entity-resolution pass required. Live values for the companies you name arrive with the sample.
What fields does the dataset include?
Ten fields span the five record families, keyed on the company spine: identity (company_name, company_key), placement (headquarters_location, sector_tags), the dated events (event_date, amount_raised, counterparty_name, role_or_stage) and the predictive layer (signal_score). One row carries one family, so a funding round and the leadership hire six months later read as two rows of the same company rather than two different databases.
Provenance, stated plainly: the operator publishes no field-level data dictionary, so these definitions trace to the platform's observable surfaces - profile pages, search filters, event feeds - rather than to a published schema. They hold across the corpus, but the exact column names and typed extract schema fold under additional fields on request and are pinned down against a delivered sample before anything depends on them.
What does coverage look like across geography, time and granularity?
Geography - global by design, unevenly deep by gravity: the strongest concentration sits in US and European venture-backed companies, which is precisely where industrial-tech scale-ups live. A Munich robotics Series C and an Austin sensor round resolve equally well.
Temporal - funding events reach from the late 1990s to the present, roughly three decades of private-market activity. That span crosses the dot-com cycle, the 2010s industrial-software wave and the recent climate-hardware run inside one consistent record, which is what makes longitudinal questions answerable at all.
Granularity - five record families descending from company profile to individual event: per company, per funding round, per investor, per person, per acquisition event. Deals and leadership changes stay individually addressable instead of dissolving into annual aggregates - the difference between reading a trend and dating one.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Cadence is your call even though deal events land when they land - load the historical depth once as a backfill, then keep new rounds, acquisitions and hires rotating into place on whatever schedule your models and alerts expect. Rows arrive flattened to the ten-field shape above with company identities already normalized, so joins against your CRM or feature store need no fuzzy-matching pass. A sample cut to your named companies comes first either way.
Who uses this data, and for what?
- Competitive early-warning - a rival's fresh round, new board seat or quiet acquisition surfaces as a row long before it becomes a press-cycle narrative; see competitive intel product teams use cases.
- Private-target screening - investors trace funding rounds, acquirers and investor networks across private industrial-tech names back to the late 1990s; see investors quants use cases.
- Outbound prioritization - sales teams rank accounts by recency of funding and leadership movement, because a company that just raised buys differently than one that just restructured; see sales growth teams use cases.
- Landscape sizing - researchers count emerging players per sector from funding and acquisition records instead of stitching trade-press anecdotes; see market researchers use cases.
- Feature enrichment - data scientists join funding and investor history onto firm records as model features; see data scientists use cases.
Which personas get the most value?
Three personas hold this dataset at relevance 3 of 3 in their industrial-conglomerates packs - sales and growth teams, investors and quant researchers, and competitive-intelligence teams - an unusually strong showing in a slice where most records serve macro analysts. The reason is structural: everything else in the vertical watches companies that already filed; this one watches the ones that have not yet.
Market researchers and consultants (relevance 2) get sector landscape maps with player counts they can defend. Data scientists (2) get event-level funding history as features rather than as prose. Developers and builders (2) wire enrichment into internal tools on a fixed per-record shape. Journalists and academics (2) get dated, attributable facts for the startup-and-industry story - see journalists academics use cases.
How does it compare to alternatives in its slice?
Within industrial conglomerates data, Crunchbase owns private-market events and concedes nearly everything else. SEC EDGAR Company Facts & Submissions API scores 10/10 in our catalog and carries standardized XBRL facts for every US-listed filer - GE alone exposes about 945 us-gaap concepts, with submissions history back to 1994 - but holds nothing on private rounds, because private companies do not file. Companies House Register Search (UK) traces officers and ownership across roughly 5 million live UK companies including private ones, but stays register-shaped and UK-bound. Fortune Global 500 ranks the world's largest companies by revenue once a year - a scale benchmark, not an event stream.
What none of them replicates: funding rounds, investor networks, acquisition events and predictive signals for companies that exist outside the filings-and-registers world. The usual pattern is to stack them - EDGAR and Companies House for structure and financials, Crunchbase for the private-market narrative between filings. The full slate sits on our best industrial-conglomerates datasets ranking.
What should I know before requesting a sample?
Three things worth settling upfront. First, schema provenance: the operator publishes no field-level dictionary, so the ten-field structure above traces to observable surfaces and gets pinned to exact column names and types against your delivered extract before anything downstream depends on it. Second, scope discipline: millions of profiles reward a named universe - bring your watchlist, your sector tags or your competitor set and the sample arrives cut to it. Third, treat the predictive scores as what they are - modeled probabilities riding beside the records, useful as triage, not as verdicts. Get a sample of this dataset scoped to the companies you actually track.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
record_type | enum | Which family the row carries: company profile, funding round, investor, person or acquisition event - one family per row. | funding round |
company_name | string | Canonical name of the company the row concerns; subsidiaries and portfolio members hold their own profiles, so a conglomerate's tree resolves row by row. | (illustrative scale-up) |
company_key | string | Stable opaque identifier binding every row of every family to one company profile, so joins run on identity rather than name matching. | on request |
headquarters_location | string | City and country on the company profile; depth concentrates on US and European venture-backed firms. | on request |
sector_tags | array | Category labels placing the company among industries - industrial automation, energy equipment, advanced manufacturing and their kin. | ["industrial automation"] |
event_date | date | Announcement date on the dated families: funding rounds, acquisitions, leadership hires and product launches. | on request |
amount_raised | number | Disclosed size of a funding round in the currency announced; absent where the round went undisclosed. | on request |
counterparty_name | string | The other side of a relationship row: lead investor on a round, acquirer or target on an acquisition, appointee on a hire. | on request |
role_or_stage | string | Qualifier on relationship rows: seed through late-stage on rounds; board, executive or investor role on people rows. | Series C |
signal_score | number | Predictive layer riding on qualifying company profiles: acquisition-prediction, funding-likelihood and growth-insight readings. | on request |
Record families - what each row type captures
| record family | what it captures |
|---|---|
| Company | Profile spine: identity, description, headquarters, sector tags, operating status |
| Funding round | Stage, disclosed amount, date and the investor roster behind the money |
| Investor / person | Firms and named individuals - founders, executives, board seats, partners - with their roles |
| Acquisition event | Who bought whom, when, joining buyer and target profiles permanently |
| Growth signal | Leadership hires, product launches and the predictive scores derived from the whole stream |
What teams do with it
- Competitive early-warning A rival's new round, a leadership hire or an acquisition rumor surfacing as records reads as market intent long before it reaches a filing.
- Private-target screening Track funding rounds, acquirers and investor networks across private industrial-tech targets back to the late 1990s.
- Account prioritization Recent funding, growth signals and leadership data turn a flat account list into a ranked call sheet.
- Sector landscape sizing Venture activity and emerging-player counts per sector give consultants a defensible market map instead of a anecdote pile.
- Feature enrichment Company-level funding and investor history join onto existing firm records as scoring features for models.
Questions buyers ask
What does the Crunchbase dataset actually contain?
Five record families hanging off company profiles: funding rounds with stage, amount and investor roster; acquisition events linking buyer and target; investor and person records covering founders, executives, board seats and partners; subsidiaries and portfolios; and the growth-signal stream of leadership hires and product launches. Predictive acquisition and funding-likelihood scores ride beside the records on qualifying companies.
Does it cover private companies or only listed conglomerates?
It is built for the companies that never file: private, venture-backed firms sit at the center of the corpus, with public companies profiled as acquirers, investors and context. Inside industrial conglomerates work that means the listed parent comes from the filings layer while the scale-ups orbiting it - targets, suppliers, spin-outs - come from here.
How far back does the funding history go?
Funding events reach from the late 1990s to the present, roughly three decades of private-market activity in one consistent record. That span crosses the dot-com cycle, the 2010s industrial-software wave and the climate-hardware run, which is what makes decade-scale longitudinal analysis possible without reconciling vintages.
How granular does the data get?
Five levels: per company, per funding round, per investor, per person and per acquisition event. Individual deals and individual leadership changes stay separately addressable rather than dissolving into annual aggregates, so a trend can be dated to specific transactions instead of gestured at.
Do the predictive scores ship with the records?
Yes - acquisition-prediction, funding-likelihood and growth-insight scores are computed over the corpus and travel beside the underlying records on qualifying company profiles. Treat them as modeled probabilities for triage rather than verdicts, and confirm the exact delivered form of the score columns against your sample extract.
Can a sample be scoped to my companies?
Name your watchlist, competitor set or sector tags and the sample arrives cut to that universe in exactly the ten-field shape documented above, with the extended structures - subsidiary linkages, investor rollups, the score set - confirmed against the extract before anything recurring is configured.
Notes on this record
- Only private-markets record in the slice Among the 23 primary industrial-conglomerates datasets, this is the sole record covering private-market financing rather than filings, macro accounts or plant operations.
- Catalog score 4/10 - documentation, not depth Scored 4 against a catalog mean of 7.81 because the operator publishes no field-level dictionary. The remedy is procedural: definitions confirmed against a delivered extract before you build.
Datasets that pair with this one
- SEC EDGAR Company Facts & Submissions API Where the listed parent's numbers live - standardized XBRL for every US filer, GE carrying about 945 us-gaap concepts. Pair the two and you see the parent and its private orbit.
- Get a sample of this dataset Samples ship cut to the companies, sectors and record families you name, in exactly the field shape documented above.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.