Electronic Components · Hugging Face

Electronic Components Supply Chain & Risk Dataset

Datadory delivers electronic components supply chain risk dataset data covering 791 MPNs linked to manufacturer, country of origin, category, datasheets and multi-distributor supplier offers, scored with China-origin risk flags and reason codes. Delivered daily, weekly, or hourly as an API, files, or straight into your warehouse.

API, files, or your warehouse. Daily, weekly, or hourly.

Where it covers
Global distributor coverage - DigiKey, Mouser, TME, Conrad, Win Source and independent brokers observed; manufacturers concentrated in the US, EU, Japan and China
How far back
Point-in-time snapshot of 791 records; prices and stock are as-of-collection values, not live quotes
How fine
One row per MPN, with nested per-seller offer detail inside each record

What is the electronic components supply chain risk dataset?

It is a component-level risk panel that answers one question per row: how exposed is this part? Each of the 791 records is keyed by manufacturer part number and stacks four layers on top of that key. An identity layer names the manufacturer and headquarters country. A catalog layer adds an Octopart-style category label, a free-text technical description and datasheet references. A commerce layer lists distributor offers with name, stock, price and SKU. And a risk layer scores the part: a boolean China-origin flag, a confidence score attached to that origin judgment (observed running roughly 0.7 to 0.95), a 0-to-1 risk score, and the reason codes behind it.

The risk layer is what makes the dataset more than another parts catalog. Three reason codes do the explaining: chinese_origin, critical_category and limited_suppliers. Because every flagged row carries both a score and the codes that produced it, the verdicts are auditable - a sourcing team can disagree with a score and see exactly why it was assigned.

The fourth layer is depth. Each record nests a full Nexar/Octopart enrichment block: manufacturer homepage, best datasheet link, and a sellers array of companies with offers carrying SKUs, price points and inventory levels. One flat JSON file holds all four layers, so a single read gets you from part number to priced, stocked, risk-scored reality.

What do sample rows look like?

Identity, commerce and risk layers for one record, then two comparison rows from elsewhere in the score range - exactly as they land in a delivery:

mpn=ESP8266EX   mfr=Espressif Systems   country=China   is_chinese=true
category=RF Receivers, Transceivers   confidence=0.95
risk_score=0.60   risk_factors=[chinese_origin]

suppliers=[{name=DigiKey, stock=6433, price=$1.60, sku=1904-1001-1-ND}, ...]
nexar_data={bestDatasheet: url, sellers[] company=DigiKey, offers[] inventoryLevel, ...}

mpn=DFR1063     mfr=DFRobot             risk_score=0.25  risk_factors=[critical_category]
mpn=GPY         mfr=Pycom               risk_score=0.10  risk_factors=[limited_suppliers]

Three rows, three different triggers. The Espressif Wi-Fi SoC flags purely on origin, with the classifier stating 0.95 confidence. The DFRobot part trips critical_category without an origin flag. The Pycom module sits at the low end because its trigger is scarcity - limited_suppliers - not geography. Scores observed in sampled rows run 0.1 to 0.6, with nulls where nothing fired.

What fields does the dataset include?

Twelve documented fields split across the same four layers: identity (mpn, manufacturer, manufacturer_country), catalog (category, description, datasheets), commerce (suppliers, with the nexar_data block deepening it into per-seller offers) and risk (is_chinese, confidence_score, risk_score, risk_factors).

The dictionary below is the verified core. The nested seller detail inside nexar_data flattens into pipeline-ready columns - those fold out on request rather than being promised blind.

How wide is the coverage?

Geography: global distributor coverage - DigiKey, Mouser, TME, Conrad, Win Source and independent brokers all appear in the suppliers arrays - against manufacturers headquartered in the US, EU, Japan and China.

Manufacturer concentration skews analog and embedded. Analog Devices leads with 195 rows, Texas Instruments follows at 184, STMicroelectronics at 100, then Espressif Systems (63), GigaDevice (48), Microchip (41), Seeed Studio (28) and Adafruit Industries (22). Eight manufacturers account for the large majority of the panel, which makes it a focused instrument for analog/embedded sourcing risk rather than a shallow index of everything.

Origin mix: only 14 of 791 rows carry the China-origin flag - roughly 1.8 percent - even though Chinese-headquartered makers such as Espressif and GigaDevice hold over 100 rows between them. That gap between headquarters country and flag rate is itself informative about how the classifier weights evidence.

Temporal: a point-in-time snapshot of 791 records. Prices and stock are as-of-collection values, so treat them as structure rather than live quotes - cadence accrues the time dimension on your side.

Granularity: one row per MPN with nested per-seller offer detail. No rollups, no category averages - every downstream cut aggregates from the atomic unit.

Set against the wider Datadory catalog - 1,744 datasets, average quality score 7.81 - this slice scores 7/10: exceptional structural richness per row, held back mainly by its static snapshot nature and undocumented scoring methodology.

How is the data delivered?

Pick the channel your team already works in: REST endpoints for live lookups against a specific MPN, flat files sized for overnight warehouse loads, or a direct pipe into Snowflake, BigQuery or Redshift. Cadence is yours to set - daily catches design-cycle churn, weekly suffices for quarterly risk reviews, hourly suits monitoring dashboards.

Every delivery ships with the full field dictionary, sample rows for validation, and a schema that holds steady between refreshes. The nested nexar_data block arrives pre-flattened if you want it relational, or intact if your stack prefers documents.

Who uses this data, and for what?

  • Supply-chain risk managers screen bills of materials for China-origin exposure part by part. Because every flag ships with reason codes and a confidence score, screening results survive audit - the explanation travels with the verdict.
  • Component and sourcing engineers run second-source checks before design freeze: how many distributors carry the part, at what price and stock levels, and whether limited_suppliers fires. A part with one offer object is a different procurement decision than one with six.
  • Competitive intelligence teams map manufacturer concentration per category - who owns the supply base, where second sources are missing, and which categories concentrate hardest.
  • Data scientists use the flag-plus-confidence structure as labeled training material for origin-classification models, and the nested seller arrays as ground truth for pricing and inventory estimation.
  • Investors and analysts quantify component-level sourcing exposure beneath the company-level disclosure layer, particularly for hardware businesses selling into categories flagged critical.

More on how these teams apply the slice lives on the data scientists, competitive intel product teams and investors & quants pages.

How does it compare within electronic components data?

Inside this industry slice, three datasets answer three different questions. The Nexar API (Octopart) is the live commerce engine - multi-distributor offers, lifecycle status and specs across the full Octopart database, but it carries no risk scoring of its own. FindChips' Aggregated Price & Inventory Search consolidates current price and stock across authorized and independent distributors - breadth of quote coverage, again without a risk layer. The TME catalog contributes 1.5 million products with documented parameters and files from one major European distributor.

This dataset is the fourth leg: judgment. It is the slice that says not just what a part costs and who stocks it, but whether sourcing it carries geographic or availability risk - and shows its work via reason codes. The practical stack joins this risk layer onto Nexar or FindChips pricing by MPN, which turns three separate feeds into one answerable question: for a given part, what does it cost, who has it, and how risky is depending on it.

What should I know before requesting a sample?

Three things, stated up front.

First, this is a snapshot, not a stream. The 791 records reflect one collection pass - prices and stock are as-of-collection values. Teams needing a time dimension start a cadence immediately and let repeated captures build the curve.

Second, the scoring methodology is undocumented upstream. No published methodology explains how risk_score, confidence_score or the reason codes are computed, and the snapshot date of the underlying price data is unstated. Treat the risk layer as a strong prior to be validated against your own sourcing rules, not as a regulatory determination - we say this plainly because a risk feed you cannot interrogate is worse than none.

Third, the verified core dictionary covers twelve fields; the deeper seller-level detail inside nexar_data is confirmed and flattened when we cut your sample, which is also where column naming gets locked down for your pipeline. Name the MPNs, manufacturers or categories you care about and the sample comes back shaped to them.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - Electronic Components Supply Chain & Risk Dataset (one row per MPN)
fieldtypedefinitionexample
mpnstringManufacturer part number of the component.ESP8266EX
manufacturerstringName of the component manufacturer.Espressif Systems
categorystringOctopart-style product category label.RF Receivers, Transceivers
descriptiontextShort free-text technical description of the part.IC ESP8266, SDIO 2.0, SPI, UART, QFN32, RF switch, balun, 24dBm PA, DCXO, PMU/Espressif
manufacturer_countrystringCountry where the manufacturer is headquartered.China
datasheetstextList of datasheet document references for the part.["ESP8266EX-Espressif-Systems-datasheet.pdf"]
suppliersarrayList of distributor/reseller offer objects with keys name, stock, price and sku.[{"name": "DigiKey", "stock": 6433, "price": 1.6, "sku": "1904-1001-1-ND"}]
is_chinesebooleanBoolean flag indicating whether the manufacturer is judged to be China-based.true
confidence_scorenumberFloat confidence attached to the origin classification (observed range roughly 0.7-0.95).0.95
risk_scorenumberFloat supply-chain risk score between 0 and 1; null when no risk factors were triggered.0.6
risk_factorsenumList of triggered risk reason codes such as chinese_origin, critical_category or limited_suppliers.["chinese_origin"]
nexar_dataobjectNested dict of Nexar/Octopart enrichment: mpn, manufacturer{name, homepageUrl}, category{name}, shortDescription, bestDatasheet{url}, sellers[{company{name}, offers[{sku, prices[{price}], inventoryLevel}]}].{"sellers": [{"company": {"name": "DigiKey"}}]}
Additional fields on request-Full nexar_data sub-structure (per-offer price ladders, inventory levels, seller SKUs) flattened into pipeline-ready columns, plus any derived columns cut during sample preparation - definitions and examples ship with your sample.-

Sample rows - three records as they land, spanning the risk-score range observed in the data

mpnmanufacturercountryrisk_scorerisk_factors
ESP8266EXEspressif SystemsChina0.60chinese_origin
DFR1063DFRobot-0.25critical_category
GPYPycom-0.10limited_suppliers

Questions buyers ask

What fields does the electronic components supply chain risk dataset include?

Twelve core fields per record: mpn, manufacturer, category, description, manufacturer_country, datasheets, suppliers, is_chinese, confidence_score, risk_score, risk_factors and a nested nexar_data block. The suppliers field carries distributor offer objects with name, stock, price and sku; nexar_data nests seller companies with offers, price points and inventory levels.

How many components does the dataset cover?

Exactly 791 records, one per manufacturer part number. Manufacturer concentration runs heavy: Analog Devices accounts for 195 rows, Texas Instruments 184, STMicroelectronics 100, Espressif Systems 63, GigaDevice 48, Microchip 41, Seeed Studio 28 and Adafruit Industries 22 - a deep analog-and-embedded panel rather than a shallow index.

How is supply-chain risk scored?

Each record carries a boolean China-origin flag, a confidence score attached to that judgment (observed around 0.7 to 0.95), a 0-to-1 risk score and triggered reason codes: chinese_origin, critical_category or limited_suppliers. Observed scores span 0.1 to 0.6 with nulls where nothing fired. The upstream methodology is unpublished, so treat scores as auditable priors, not rulings.

Does the data include distributor pricing and stock?

Yes. The suppliers field lists distributor offers with name, stock, price and SKU - DigiKey, Mouser, TME, Conrad, Win Source and independent brokers all appear. The nested Nexar/Octopart block goes deeper with per-seller companies, multiple offers, individual price points and inventory levels. Values are as-of-collection snapshots rather than live quotes.

Which manufacturers and countries appear most?

Analog Devices (195 rows), Texas Instruments (184) and STMicroelectronics (100) dominate, with Espressif Systems, GigaDevice, Microchip, Seeed Studio and Adafruit filling out the top eight. Manufacturers concentrate in the US, EU, Japan and China, while distributor coverage is global - only 14 of 791 rows carry the China-origin flag despite that footprint.

Who uses electronic components supply chain risk data?

Supply-chain risk managers screen BOMs for geographic exposure with auditable reason codes. Sourcing engineers check second-source depth before design freeze. Competitive-intelligence teams map manufacturer concentration per category, and data scientists train origin classifiers on the flag-plus-confidence structure while using seller arrays as pricing ground truth.

Datasets that pair with this one

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing