Datadory notebook
China Origin Component Risk Score: How to Screen a Bill of Materials
Datadory delivers electronic components data covering the China origin component risk score end to end: 791 part-number rows carrying an is_chinese flag, a risk_score observed running 0.1 to 0.6, confidence between 0.7 and 0.95, and reason codes chinese_origin, critical_category and limited_suppliers - joined on the MPN to live price and stock layers, delivered daily, weekly, or hourly.
1,744 datasets. Pick your catch.
What is a China origin component risk score?
A China origin component risk score condenses several sourcing questions about one part number into a single value: where the manufacturer sits, whether the part's category counts as critical, and how many suppliers can actually fulfil orders. In the Electronic Components Supply Chain & Risk Dataset the score never travels alone - each row pairs it with an is_chinese boolean origin flag, a confidence_score observed between 0.7 and 0.95, and a risk_factors list of reason codes. Three codes appear across the published rows: chinese_origin, critical_category and limited_suppliers.
Observed risk_score values span 0.1 to 0.6, with nulls on rows where nothing triggered. The spread matters operationally. A 0.6 tagged chinese_origin reads nothing like a 0.1 whose only trigger was limited_suppliers: one is a geography verdict, the other is a scarcity count. No published documentation explains the weighting behind the decimal, so treat the number as triage input that orders human attention - not as an auditable rating.
Where does a published China-origin score actually exist?
Exactly once in the electronic-components slice: the Electronic Components Supply Chain & Risk Dataset, cataloged by Datadory from source Hugging Face. It holds 791 rows, one per manufacturer part number, each stacking four layers on the same key. Identity: mpn, manufacturer, manufacturer_country. Catalog: category label, free-text technical description, datasheet references. Commerce: a suppliers array deepened by a nested seller-detail block down to individual offers with SKUs, price points and inventory levels. Risk: is_chinese, confidence_score, risk_score and risk_factors.
Coverage notes set expectations honestly. Manufacturer origins span the US, EU, Japan and China; distributor offers come from DigiKey, Mouser, TME, Conrad, Win Source and independent brokers. Only 14 of the 791 rows carry is_chinese=true, so the panel screens a curated analog-and-embedded shortlist rather than the whole market - the Nexar API (Octopart) alone reaches millions of components, FindChips Aggregated Price & Inventory Search spans roughly 41,473 public MPN landing points, and the LCSC Electronics Catalog lists about eight million product pages per storefront language. Within its own scope the record is unusually well documented: Datadory scores it 7 out of 10 against a cross-catalog mean of 7.81 across all 1,744 datasets, with field definitions verified against published sample rows.
Which parts actually score highest in the panel?
Manufacturer concentration is heavy and easy to quote: Analog Devices holds 195 of the 791 rows, Texas Instruments 184, STMicroelectronics 100, Espressif Systems 63, GigaDevice 48, Microchip 41, Seeed Studio 28 and Adafruit Industries 22. Three sampled rows show how differently parts reach the same score scale:
mpn=ESP8266EX mfr=Espressif Systems country=China is_chinese=true
category=RF Receivers, Transceivers confidence=0.95
risk_score=0.60 risk_factors=[chinese_origin]
mpn=DFR1063 mfr=DFRobot risk_score=0.25 risk_factors=[critical_category]
mpn=GPY mfr=Pycom risk_score=0.10 risk_factors=[limited_suppliers]ESP8266EX is the only sampled row pairing a high score with a China origin flag: 0.6 at 0.95 confidence on chinese_origin. DFR1063 reaches 0.25 through critical_category, and GPY sits at 0.1 through limited_suppliers - neither trigger involves geography at all. That is exactly why the reason codes deserve reading beside the number: the same decimal can mean origin exposure, category criticality or plain supplier scarcity.
Ranking a whole bill of materials on the raw score alone therefore misorders work. Sort by risk_score descending within each reason-code group - chinese_origin rows first, then critical_category, then limited_suppliers - and attach the manufacturer_country value so a reviewer sees geography and scarcity in one screen. Across 791 rows that sort is a one-pass job, and Datadory returns it already applied to whatever subset you name.
How do scored rows become a screened bill of materials?
You send the part-number list; the screen comes back as typed rows. The mechanics, in the order they matter:
- Match on the MPN. Both sides normalize to uppercase with package-code suffixes stripped before joining; matched rows return the five working columns - manufacturer_country, is_chinese, confidence_score, risk_score and risk_factors. The manufacturer part number is the join key across the entire slice.
- Read nulls as unscored, not safe. With only 14 flagged rows out of 791, most joins come back clean by absence. The delivered sort puts nulls last, score descending, confidence alongside so a reviewer can weigh how loudly each verdict speaks.
- Take the seller detail pre-flattened. The nested seller block carries full offer objects with price-break arrays and inventory levels; delivered relationally it arrives as pipeline-ready columns instead of a nesting problem that explodes row counts.
- Get the verdict and its evidence together. Every scored row ships with the category label, description and supplier list that sit behind it, so a challenged score can be argued with on the merits rather than taken on faith.
Which live layers confirm the verdict?
The score orders attention; four live layers decide whether acting on it is actually possible. All join on the same MPN key:
The Nexar API (Octopart) contributes multi-distributor depth - per-seller offers, quantity-break pricing, inventory levels, lifecycle status and factory lead days on the same part record. FindChips Aggregated Price & Inventory Search spreads the same question across authorized, independent and industrial/MRO channels, with alternates and watch alerts riding alongside. The LCSC Electronics Catalog brings Shenzhen-side volume - per-warehouse stock rows and USD/CNY price ladders skewed toward Asian domestic brands Western catalogs underrepresent. The TME Electronic Components Catalog & API attaches forward-looking supply dates and waiting periods from one major European distributor's own warehouse positions.
Together they answer the question a flag cannot: if this part needs a second source, does one exist on a shelf near your build, at what price and lead time? The second table below lays the stack out layer by layer.
What are the limits of the score before you act on it?
Four limits bound what the number can tell you.
Breadth. 791 rows against catalogs of millions means absence of a flag is not evidence of safety - the panel is a shortlist, and unlisted parts simply carry no verdict. Methodology comes second: no published documentation explains how risk_score, confidence_score or the three reason codes are computed, and sampled scores run only 0.1 to 0.6 with some rows null. Third, freshness: the scored panel is a single-pass snapshot whose prices and stock are as-of-collection values, while the live layers beside it reflect current state - which is precisely why the verdict and the confirmation should arrive in the same delivery. Fourth, semantics: because triggers mix origin, category and supplier count, a mid-score part may need a second source rather than a redesign.
Use the score to order human attention, then confirm supplier reality against live rows before committing to a redesign or a dual-source qualification. The electronic component price history data walkthrough shows how repeated capture timestamps those live quotes into a series worth reviewing monthly.
Who screens bills of materials with China-origin data?
Five teams run on these rows, ranked by how directly a scored panel settles their day job:
- Supply-chain risk managers screen BOMs for geographic exposure part by part, and because every flag ships with reason codes and a confidence score, the result survives audit - the explanation travels with the verdict.
- Component and sourcing engineers run second-source checks before design freeze: how many distributors carry the candidate part, at what price and stock. A part with one offer object is a different procurement decision than one with six, which is what
limited_suppliersfires on. - Competitive intelligence and product teams map manufacturer concentration per category - who owns the supply base and where second-source gaps sit across RF, MCU, sensor and power lines. Their workflow continues at competitive intel product teams.
- Data scientists and quant teams calibrate origin-classification models on the flag-plus-confidence structure and treat the nested seller arrays as ground truth for pricing work. Their workflow: data scientists.
- Investors and analysts quantify semiconductor sourcing exposure at the component level, beneath company-level disclosure. Their workflow: investors & quants.
Why get China-origin screening through Datadory?
Because the hard part was never locating a score - it is reconciling a static judgment layer with live commerce layers that define stock, currency and granularity differently, then keeping the join alive as catalogs churn.
Every delivery ships with the field dictionary unchanged and sample rows for validation. Name the MPNs, manufacturers or categories you care about and request a sample scoped to your own bill of materials: the schema in the sample is the schema you ship against.
Where to go next
Start with the electronic components data guide, the pillar that scores and groups all eight records in the slice, or browse the ranked best electronic-components datasets. For adjacent workflows, see the electronic component MPN search database comparison on the resolution side of the same join and the electronic component price history data walkthrough on turning quotes into series. The full pooled view lives on the electronic components data hub, and the Electronic Components Supply Chain & Risk Dataset record stands behind every figure above.
| MPN | Manufacturer | risk_score | risk_factor | confidence_score | Reading |
|---|---|---|---|---|---|
| ESP8266EX | Espressif Systems | 0.6 | chinese_origin | 0.95 | Origin-flagged; the highest sampled score in the panel |
| DFR1063 | DFRobot | 0.25 | critical_category | not published | Category criticality, not origin |
| GPY | Pycom | 0.1 | limited_suppliers | not published | Supplier scarcity, not origin |
| Layer | Record | Grain | What it adds to the screen |
|---|---|---|---|
| Verdict | Electronic Components Supply Chain & Risk Dataset | One row per MPN with nested per-seller offer detail | is_chinese flag, 0.1-0.6 risk score, 0.7-0.95 confidence and reason codes chinese_origin, critical_category, limited_suppliers |
| Live quote depth | Nexar API (Octopart) | One record per MPN, nested to per-seller, per-offer breaks | Inventory levels, quantity-break pricing, lifecycle status and factory lead days on the same part record |
| Channel spread | FindChips Aggregated Price & Inventory Search | One result row per distributor offer per MPN | Low/high unit-price range per channel, Authorized/Independent/MRO buckets, alternates and watch alerts |
| Asia-side volume | LCSC Electronics Catalog | One row per LCSC part code | Per-warehouse stock rows and USD/CNY price ladders skewed toward Asian domestic brands Western catalogs underrepresent |
| First-party lead time | TME Electronic Components Catalog & API | One row per product symbol | Typed availability statuses with waiting periods and named supply dates from one major European distributor's own warehouse |
Pick up where this leaves off
Every one of these ships with sample rows before you commit to anything.
Electronic Components Supply Chain & Risk Dataset
Nexar API (Octopart) Data
FindChips Aggregated Price & Inventory Search
LCSC Electronics Catalog
TME Electronic Components Catalog & API
symbol · mpn · country …+6 more
Want rows instead of a pitch? Name the datasets.
API, files, or your warehouse. Daily, weekly, or hourly.
Get a sampleQuestions worth asking
Is there a dataset with a China-origin risk score per component?
Yes - the Electronic Components Supply Chain & Risk Dataset, cataloged by Datadory from source Hugging Face, ships exactly 791 MPN rows carrying is_chinese, risk_score (observed 0.1-0.6), confidence_score (0.7-0.95) and risk_factors reason codes: chinese_origin, critical_category and limited_suppliers. Datadory delivers it daily, weekly, or hourly.
How is the China-origin risk score calculated?
The publisher documents no formula. Observed evidence shows three trigger codes - chinese_origin, critical_category and limited_suppliers - sampled scores from 0.1 to 0.6, nulls where nothing triggered, and confidence values between 0.7 and 0.95. ESP8266EX scores 0.6 at 0.95 confidence on chinese_origin. Treat the score as a strong prior that orders attention, not a ruling.
Which parts in the panel are flagged as Chinese-origin?
Only 14 of the 791 rows carry is_chinese=true. The rest of the panel concentrates on US, EU and Japanese manufacturers - Analog Devices leads with 195 rows, then Texas Instruments (184), STMicroelectronics (100), Espressif Systems (63), GigaDevice (48), Microchip (41), Seeed Studio (28) and Adafruit Industries (22).
Can the risk score be combined with live stock and pricing?
Yes, by joining on the MPN, which keys every record in the slice. Nexar/Octopart rows contribute per-seller offers, inventory levels and lifecycle status; FindChips adds alternates and watch alerts; LCSC serves USD/CNY price ladders; TME exposes waiting periods and supply dates. The scored panel supplies the verdict while the live layers confirm it.
Can the screen run against my own bill of materials?
That is the default shape of a delivery. Send your part-number list and the matched rows return with the five working columns - manufacturer_country, is_chinese, confidence_score, risk_score, risk_factors - sorted score descending within each reason-code group, nulls last, seller detail pre-flattened. Delivered through API, files, or a warehouse write, daily, weekly, or hourly.