Diversified Financial Services · U.S. Securities and Exchange Commission
SEC EDGAR Full-Text Search API (EFTS)
Datadory delivers sec edgar full text search api efts data covering the full text of every EDGAR filing and exhibit since 2001 - roughly two million submissions a year resolving to tens of millions of indexed documents, each carrying eighteen typed fields from accession number and filer identity through form type, SIC codes and Regulation items, delivered daily, weekly, or hourly.
API, files, or your warehouse. Daily, weekly, or hourly.
What is the SEC EDGAR Full-Text Search API (EFTS)?
It is the difference between knowing what a company filed and knowing what a company said. SEC EDGAR Full-Text Search API (EFTS) is the search layer under the EDGAR full-text search experience: an Elasticsearch-built index whose coverage begins in 2001, when electronic filing became mandatory, and runs to every submission landing today. Unlike a registry view, it reads words - the filing body plus every attachment and exhibit.
That single design choice explains the record's job on the diversified financial services shelf. A phrase query for "financial services" restricted to Form 10-K between 2024-01-01 and 2026-08-20 returned more than ten thousand hits. The same machinery, narrowed to a tighter phrase over 2026, returned exactly 184 documents - precise enough to read by hand. And because SIC-coded results put securitization vehicles beside operating companies, the auto-receivables and mortgage trusts that file servicer compliance statements every spring sit one filter away from the bank holding companies everyone already watches.
Get a sample of this dataset and Datadory returns the slice you name - phrases, form types, SIC codes, windows - as populated records rather than a search box.
What do sample rows look like?
Three verified hits from a Form 10-K servicer-filing probe, exactly as they arrive:
_id : 0001853620-25-000098:e331_mbenserv23.htm
form : 10-K
display_names : Mercedes-Benz Auto Receivables Trust 2023-2 (CIK 0001993977)
ciks : 0001993977
file_date : 2025-03-26
period_ending : 2024-12-31
sics : 6189
biz_locations : Farmington Hills, MI
_id : 0000950170-24-039213:ck0001570440-ex33_3.htm
form : 10-K
display_names : Sequoia Mortgage Trust 2013-4 (CIK 0001570440)
ciks : 0001570440
file_date : 2024-04-01
period_ending : 2023-12-31
sics : 6189
biz_locations : Mill Valley, CA
_id : 0001853620-24-000096:e331_daimserv.htm
form : 10-K
display_names : Daimler Trucks Retail Trust 2023-1 (CIK 0001971264)
ciks : 0001971264
file_date : 2024-03-28
period_ending : 2023-12-31
sics : 6189
biz_locations : Fort Worth, TXRead the trio together and the texture shows. All three are annual reports from SIC 6189 asset-backed securities vehicles - a Mercedes-Benz auto receivables trust, a Sequoia mortgage trust, a Daimler truck retail trust - each hit pointing at a specific document inside the submission rather than the filing as a whole. The first _id ends in mbenserv23.htm, a servicer compliance exhibit; the second lands on an ex33_3.htm certification. Same form type, three different documents, each independently addressable - which is what turns a search result into pipeline input rather than reading homework.
What fields does the dataset include?
Eighteen fields ride on every hit, and every definition below was confirmed against live responses during the August 2026 research pass rather than copied from a spec page. Six identify the document and its filing - _id, adsh, ciks, display_names, file_num and film_num - with the compound identifier doing the heavy lifting: accession number plus file name resolves any hit to its original document in one step.
Three place the hit in time: file_date for acceptance, period_ending for the period actually being reported on, and sequence for the document's position inside its submission. Six classify the disclosure: form and root_forms at filing level, file_type and file_description at document level - the pair that separates a main 10-K body from the EX-33 certification beside it - plus sics for the filer's industry code and items for Regulation items such as material-agreement disclosures. Three locate the filer: biz_locations, biz_states and inc_states, where the incorporation column doubles as a quick securitization tell.
Datadory-side enrichment folds under request a sample: ticker resolution onto CIKs, decoded SIC labels, form-family rollups, lead-registrant entity rollups and document text retrieved alongside the metadata.
What does coverage look like across geography, time and granularity?
Geography - every domestic and foreign private issuer that files with the SEC. For diversified financial services work that means the full spread: bank and thrift holding companies, finance subsidiaries, broker-dealer parents, insurers, and the securitization trusts that exist only to file.
Temporal - 2001 to the present, anchored to mandatory electronic filing. Anything filed with EDGAR since then is searchable in full text, which makes the archive deep enough for longitudinal studies of disclosure language while staying complete enough that absence of hits means something.
Granularity - one indexed record per document or exhibit within each filing submission. A single trust's 10-K might contribute several records: the main document, its certifications, its servicer statements. Tens of millions of records resolve from roughly two million-plus filings a year, with result pages running ten to a hundred hits.
Set against the wider Datadory catalog - average quality score 7.81 across all 1,744 datasets - this record scores 10/10, one of the few to reach the ceiling, carried by fully verified field documentation and unmatched archive breadth.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Channel and cadence are settings, not projects. Files suit a research team loading the full 2001-forward archive once and building a disclosure-language corpus on top of it; feeds suit monitoring products that watch named issuers or SIC codes for new mentions; warehouse delivery suits quant and compliance teams joining hit records against holdings, prices or credit files in SQL. Every delivery travels with the eighteen-field dictionary above, the sample rows and the coverage profile mapped to the phrases, forms and windows you named - so the schema you see in the sample is the schema you ship against.
Who uses this data, and for what?
- Equity and credit analysts - sweep risk-factor and liquidity language across an industry before earnings season, catching shifts in wording that precede shifts in guidance (investors & quants use cases).
- Data scientists & ML engineers - assemble labeled disclosure corpora filtered by form, SIC code and date, the substrate for classification and extraction models (data scientists use cases).
- Developers & data-product builders - wire disclosure alerts and screening tools into products on typed, documented fields (developers builders use cases).
- ABS and structured-finance desks - track servicer compliance filings from auto, mortgage and equipment trusts each spring, by SIC code rather than by hand-built watchlists.
- Compliance and competitive-intel teams - isolate material-agreement disclosures through the Regulation items column and watch counterparties' filings as they land.
- Journalists, academics & students - anchor reporting and research in the filings themselves, every quote resolvable to an accession number (journalists academics use cases).
Which personas get the most value?
Analysts and quant researchers come first: full-text discovery is the front door to every other filings workflow, and the eighteen-field hit record means what they find arrives structured enough to rank, deduplicate and join. Structured-finance and surveillance teams get the only practical route into the servicer-filing tier of the market, where hundreds of trusts file near-identical annual compliance documents that no headline will ever surface. Data scientists and ML engineers inherit a corpus whose typing never surprises them and whose filters double as labels. Journalists, academics and students cite primary documents rather than aggregators' summaries. Developers and product builders wrap discovery into monitoring tools because the hit record already carries the identity, date and classification columns those tools need.
How does it compare to alternatives in its slice?
Within the diversified financial services shelf, this record owns words: what issuers actually said, document by document, exhibit by exhibit. Neighbors own different jobs. FDIC Bank Data & Statistics quantifies insured institutions - branch networks, call-report financials, deposit concentration; the direct comparison lives at vs FDIC Bank Data & Statistics. Federal Reserve Economic Data Portal (Board Releases) carries policy and monetary series. Nasdaq Stock Screener API covers quotes and screenable fundamentals, and OpenSanctions Dataset handles risk screening. On the structured side of the SEC family, the XBRL financial statements record carries the tagged numbers this index surrounds.
When the question is who said what - and in which exhibit - nothing else on the shelf answers. When the question is a balance-sheet figure or a branch count, the neighbors are faster.
What should I know before requesting a sample?
Three things worth having in hand.
First, hits describe documents, not companies. One issuer's year produces many records - main documents, certifications, servicer statements - so expect many-to-one shapes when rolling hits up to entities, and use the entity rollup enrichment if a holding-company view is the goal.
Second, broad queries hit a display ceiling. Unbounded popular phrases return counts marked as greater-than-or-equal at ten thousand, which is a display convention rather than a wall: narrowing the window or the form set enumerates cleanly, as the 184-document probe above shows. Scope first, enumerate second.
Third, verification went to the field level. All eighteen definitions were confirmed against live responses during the August 2026 research pass, and because exhibits are indexed in their own right, exhibit-only records carry their own file_type and file_description values rather than inheriting the parent document's. Name the scope you need and the sample settles the rest empirically.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
_id | string | Document identifier combining EDGAR accession number and file name. | 0001853620-25-000098:e331_mbenserv23.htm |
adsh | string | EDGAR accession number of the parent filing. | 0001853620-25-000098 |
ciks | string | Central Index Keys for the filers associated with the document. | 0001993977 |
display_names | text | Human-readable entity names with CIK suffixes. | Mercedes-Benz Auto Receivables Trust 2023-2 (CIK 0001993977) |
file_num | string | SEC registration or file number of the filing. | 333-266303-03 |
film_num | string | SEC film number for the filing. | 25771619 |
file_date | date | Date the filing was accepted into EDGAR. | 2025-03-26 |
period_ending | date | Reporting period end date of the filing. | 2024-12-31 |
sequence | integer | Sequence position of the document inside the filing submission. | 2 |
form | enum | EDGAR form type of the filing. | 10-K |
root_forms | string | Root form classification(s) for the document. | 10-K |
file_type | string | Specific document or exhibit type within the filing. | EX-33 |
file_description | text | Description line for the document within the filing. | MERCEDES-BENZ FINANCIAL SERVICES USA LLC, AS SERVICER. |
sics | string | Standard Industrial Classification codes assigned to the filer. | 6189 |
items | string | Regulation items reported on the filing. | 1.01 |
biz_locations | string | Principal business location(s) of the filer. | Farmington Hills, MI |
biz_states | string | State(s) where the filer does business. | MI |
inc_states | string | State(s) of incorporation. | DE |
Questions buyers ask
What does the SEC EDGAR Full-Text Search API (EFTS) cover?
The complete text of every EDGAR filing and its attachments since 2001: annual and quarterly reports, current reports, adviser and fund filings, and the exhibits beside them. Each match returns as a structured record carrying filer identities, dates, form type, SIC codes and Regulation items.
How far back does the history run?
To 2001, the year electronic filing became mandatory for EDGAR registrants. Everything filed electronically since then is in the index, which makes 2001 the practical starting line for any whole-market disclosure study.
Do exhibits get searched, or only main filing documents?
Both. The index covers the filing body plus every attachment and exhibit, which is precisely why servicer agreements, certifications and material-contract exhibits surface here when they would never appear in a structured-facts extract.
Which fields identify each hit?
The _id combines the accession number and file name into one pointer, backed by adsh, the ciks array and display_names. Together they take a hit from search result to original document without ambiguity, even inside combined filings with multiple co-filers.
How precise can a query be?
Narrow enough to enumerate by hand. A phrase query for financial services limited to Form 10-K across a two-and-a-half-year window returned more than ten thousand hits; narrowed to a tighter phrase and a single year, the same probe returned exactly 184 documents.
Can a sample be cut to specific forms, industries or date windows?
Yes. Name the phrase or topic, the form types, the SIC codes and the date range, and the sample arrives shaped to that scope with the full eighteen-field dictionary attached - the schema you see in the sample is the schema you ship against.
Notes on this record
- Scored at the top of the catalog Datadory scores this record 10/10 against a catalog mean of 7.81 - one of the small set of records out of 1,744 cataloged to reach the ceiling, on the strength of fully verified field documentation and archive depth.
- Sample policy Samples ship in the exact eighteen-field schema above, cut to the phrases, form types, SIC codes and windows you name; the enrichment columns under additional fields on request confirm at sample time.
Datasets that pair with this one
- FDIC Bank Data & Statistics Words versus registers: this index finds what issuers disclosed; FDIC statistics quantify the insured institutions beneath them. The head-to-head in [vs FDIC Bank Data & Statistics](/compare/sec-edgar-full-text-search-api-efts-vs-fdic-bank-data-statistics) settles which unit of analysis you need.
- SEC EDGAR XBRL Financial Statements API XBRL company facts carry the tagged figures; this index carries everything the tags leave out - narrative, risk factors and the exhibits where agreements actually live.
- CIK identifier Every hit carries filer CIKs, the stable key that holds through renames, mergers and reorganizations - which is what lets discovery here feed entity-resolved pipelines elsewhere.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.