Diversified Support Services · U.S. Securities and Exchange Commission (EDGAR)

SEC EDGAR Full-Text Search API

Datadory delivers sec edgar full text search api data covering every EDGAR filing since 2001 - tens of millions of indexed documents searchable by keyword and form type, each hit carrying CIK, company name, SIC code, filing date and accession number at individual-document grain, delivered daily, weekly, or hourly.

API, files, or your warehouse. Daily, weekly, or hourly.

Where it covers
United States - all domestic and foreign private issuers that file with the SEC
How far back
Filings from January 1, 2001 through the present, one continuous frame with no gap years
How fine
Individual document level within each filing, with filer-level attributes attached to every hit

What is the SEC EDGAR Full-Text Search API?

Every word every public filer chose to publish, made addressable. SEC EDGAR Full-Text Search API is the search layer under EDGAR - the same index that powers the SEC's own filing-search screens - opened up here as structured rows instead of raw responses. Since January 1, 2001 the index has absorbed essentially the whole filing stream: annual reports, quarterly and interim disclosures, ownership statements, proxy materials, registration statements, and the exhibits that trail behind them. Tens of millions of documents sit inside one schema.

Two things distinguish it from a document dump. First, the grain: hits resolve at individual document level, so a subsidiary list filed as Exhibit 21.1 is its own row, findable separately from the 10-K that carries it. Second, the filer context riding on every hit - CIK, display name with ticker, business and incorporation states, SIC classification, accession number - means a match arrives already identified, not as a pointer you have to chase. Within Datadory's diversified support services shelf this is the primary-record layer everything else annotates, and it scores 9/10 on our quality rubric.

What do sample rows look like?

Three real hits spanning the industry's breadth - a workforce giant, an IT-staffing specialist and a facilities operator, the last reaching back to 2002:

# hit 1 - annual report, workforce services (Feb 2026)
_id           : 0001193125-26-064113:mmi-20251231x10k.htm
display_names : ["ManpowerGroup Inc.  (MAN)  (CIK 0000871763)"]
file_date     : 2026-02-23
form          : 10-K
biz_locations : ["Milwaukee, WI"]
sics          : ["7363"]   # help supply services

# hit 2 - annual report, IT staffing (Mar 2025)
_id           : 0001193125-25-054447:mhh-20241231x10k.htm
display_names : ["Mastech Digital, Inc.  (MHH)  (CIK 0001437226)"]
file_date     : 2025-03-14
form          : 10-K
biz_locations : ["Moon Township, PA"]
sics          : ["8742"]

# hit 3 - annual report, facility services (Dec 2002)
_id           : 0000950149-02-002436:f86502e10vk.htm
display_names : ["ABM INDUSTRIES INC /DE/  (ABM)  (CIK 0000771497)"]
file_date     : 2002-12-17
form          : 10-K
biz_states    : ["CA"]
sics          : ["7340"]   # services to dwellings and other buildings

Read the _id once and the design clicks: accession number, colon, document filename. Every hit therefore resolves back to its exact place inside its parent submission - the row is simultaneously a search result and a citation. The display_names column fuses company name, ticker and CIK into one string, which is why entity resolution takes subtraction rather than fuzzy matching. And the third hit makes the archive argument for itself: a facilities-services annual report from December 2002 sits in the same schema as a February 2026 filing, so a twenty-year panel needs no stitching across format eras. Bracketed arrays arrive typed; concrete cuts ship with your sample.

Which fields does the dictionary define?

Twenty-four verified fields, mapped below with types, definitions and examples checked during the August 2026 research pass - nothing inferred from column names alone. Three groupings carry the analytical weight. The identity block (ciks, display_names, adsh, file_num, film_num) pins each hit to a filer and a submission. The document block (_id, form, root_forms, file_type, file_description, sequence, items) says exactly what kind of paper matched and where it sits in the stack. The context block (file_date, period_ending, biz_locations, biz_states, inc_states, sics) dates the disclosure and locates the filer geographically and industrially.

The distinction worth internalizing is form versus root_forms versus file_type: the parent submission may be a 10-K while the matched document is an exhibit inside it, and conflating the two quietly corrupts any count of "how many annual reports mention X." The sics column is the industry lever - 7363 flags help-supply services and 7340 flags services to dwellings and other buildings, the two codes under which much of this industry files.

Where does coverage sit, and at what grain?

Geography - United States: all domestic and foreign private issuers that file with the SEC. For this industry that means the national staffing majors, regional facilities operators, uniform-rental and security-services firms, plus overseas issuers listing into US markets - one frame covering everyone obliged to disclose.

Temporal - filings from January 1, 2001 through the present, one continuous frame. The 2001 start matters analytically: it reaches back through dot-com-era outsourcing booms, the 2008 labor-market break and the post-2020 remote-work rewrite, all in a single schema, which is why longitudinal language studies start here rather than by splicing decade-old snapshots.

Granularity - individual document level within each filing, with filer-level attributes attached. Scale check: tens of millions of indexed documents, and one industry phrase - "staffing services," restricted to annual reports - alone matched 2,486 filings. Set against the wider catalog, where the average quality score across all 1,744 datasets is 7.81, this record's 9/10 rests on verified definitions and a schema that has not moved underfoot.

How is the data delivered through Datadory?

API, files, or your warehouse. Daily, weekly, or hourly.

Cadence is yours to set and to change: a one-time historical pull for a study, or a warehouse kept current so new filings diff cleanly into yesterday's rows. Deliveries arrive normalized to the dictionary above - typed arrays preserved as arrays, dates typed as dates, the accession-to-document identifier intact as the join path - so nothing needs reconstructing after the fact.

Name the phrases, form types, SIC codes, states or issuer list when you scope the sample and the extract arrives shaped to them. The schema you validate in the sample is the schema you ship against.

Who builds on this data, and for what?

Full-text reach over the primary record earns its keep on specific jobs:

  • Service-industry market mapping - phrase searches over annual reports settle who actually competes in staffing, facility and workplace services, in issuers' own words rather than a third party's taxonomy.
  • Competitive structure reads - subsidiary exhibits (EX-21.1) and charter exhibits (EX-3.x) surface operating brands and ownership chains ahead of announcements.
  • Event studies and announcement windows - separating annual reports from interim disclosures by form type keeps windows clean and replicable.
  • Counterparty and credit screening - a per-issuer filing history shows risk-factor drift and disclosure changes as disclosed, quarter after quarter.
  • Disclosure-language analytics - two decades of annual-report text turns a phrase's rise or fall into a datable series.

Each of these continues naturally on the competitive intel product teams use cases and data scientists use cases pages.

Which personas get the most value?

Market researchers and consultants get competitor sets named by the competitors themselves, with locations attached. Competitive intelligence and product teams read corporate structure out of exhibits before it becomes a headline. Data scientists and analysts get a fixed 24-field schema keyed on CIK and accession number - features on day one, not after a parsing project; see data scientists use cases. Investors and quants build event studies on form-type-separated windows with honest match counts; see investors and quants use cases. Developers and builders wire typed rows into enrichment pipelines keyed on stable identifiers; see developers builders use cases. Journalists, academics and students round out the list with primary-source evidence whose citation metadata arrives pre-resolved.

How does it compare within diversified support services data?

Inside this slice, the neighbors measure different things. QCEW NAICS 56 counts establishments, jobs and wages - the physical footprint of the industry, county by county. OEWS support occupations prices the workforce. Statistics Canada NAICS 56 repeats the exercise for Canada. This record contributes what none of the others hold: the issuers' own narrative, at document grain, back to 2001. Headcount series tell you the industry's size; filing text tells you what the industry says about itself - pricing pressure, contract mix, automation bets - while it happens.

The head-to-head with the Canadian statistical view is worked through in SEC EDGAR Full-Text Search API vs Statistics Canada NAICS 56 data tables. Used together, establishment counts size the market and filing text explains its direction.

What should I know before requesting a sample?

Four honest caveats. First, the 2001 floor: nothing earlier exists in full text, so pre-2001 work falls back to abstracts and indexes outside this dataset. Second, hits carry metadata about the document, not the document's prose - the row identifies who filed what, when, and where it sits; the underlying text rides separately on request, and your sample confirms both halves together. Third, match totals arrive flagged exact (eq) or lower-bound (gte) - build counts off the relation column or a deduplication sweep will inflate them. Fourth, industry cuts follow SIC, not NAICS or GICS: 7363 and 7340 cover much of this industry, but map codes before joining to modern sector schemes. None of these bite unexpectedly; they ship flagged against the analysis you plan to run.

Which notes and datasets pair with it?

Notes that pair well with this page:

Source: U.S. Securities and Exchange Commission (EDGAR).

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - the 24 verified fields riding on every hit (definitions and examples)
FieldTypeDefinitionExample
_idstringDocument identifier combining the accession number and the filed document filename, separated by a colon.0001193125-25-054447:mhh-20241231x10k.htm
_indexstringSearch-index name identifying the corpus a hit came from; always edgar_file here.edgar_file
_scorenumberRelevance score for the match against the query.-
tookintegerMilliseconds the server took to execute the query.-
timed_outbooleanWhether the query exceeded the server's time budget.-
hits.total.valueintegerTotal number of matching filings across all pages.2486
hits.total.relationstringWhether total.value is exact (eq) or a lower bound (gte).-
cikstextArray of 10-digit Central Index Key numbers identifying the filer(s).["0000871763"]
display_namestextFiler display strings combining company name, ticker symbol and CIK.["ManpowerGroup Inc. (MAN) (CIK 0000871763)"]
file_datedateDate the filing was accepted into EDGAR.2026-02-23
period_endingdateEnd date of the reporting period covered by the document.-
formstringSpecific form type of this document within the filing.10-K
root_formstextRoot form type(s) of the parent submission.-
file_typestringDocument type within the submission, e.g. 10-K, EX-21.1, EX-3.28.EX-21.1
file_descriptionstringHuman-readable description of the document.FORM 10-K
adshstringAccession number of the parent filing, used to locate the original document set.0001193125-26-064113
file_numtextSEC file number(s) assigned to the registrant.-
film_numtextFilm number from the SEC's document microfilm records.-
biz_locationstextPrincipal business location(s) of the filer.["Milwaukee, WI"]
biz_statestextState code(s) of the filer's business locations.-
inc_statestextState(s) where the filer is incorporated.-
sicstextStandard Industrial Classification code(s) for the filer; 7363 = help supply services, 7340 = services to dwellings and other buildings.["7363"]
itemstext10-K item numbers included in the filing, when applicable.-
sequenceintegerSequence position of the document within its submission.-

Questions buyers ask

What does the SEC EDGAR Full-Text Search API dataset contain?

Structured rows for every document EDGAR has indexed since January 1, 2001 - tens of millions of hits, each carrying accession number and filename, CIK, company display name with ticker, form and root-form types, document type, filing and period dates, business and incorporation states, SIC codes and 10-K item numbers, at individual-document grain.

How far back does the data go?

To filings accepted from January 1, 2001 onward - the earliest year full text entered the index. The frame is continuous from there, so a two-decade-plus panel builds without splicing formats; the sample row for a December 2002 facilities-services annual report shows the earliest era reading identically to a 2026 filing.

What is the difference between form, root_forms and file_type?

Root form is the parent submission's type (say, a 10-K); file_type is the specific document inside it that matched (an EX-21.1 subsidiary list, for instance); form is the matched document's own type. Keeping the three apart is what stops an exhibit count being reported as an annual-report count.

Can results be narrowed to one industry or state?

Yes - the sics column carries Standard Industrial Classification codes (7363 for help-supply services and 7340 for services to dwellings and other buildings cover much of this industry), while biz_states, biz_locations and inc_states locate the filer operationally and legally. Name the codes and states when you scope a sample and the cut comes back pre-filtered.

Does a hit include the filing's actual text?

No - a hit is a fully identified pointer: who filed it, which document matched, when, and where it sits inside the submission. The underlying document travels separately on request, and because the accession-to-filename identifier resolves exactly, text and metadata reunite losslessly in your warehouse.

Can I get a sample scoped to my companies or phrases?

Yes. Name the phrases, form types, SIC codes, states or issuer list - say, every annual report naming a particular service line since 2015 - and the sample returns rows in exactly the schema above, cut to that scope, with the field dictionary and match-count flags documented alongside.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing