Hugging Face Datasets - Financial Services

Datadory delivers hugging face datasets financial services data as one governed feed over the Hub's finance shelf: roughly 19 exact-match repositories plus 100-plus broader financial collections spanning expert-annotated news sentences, tweet sentiment pairs, filing excerpts and 100-million-row market bundles, delivered daily, weekly, or hourly.

What is the Hugging Face Datasets - Financial Services collection?

A shelf with real depth and no gatekeeper - which is simultaneously its charm and its hazard. Hugging Face Datasets - Financial Services is Datadory's governed feed over the machine-learning community's finance collections on the Hugging Face Hub: at the August 2026 research pass, about 19 repositories answer a 'financial services' query exactly, and the wider 'financial' searches clear 100. Row counts run from under a thousand to beyond a hundred million, and the named residents give the shelf its texture.

Three carry most of the weight. takala/financial_phrasebank holds 4,840 financial-news sentences annotated by 16 people with adequate background knowledge in financial markets, split into four agreement tiers so strictness becomes a filter rather than a compromise. zeroshot/twitter-financial-news-sentiment and its topic sibling supply text-label pairs harvested from finance Twitter. And defeatbeta/yahoo-finance-data bundles prices, earnings-call transcripts, SEC-filing text, dividend events and Treasury yields into a single repository banded above 100 million rows, with about 137,896 recorded downloads - the collection's workhorse.

Get a sample of this dataset and we return catalog rows plus payload records cut to the repositories your project actually names.

What do rows from this collection look like?

Two layers, deliberately kept apart. The catalog layer is one observation per repository and never changes shape; payload records arrive in whatever columns their uploader chose. From the August 2026 research pass:

# catalog observation - one row per repository
id            : defeatbeta/yahoo-finance-data
size_tag      : 100M<n<1B
downloads     : 137896          # cumulative at the August 2026 research pass

# payload observations - classification pairs from inside a repository
repo          : zeroshot/twitter-financial-news-sentiment
text          : $BYND - JPMorgan reels in expectations on Beyond Meat
label         : 0               # class index; readable mapping lives on the card

repo          : zeroshot/twitter-financial-news-sentiment
text          : $CCL $RCL - Nomura points to bookings weakness at Carnival and Royal Caribbean
label         : 0

# card excerpt - the shelf's most-cited sentiment benchmark
id            : takala/financial_phrasebank
excerpt       : collection of 4840 sentences ... annotated by 16 people
                with adequate background knowledge on financial markets
size_tag      : 1K<n<10K

Read the counter honestly: 137,896 downloads accumulates across versions, making it a popularity signal and nothing more - several heavily downloaded repositories turn out to be mirrors whose upstream sourcing cannot be traced. Note too what the label column does: 0 is an index, and the human-readable class meaning sits on each dataset card, which is exactly the decode work Datadory ships finished. (Catalog observations from the August 2026 research pass; payload records ship with your sample.)

Which fields does the field dictionary define?

Eight fields carry the catalog layer plus the commonest payload columns, defined below with examples taken from live records. id keys onto any model-asset registry you already maintain; task_categories and size_categories make the shelf queryable by intended job and declared heft; downloads and lastModified timestamp each repository's standing.

Everything else is per-repository and therefore folded rather than promised: the reuse declaration each uploader attaches to their card, the inner column sets behind the market-tabular and filing-excerpt repositories, and card configuration blocks get pinned with concrete examples when your sample names its targets.

Where does coverage run, and at what grain?

  • Geography: a global contributor base with English-dominant content - US and EU market material concentrates because that is where the harvested sources live - while country-level financial inclusion subsets push reach across Africa and Asia.
  • Temporal: author-deposited snapshots rather than a maintained series. Observed last-modified stamps span early 2025 through August 2026, each repository carrying its own version history, and the matched set shifts as uploads land - any count on this page is a photograph, not a portrait.
  • Granularity: one catalog observation per repository at the wrapper, descending to per-record grain inside them - sentence, tweet, filing excerpt or market observation - within independently versioned repositories.

Set against the wider Datadory catalog - where the average quality score across all 1,744 datasets is 7.81 - this slice scores 6/10: carried by breadth, named anchor corpora and fully-typed catalog columns, held back by unvetted community provenance and payload schemas that change from one repository to the next.

How is the data delivered through Datadory?

API, files, or your warehouse. Daily, weekly, or hourly.

Pick the channel your stack already speaks and set the cadence to match the decision being fed. Bulk files suit research benches loading whole corpora overnight; structured payloads suit products scoring headlines as they cross the tape; warehouse delivery suits teams joining these corpora to their own price and fundamentals tables without a second integration. The catalog layer, the common payload columns and the inner schemas of your chosen repositories travel unchanged across all three - cadence is a settings conversation, not a re-engineering project.

Who builds on this collection?

Five jobs it does better than any single official publisher.

Train and grade finance language models. Expert-labelled news sentences, tweet sentiment pairs and domain corpora sit one shelf apart, so tuning material and evaluation material arrive in one catalog filter instead of a procurement cycle.

Calibrate analyst-tone scoring. financial_phrasebank's four agreement tiers let you test a sentiment model against progressively stricter human consensus before trusting any vendor's number on a results-day headline.

Prototype document-understanding pipelines. OCR'd annual-report table collections and filing-derived corpora give extraction and retrieval work raw material that official publishers ship only as polished PDFs.

Screen the long tail before commissioning annotation. The discovery layer's modality, format, size-band and task filters turn 'what already exists' into a query - far cheaper to learn in August than after budgeting a labeling effort.

Stand up a market-data sandbox. yahoo-finance-data's single bundle - prices, transcripts, filings, dividends, Treasury yields - stocks a prototype environment with the joins already cohabiting one repository.

Which personas get the most value?

Data scientists and ML engineers get the shortest idea-to-experiment path in finance NLP: labelled text, typed catalog columns and versioned repositories. See data scientists use cases.

Investors and quants get the calibration set for tone: expert-agreement-tiered sentiment labels to hold their own scoring against, plus transcript and filing text for event studies. See investors quants use cases.

Developers and builders get stable repository identifiers and typed catalog rows to wrap into products, demos and retrieval features. See developers builders use cases.

Sales and growth teams selling into financial institutions get transcript and filing corpora that reveal what a prospect firm actually talks about - lending, payments, treasury - before the first call. See sales growth teams use cases.

What should I know before requesting a sample?

Four things, stated plainly.

First, quality varies by uploader. Nobody curates this shelf; discipline ranges from institutional releases to weekend projects, and download counts will not tell you which is which. Sample preparation vets the repositories you name and flags the ones whose upstream sourcing cannot be traced.

Second, legacy packaging lingers. Some older repositories shipped with loading scripts that current tooling declines to execute; confirm a converted format exists before building anything on one, which is part of what a scoped sample settles.

Third, counts are photographs. Roughly 19 exact matches versus 100-plus wider ones reflects query width as much as inventory, and the matched set moves as uploads land - your sample states its own census date.

Fourth, the dictionary covers the wrapper plus the common payload columns. Inner schemas differ by design between a tweet pair and a hundred-million-row market bundle, and they get defined row by row once your sample fixes the repository list.

Which notes pair with this dataset?

Notes that pair well with this page:

  • NASDAQ Stock Screener API - the exchange's own listing universe, one row per security with sector and industry labels; the structured counterpart to community-built market bundles (profile).
  • FINRA Data Catalog and API - regulator-grade broker-dealer transparency, where compliance work goes when community provenance is not enough (catalog).
  • BIS Data Portal - Global Banking and Financial Statistics - central-bank statistics with institutional certainty behind every cell (portal).
  • Diversified financial services data hub - the full pooled view of the industry, from ticker universes to bank regulators (hub).
  • Best diversified-financial-services datasets - the ranked shortlist this slice sits within (ranking).

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - catalog layer plus common payload columns, verified against the August 2026 research pass (examples from live records)
fieldtypedefinitionexample
idstringFull repository identifier in author/name form within the Hub namespace - the join key onto any model-asset registry you already maintain.takala/financial_phrasebank
task_categoriesenumCard tag declaring the machine-learning job a repository was built for; text-classification and sentiment-classification sit among the finance values.text-classification
size_categoriesenumRow-count band declared on the dataset card, running from under 1K records to over 100M.100M<n<1B
downloadsintegerCumulative download counter carried on the repository - a popularity signal, explicitly not a quality measure, since versions stack.137896
lastModifieddatetimeTimestamp of the repository's most recent revision; observed values on this shelf span early 2025 through August 2026.-
texttextCommon column across the finance NLP repositories: the headline, tweet or snippet being classified.$BYND - JPMorgan reels in expectations on Beyond Meat
labelintegerEncoded class index travelling alongside text in classification corpora; the human-readable mapping lives on each dataset card.0
sentencetextThe annotated financial-news sentence column in expert-labelled sentiment benchmarks.-
Additional fields on request-Payload structure beyond the common three: per-repository column sets behind the market-tabular and filing-excerpt repositories, card configuration blocks and citation metadata, plus the reuse declaration each uploader attaches to their card - each pinned with examples when your sample names its repositories.-

Coverage at a glance

dimensioncoverage
GeographyGlobal contributor base, English-dominant content; country-level financial inclusion subsets extend reach across Africa and Asia
TemporalAuthor-deposited static snapshots rather than a maintained series; observed last-modified stamps span 2025 through August 2026, each repository keeping its own version history
GranularityOne catalog observation per repository, descending to per-record grain inside them - sentence, tweet, filing excerpt or market observation

Questions buyers ask

What does one record of hugging face datasets financial services data contain?

A catalog observation per repository - identifier, task tags, size band, download counter and last-modification timestamp - joined to payload records from inside the repositories you select: labelled news sentences, tweet sentiment pairs, filing excerpts or market observations. Payload columns are pinned to your named repositories when the sample is cut.

How many financial services datasets are in the collection?

About 19 repositories answer a 'financial services' query exactly at the August 2026 research pass, while broader 'financial' searches clear 100. Individual repositories run from under 1,000 rows to beyond 100 million, anchored by takala/financial_phrasebank, the zeroshot twitter sentiment pair and defeatbeta/yahoo-finance-data.

What makes financial_phrasebank useful for sentiment work?

Its 4,840 sentences were annotated by 16 people with adequate background knowledge in financial markets and released in four agreement tiers, so you choose how much inter-annotator consensus a training or evaluation label requires. That graded strictness is rare anywhere in finance text and unmatched at zero procurement effort.

What does the yahoo-finance-data repository hold?

A single bundled workspace of market material: prices, earnings-call transcripts, SEC-filing text, dividend events and Treasury yields, banded above 100 million rows with roughly 137,896 recorded downloads. Because the pieces cohabit one repository, prototype joins arrive pre-assembled instead of negotiated across five vendors.

How current are the numbers on this page?

They are photographs taken during the August 2026 research pass. Observed last-modification stamps span early 2025 through August 2026, the matched repository set shifts as uploads land, and download counters accumulate across versions - delivered rows carry their own observation date so anything you build states its census.

How does this differ from institutional finance statistics?

Institutional publishers - BIS, FINRA, Treasury - deliver curated, authoritative series with slow publication cycles. This collection delivers community-built training material at arrival speed with no curation gate. They complement rather than compete: models train here, decisions cite there.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing