Glossary
text classification benchmark
text classification benchmark is frozen dataset with fixed train/validation/test splits used to score text-classification models across named tasks - TweetEval ships eleven configurations over seven Twitter tasks, one labeled row per tweet, with license terms set per subset. Datadory's catalog of 1,744 datasets documents it directly in TweetEval Benchmark.
What is text classification benchmark?
A frozen dataset with fixed train/validation/test splits used to score text-classification models across named tasks - TweetEval ships eleven configurations over seven Twitter tasks, one labeled row per tweet, with license terms set per subset.
In this catalog it appears concretely: - TweetEval Benchmark — "Unified seven-task Twitter classification benchmark from Cardiff NLP with fixed train/val/test splits covering irony, hate, offensive, stance, emoji, emotion and sentiment".
Why does text classification benchmark matter when choosing a dataset?
A label on a listing is not a deliverable. Teams that license on the strength of a product-name match routinely find the shipped files cover a narrower slice than they assumed, and backfilling history afterwards costs more than the license ever did.
The failure mode is concrete: the label appears in a listing, the delivered files tell a different story, and the gap surfaces mid-project when fixing it is most expensive.
You rarely have to take a vendor's word for it. 82.4% of the 1,744 datasets Datadory catalogs are free to access, and TweetEval Benchmark lets you inspect the real artifact before any budget is committed.
How do you evaluate text classification benchmark in a data source?
Treat every claim of this attribute as testable:
- Open TweetEval Benchmark and confirm its record — "Unified seven-task Twitter classification benchmark from Cardiff NLP with fixed train/val/test splits covering irony, hate, offensive, stance, emoji, emotion and sentiment" — against the files you actually receive.
- Pin down update cadence in writing. Across this catalog, 22.6% of 1,744 datasets refresh daily and 64 still arrive only through a manual request form, so ask exactly how fresh each release is.
- Check whether definitions are verified at all. Field definitions are verified for 1495 of 1,744 datasets (85.7%), and any source you license should meet that bar.
- Price the delivery route before the license. In this catalog bulk download is the most common access method (725 datasets) ahead of official APIs (574), and 379 sources still require scraping — a maintenance cost that lands on you, not the vendor.
See the term applied to real records: interactive-media-services data.
Related terms
Adjacent concepts worth reading next: - social network graph dataset - upstream platform terms
Frequently asked questions
What is an example of text classification benchmark?
TweetEval Benchmark is the clearest example in this catalog. Its record states: "Unified seven-task Twitter classification benchmark from Cardiff NLP with fixed train/val/test splits covering irony, hate, offensive, stance, emoji, emotion and sentiment". Across all 1,744 datasets Datadory averages a quality score of 7.81 out of 10, so a named example can be weighed rather than trusted blindly.
Is data described as "text classification benchmark" free to use?
Treat access and permission separately. 82.4% of the 1,744 datasets in this catalog are free to access, but 235 are freemium and 61 are paid outright, so confirm both the price and the license on the exact distribution before building on it.
Datasets containing this field
Datasets containing text classification benchmark
6 datasets carry text classification benchmark in the catalog. Open one, count the fields, judge for yourself.
APNIC Labs Measurements & Dashboards
Domain Name Industry Brief & DNIB Data
quarter / reporting period · total_domain_name_registrations · com_net_domain_name_base …+3 more
Hurricane Electric BGP Toolkit
IANA Root Zone Database
OONI Explorer – Global Internet Censorship Measurements
Every listing shows the field dictionary, sample rows, and coverage before you commit. API, files, or your warehouse. Daily, weekly, or hourly.
Get sample rows