Education Services · Kaggle

Kaggle — curated education datasets collection

Datadory delivers kaggle curated education datasets collection data covering student performance, retention prediction and international schooling comparisons: the 155 MB US Education Unification Project of NCES-derived district enrollments back to 1986, a 79 MB World Bank education mirror, India-specific schooling sets and more - each versioned, usability-scored and delivered daily, weekly, or hourly.

API, files, or your warehouse. Daily, weekly, or hourly.

Where it covers
Global mix by construction: United States-focused NCES-derived collections, India-specific schooling sets, and international comparison series from World Bank-derived mirrors covering national systems worldwide.
How far back
Vintages frozen at capture rather than flowing series: the Unification Project's NCES-derived district enrollments span school years 1986 through 2018, comparison sets carry the years their maintainers published, and every row quotes the latest version stamp at capture.
How fine
Three grains share the shelf - student-level records in performance and outcomes collections, district or institution level in administrative mirrors, country level throughout the international comparisons.

What is the Kaggle curated education datasets collection?

The community-published pool of education data on Kaggle, gathered into one cataloged record instead of scattered across search results. A live capture during the August 2026 research pass surfaced the recurring anchors: U.S. Education Datasets: Unification Project - 155,201,337 bytes of NCES-derived district enrollment and assessment tables tracing school years back to 1986 - alongside a 79,282,075-byte World Bank education statistics mirror, Education & Career Success, Education in India, Indian School Education Statistics, Cost of International Education, Education Inequality Data, Global Education and World Educational Data.

Every upload is versioned, stamped with its latest revision, sized in bytes and graded 0-1 for usability on documentation and file structure, with community notebooks attached wherever analysis already exists. Quality is deliberately mixed: documented mirrors of official statistics sit beside one-off CSVs with thin provenance, so this record ships the provenance signals per row instead of averaging them away. See the Kaggle source profile for how the platform sits inside our wider catalog of 1,744 datasets.

What do sample rows look like?

One row per matched upload, typed identically whether the underlying set runs 79 MB or a single CSV. The leading records, exactly as captured:

# one row per matched upload; leading records as captured in the August 2026 pass
title        : U.S. Education Datasets: Unification Project
ref          : noriuk/us-education-datasets-unification-project
bytes        : 155201337              # roughly 148 MB, the flagship set
members      : hundreds of NCES-derived CSV tables
span         : school years 1986 through 2018

title        : Education Statistics
ref          : organizations/theworldbank
bytes        : 79282075               # roughly 76 MB, World Bank education mirror

title        : Education & Career Success
version      : 2025-08-07             # latest revision stamp at capture

title        : Education in India
scope        : India-specific schooling indicators

title        : Education Inequality Data
scope        : cross-country inequality measures

# member-file header inside the flagship set:
column       : Grade 1 Students - American Indian/Alaska Native - male [District] 2017-18

Notice what the rows settle that raw listings leave open: byte sizes for capacity planning, revision stamps so nobody argues about vintages, and the usability grade that separates documented mirrors from quick uploads before anyone commits a pipeline. Member-file columns travel separately - the Unification Project alone unpacks hundreds of NCES-derived tables whose headers read like the per-grade, per-subgroup enrollment column quoted above.

Which fields does the field dictionary define?

Seven fields make up the confirmed core of every delivered row: two identify (title, ref), two size and grade the set (totalBytes, usabilityRating), one timestamps it (lastUpdated), one carries the uploader-declared usage classification (terms_posture), and one exposes the column vocabulary of multi-table member files (member_columns). Definitions were checked against live captures during the August 2026 pass rather than reconstructed from memory, which is why anything beyond the core folds under the request note instead of being padded with guessed columns.

What does coverage look like across geography, time and granularity?

Geography - a global mix by construction: United States-focused NCES-derived collections, India-specific schooling sets, and international comparison series from World Bank-derived mirrors that put national systems side by side.

Temporal - vintages frozen at capture rather than flowing series. The deepest run belongs to the Unification Project, whose NCES-derived district enrollments span school years 1986 through 2018; comparison sets carry whatever years their maintainers published, and every row quotes the stamp of the latest version current when the snapshot was taken.

Granularity - three grains share the shelf: student-level records in the performance and outcomes collections, district or institution level in administrative mirrors such as the enrollment tables, and country level throughout the international comparisons.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

You pick the channel and the cadence; the field dictionary above travels unchanged through all three. Rows arrive flattened - one row per matched upload, member-file schemas joined or shipped alongside - so joining enrollment counts against your own outcome data is a join statement rather than a reconstruction project. Cadence changes are a settings conversation, not a re-integration, and a sample cut to your named outcomes and geographies comes first either way.

Who uses this data, and for what?

  • Student-outcome and retention model prototyping - benchmark churn and completion predictors against real community baselines before budgeting a bespoke build.
  • District enrollment and demographic panels - fold the Unification Project's hundreds of NCES-derived tables into one panel running 1986 through 2018 instead of stitching files by hand.
  • International education comparisons - World Bank-derived mirrors put attainment and spending for national systems into a comparable frame at country grain.
  • India education-market sizing - Education in India and Indian School Education Statistics give subcontinent demand studies a starting corpus.
  • Teaching corpora and analyst onboarding - small self-contained sets that let a classroom or a new hire practice the whole modeling loop on real data.
  • Fitness-for-purpose audits - usability grades and provenance signals turn 'can we rely on this' into a filterable column before anything reaches production.

Which personas get the most value?

Ranked by relevance in our persona tagging:

  1. Data Scientists & ML Engineers (relevance 3/3) - the natural home: versioned, usability-graded collections built for prototyping outcome and retention models.
  2. EdTech Product & Growth Teams (2/3) - enrollment trend panels and India-specific sets for market sizing and feature discovery.
  3. Market Researchers & Consultants (2/3) - cost-of-international-education figures and country comparisons behind client deliverables.
  4. Journalists, Academics & Students (2/3) - citable starting points with provenance flags keeping claims honest about what a community upload can support.
  5. Developers & Builders (1/3) - a fixed row schema to parse once while powering education search and directory features.
  6. Investors & Quants (1/3) - sector demand backdrop of cohort sizes and participation trends for diligence screens.

How does it compare within education services data?

Inside our education services shelf, this record plays the breadth-and-speed role. Where the ONS Education & Childcare collection is the UK's register of record and the OECD Education at a Glance tables are the harmonized cross-country frame, this collection is the fastest route to a working dataset - dozens of ready sets spanning student performance, district enrollments and international comparisons, graded for documentation quality on arrival. The trade is authority for velocity: registers carry the legally mandated counts, this carries the prototyping corpus, and the two join cleanly once a model graduates to production. For establishment-level directory work, the data.gouv.fr French education collection shows what an official register looks like at school grain.

What should you know before requesting a sample?

Three things. First, scope wins: name the outcomes, geographies and grade bands you care about and the sample arrives cut to them, in the exact schema shown above. Second, expect heterogeneity and get ahead of it - usability grades and provenance signals ship on every row precisely because a community marketplace mixes thorough mirrors with weekend projects. Third, usage classifications are declared per upload; the exact terms for the sets in your sample are confirmed alongside it, so fitness-for-purpose decisions rest on paperwork rather than assumption. From there, api, files, or your warehouse. daily, weekly, or hourly.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - seven confirmed-core fields in the kaggle curated education datasets collection record
fieldtypedefinitionexample
titlestringDisplay title of the uploaded set exactly as published on its card.U.S. Education Datasets: Unification Project
refstringOwner-plus-slug identifier naming the uploader and the set inside the platform's catalog.noriuk/us-education-datasets-unification-project
totalBytesintegerCompressed byte size of the set, so capacity planning happens before acquisition rather than after.155201337
usabilityRatingnumberPlatform's 0-1 usability score grading documentation quality and file structure on every upload.Score between 0.0 and 1.0, higher means better documented
lastUpdateddatetimeTimestamp of the most recent version of the set, letting two teams quote the same vintage unambiguously.2025-08-07T11:03:17.003Z
terms_posturestringUploader-declared usage classification carried on every card, normalized onto the row so fitness-for-purpose screening happens up front; exact wording confirms with your sample.Declared per upload, reported per row
member_columnstextColumn headers inside multi-table member files - district name, state, and per-grade per-year per-subgroup counts in the NCES-derived sets.Grade 1 Students - American Indian/Alaska Native - male [District] 2017-18

Kaggle curated education datasets collection - product specification

AttributeValue
IndustryEducation Services
PlatformKaggle community dataset marketplace
Flagship setU.S. Education Datasets: Unification Project - 155,201,337 bytes compressed
Largest mirrorWorld Bank education statistics - 79,282,075 bytes compressed
Deepest historical runNCES-derived district enrollments, school years 1986 through 2018
GrainStudent, district or institution, or country depending on member set
Delivered formatsCSV, XLSX, JSON or SQLite underneath, flattened to your schema
Quality score6/10 against a catalog mean of 7.81 across 1,744 datasets
Field definitionsVerified against live captures, August 2026 research pass

What teams do with it

  • Student-outcome and retention model prototyping Benchmark churn and completion predictors against real community baselines before budgeting a bespoke build.
  • District enrollment and demographic panels Fold the Unification Project's hundreds of NCES-derived tables into one panel running 1986 through 2018 instead of stitching files by hand.
  • International education comparisons World Bank-derived mirrors put attainment and spending for national systems into a comparable frame at country grain.
  • India education-market sizing Education in India and Indian School Education Statistics give subcontinent demand studies a starting corpus.
  • Teaching corpora and analyst onboarding Small self-contained sets that let a classroom or a new hire practice the whole modeling loop on real data.
  • Fitness-for-purpose audits Usability grades and provenance signals turn reliability questions into filterable columns before anything reaches production.

Questions buyers ask

What does one delivered row of kaggle curated education datasets collection data contain?

One row per matched upload: display title, owner-plus-slug reference, compressed byte size, the 0-1 usability grade, the latest version timestamp and the uploader-declared usage classification - with member-file column schemas joined alongside for multi-table sets such as the NCES-derived enrollment project.

Which datasets anchor the collection?

Eight named members recur across captures: U.S. Education Datasets: Unification Project at 155,201,337 bytes, the 79,282,075-byte World Bank education statistics mirror, Education & Career Success, Education in India, Indian School Education Statistics, Cost of International Education, Education Inequality Data and World Educational Data.

How far back does the US district enrollment data go?

The Unification Project's NCES-derived tables span school years 1986 through 2018, carrying per-grade, per-year, per-subgroup counts for districts alongside state identifiers - one of the longest continuous administrative runs available anywhere in the education shelf.

Is quality consistent across the collection?

No, and the record says so plainly: documented mirrors of official statistics sit beside one-off CSVs with minimal provenance. Every upload carries a 0-1 usability grade on documentation and file structure, and deliveries repeat those signals per row so triage happens in the data rather than in a review meeting.

Does the collection cover India and international comparisons?

Both. Education in India and Indian School Education Statistics cover the subcontinent directly, Cost of International Education prices study-abroad markets, and World Bank-derived mirrors plus Global Education and World Educational Data put national systems side by side at country grain.

Can a sample be scoped to my use case?

Yes. Name the outcomes, geographies and grade bands and the sample arrives in exactly the schema shown above, extended across whichever slice answers the question. Delivery runs through API, files, or your warehouse on a daily, weekly, or hourly cadence, with per-set details documented alongside the sample.

Notes on this record

  • Provenance Compiled from live platform captures during the August 2026 research pass - definitions map one-to-one to the documented record rather than inferred conventions.
  • Uneven by design A community marketplace mixes thorough mirrors of official statistics with weekend projects; deliveries carry the usability grade and provenance signals per row so triage precedes production.
  • Breadth first, depth second This record is the fast, wide prototyping layer; registers such as the NCES and UNESCO digests hold deeper authoritative counts on their own slices and join cleanly alongside.
  • The grade travels with the row The 0-1 documentation-and-structure score ships on every delivered row, turning 'is this any good' into a filter instead of an afternoon of inspection.
  • Scored against the catalog Datadory rates this record 6/10 against a catalog mean of 7.81 across 1,744 datasets - breadth is the asset here, per-set depth varies.
  • Sample policy Samples ship in the exact schema shown above, scoped to your named outcomes, geographies and grade bands; per-set usage terms confirm with the sample.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing