Datadory notebook

Superstore Sales Dataset Csv Download Data: Dataset Structure and Field Coverage

1,744 datasets. Pick your catch. Every guide here is built on what the catalog can actually prove.

1,744 datasets. Pick your catch.

What columns come in the Superstore train.csv?

Eighteen columns cover the full order-line story, and Datadory verified each field definition from the CSV header and observed values rather than documentation:

Column groupFieldsExample value
IdentifiersRow ID, Order ID1; CA-2017-152156
DatesOrder Date, Ship Date (day-first)08/11/2017; 11/11/2017
FulfilmentShip ModeSecond Class
CustomerCustomer ID, Customer Name, SegmentCG-12520; Claire Gute; Consumer
GeographyCountry, City, State, Postal Code, RegionHenderson; Kentucky; 42420; South
ProductProduct ID, Category, Sub-Category, Product NameFUR-BO-10001798; Furniture; Bookcases; Bush Somerset Collection Bookcase
MeasureSales (US dollars per line)261.96

How big is the Office Supplies slice of the dataset?

Office Supplies dominates the file: 5,909 of 9,800 rows, or about 60 percent, against Furniture at 2,078 and Technology at 1,813. Regionally the lines spread West (3,140), East (2,785), Central (2,277) and South (1,598) across United States shipping addresses, so a four-region by three-category cross-tab - twelve cells, all populated - falls straight out of the header row.

Coverage runs from 3 January 2015 to 30 December 2018: exactly four calendar years of daily-dated transactions, static since publication. That window is long enough for monthly seasonality decomposition and annual comparisons, and short enough that every model trained on it inherits a pre-2019 retail world.

The category's weight inside the file is itself the reason forecasters keep it: Office Supplies lines are small-ticket, high-frequency purchases, which produces a steadier monthly series than furniture's lumpy big-ticket orders. If your exercise needs noise, model Furniture; if it needs a stable seasonal signal, the 5,909-row subset delivers it.

How do you go from ZIP to a working forecast in six steps?

The file loads fast enough that a first forecast lands inside an hour:

That last point is the honest boundary: the dataset remains the classic forecasting benchmark because it is small, clean and fully labeled, not because it predicts anything current.

What can you join to the Superstore CSV when four years is not enough?

Three complements extend the file without changing its schema:

Joins are fuzzy by nature. Match Office Depot SKUs to Superstore products on brand and name tokens and expect partial overlap, since one side is a fictionalized 2015-2018 sample and the other is a live chain's current shelf. For the market series, align on month and note that NAICS category definitions will not map one-to-one onto the file's own three categories.

The superstore sales dataset CSV download at a glance (as of August 2026)
AttributeValue
Rows x columns9,800 x 18, one row per order line
Order date range3 January 2015 to 30 December 2018
Category splitOffice Supplies 5,909; Furniture 2,078; Technology 1,813 rows
Region splitWest 3,140; East 2,785; Central 2,277; South 1,598
Compressed size491,942 bytes (about 2.1 MB uncompressed)

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Office Services & Supplies United States

Kaggle - Superstore Sales Dataset

Office Services & Supplies United States retail assortment

Office Depot / OfficeMax Product Catalog (Structured)

14 verified fields · plus additional fields on request …+11 more

Automotive Retail United States

U.S. Census Monthly Retail Trade Survey - Motor Vehicle and Parts Dealers

NAICS Code · Kind of Business · Month column (e.g. 'Jul. 2026') …+6 more

Commercial Printing United States national series, with state and…

US BLS Public Data API — PPI, CES & OES for Printing

status · responseTime · message …+10 more

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

How many rows are in the Superstore dataset?

9,800 order-line rows by 18 columns. The Office Supplies category holds 5,909 rows, Furniture 2,078 and Technology 1,813, spread over four US regions - West (3,140), East (2,785), Central (2,277) and South (1,598). Coverage spans order dates from 3 January 2015 to 30 December 2018.

Can I use the Superstore dataset commercially?

Carefully. The Kaggle page declares GPL 2 under Kaggle's own terms, while the underlying Tableau sample circulates in copies under commercial delivery terms, Apache 2.0 and commercial delivery terms. Internal analysis and teaching are uncontroversial; for shipping the data in a product, prefer a commercial delivery terms mirror such as bravehart101/sample-supermarket-dataset or verify terms first.

Does the Superstore dataset include profit or quantity columns?

No. This version carries a single measure, Sales in US dollars per order line - there is no Profit, Quantity or Discount column. Those fields appear in other Superstore copies on Kaggle, so check a variant's column list before planning margin analysis on this file.