eCommerce Behavior Data from Multi-Category Store (Apparel Segment)

Datadory delivers ecommerce behavior data from multi category store apparel segment data: clickstream events - views, cart adds, cart removals and purchases - from one large multi-category online store, October 2019 through April 2020, nine documented fields per row including product, brand, price, permanent user ID and session ID. Filtered to the apparel taxonomy, it is a ready-made browse-to-purchase funnel.

What is eCommerce Behavior Data from Multi-Category Store (Apparel Segment)?

One large online store, seven months, every click recorded. Between October 2019 and April 2020 the REES46 Open CDP project captured four event types - view, cart, remove_from_cart and purchase - from a single multi-category e-commerce operation whose taxonomy includes apparel among its top categories. Each row carries nine fields, which means a complete browse-to-purchase funnel can be reconstructed per shopper without joining a second table.

Apparel teams usually inherit purchase logs that start too late: they know what sold, not what was looked at and abandoned. This dataset starts at the view. Because user_id is permanent and user_session resets after a pause, sessions stitch themselves back together - a shopper who returns three days later is still the same user_id. That is the raw material for cart-abandonment models, recommendation training sets and price-sensitivity studies on clothing and footwear.

Datadory delivers this as a production feed: filter rows where category_code begins with apparel and you have the clothing-and-footwear segment isolated from the rest of the store's traffic.

What does a sample row look like?

Every event is one flat row - no nesting, no event-envelope JSON to unwrap. A single apparel view event:

event_time   : 2019-10-01 00:00:00 UTC
event_type   : view
product_id   : 1005115
category_code: apparel.shoes
price        : 21.47

The full row adds four more columns (category_id, brand, user_id, user_session). The first midnight row of the archive happens to be a shoe view at $21.47, which is as tidy an illustration as funnel data gets: timestamp, intent, SKU, taxonomy position, price point - five facts in five fields, and eight more months of them behind it.

What fields does the dataset include?

Nine documented fields, verified against the source card. The dictionary below is the complete schema for every event row; nothing is inferred and no field requires a lookup table.

What does coverage look like across geography, time and granularity?

Geography - events come from a single large multi-category online store operated by REES46; the source documentation does not state a country, so treat locale and currency as inferable from prices and behavior rather than confirmed metadata.

Temporal - seven consecutive months, October 2019 through April 2020, delivered as month-by-month slices so pre-holiday, holiday and post-holiday apparel behavior stay separable.

Granularity - one row per clickstream event. Four event types (view, cart, remove_from_cart, purchase) multiply out to roughly 285 million users' events across the full corpus, which is why funnel work at session level holds up statistically even after filtering to the apparel slice.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Who uses this data, and for what?

  • Cart-abandonment analysis - cart versus remove_from_cart versus purchase gives the exact leak points in the apparel funnel, by brand and price band rather than in aggregate.
  • Recommender-system training - permanent user_id plus product interaction history supplies implicit-feedback pairs for "viewed together / bought together" models without any personal identifiers beyond the numeric key.
  • Price-elasticity studies - price rides along on every event, so demand response to price changes can be read off observed behavior instead of survey answers.
  • Merchandising and assortment - category codes like apparel.shoes show where browse depth fails to convert, informing which categories need better landing pages or sharper pricing.

Which personas get the most value?

Data scientists and ML engineers get a labeled, schema-stable event stream big enough to train on without augmentation. E-commerce operators get a benchmark funnel from outside their own analytics stack - useful when internal numbers need an external sanity check. Market researchers and consultants get seven months of real transactions-and-browsing to cite, at session granularity, for apparel category reports. Developers building data products get a flat CSV-shaped schema that loads into anything.

What should I know before requesting a sample?

Three things worth knowing upfront. First, the source card's file-structure table contains a copy-paste leftover describing event_type as 'only one kind of event: purchase' while its own Event types section documents all four values - the four-value reading matches the actual files. Second, whether the headline figure means 285 million events or 285 million users could not be disambiguated from the documentation alone; either way the volume is far beyond what apparel funnel modeling needs. Third, category_code is present for meaningful categories and often skipped for miscellaneous accessories, so treat it as high-precision but incomplete and lean on brand and category_id when recall matters.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - nine documented fields, one row per clickstream event
fieldtypedefinitionexample
event_timedatetimeTime when the event happened, in UTC.2019-10-01 00:00:00 UTC
event_typeenumOne of view (user viewed a product), cart (added product to shopping cart), remove_from_cart (removed a product from cart) or purchase (purchased a product).cart
product_idintegerID of the product involved in the event.1005115
category_idintegerThe product's category ID.2053013555631882655
category_codestringProduct's category taxonomy code name, usually present for meaningful categories (e.g. apparel.* segments) and skipped for miscellaneous accessories.apparel.shoes
brandstringDowncased string of the brand name; can be missing.zara
pricenumberFloat price of the product; present on each event row.21.47
user_idintegerPermanent user ID, constant across sessions for the same shopper.1515915625519845894
user_sessionstringTemporary session ID, the same within one visit; changes when the user returns to the store after a long pause.c6bd7419-ae5e-4b1c-a6d6-c63c1a7c28d1

Questions buyers ask

How far back does the data go?

October 2019 through April 2020 - seven consecutive months of clickstream from one multi-category online store. The window spans a full holiday season plus the post-holiday clearance period, so seasonal apparel patterns are observable inside a single delivery.

How do I isolate the apparel segment?

Filter rows where category_code starts with 'apparel'. The taxonomy marks meaningful categories such as apparel.shoes, though it is often skipped for miscellaneous accessories - so keep brand and category_id as fallback join keys. Price, user_id and user_session arrive on every row regardless of category.

Can I track the same shopper across visits?

Yes. user_id is permanent for each shopper across sessions, while user_session identifies a single visit and changes after a pause. Joining on user_id therefore reconstructs cross-visit journeys - a shopper who views shoes on Monday and buys on Thursday stays connected.

Is the field schema stable across months?

The nine-field schema - event_time, event_type, product_id, category_id, category_code, brand, price, user_id, user_session - is documented once and applies to every monthly file, so pipelines written against one month run against all seven unchanged. One caveat: category_code is sparse for accessory items.

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing