Tracebase — Appliance-Level Power Traces
Datadory delivers tracebase appliance level power traces data: roughly one power reading per second from individual household appliances - refrigerators, freezers, dishwashers, microwave ovens, laundry dryers, coffeemakers, irons, desktop PCs, printers and games consoles - each measured through a metered wall outlet in Darmstadt, Germany through 2012 with a smaller 2013 Sydney, Australia set, organised into roughly 43 appliance-type folders across complete, incomplete, synthetic and Australia branches, with every row carrying watts averaged over one-second and eight-second windows, delivered daily, weekly, or hourly.
What is Tracebase — Appliance-Level Power Traces?
A whole-house meter tells you the dwelling drew power. A lab bench tells you what a brand-new unit draws on a controlled bench supply. Tracebase occupies the ground between: individual appliances - refrigerators, freezers, dishwashers, microwave ovens, laundry dryers, coffeemakers, irons, desktop PCs, printers, games consoles - each measured at the wall outlet it actually occupies, at a nominal one reading per second, in situ rather than in a test cell.
The collection grew out of Andreas Reinhardt's work with TU Darmstadt's KOM and SEEMOO labs together with TU Clausthal, and its founding study asked the question the energy-analytics world still asks: how accurately can you identify an appliance from its load signature alone ("On the Accuracy of Appliance Identification Based on Distributed Load Metering Data", SustainIT 2012)? The traces that answered it went on to serve as a standing benchmark for appliance identification and non-intrusive load monitoring.
Curation is the quiet luxury. Traces ship in four flavours: complete/ holds uninterrupted midnight-to-midnight recordings that keep logging even while the appliance sits switched off; incomplete/ preserves traces with missing measurement points, left un-interpolated rather than papered over; synthetic/ pads real usage fragments with zero-consumption readings to form full day-length traces; and australia/ adds a smaller 2013 Sydney set alongside the main Darmstadt haul. The complete/ branch alone sorts into roughly 43 appliance-type folders - the Refrigerator folder carries 206 day-files on its own - and the entire corpus fits in about 315 MB. On Datadory's quality rubric it scores 7 out of 10.
What do sample rows look like?
Rows verified during the August 2026 research pass, exactly as they land in your warehouse - one row per reading tick, three columns wide.
timestamp | power_1s | power_8s
28/11/2011 00:00:01 | 0 | 0
28/11/2011 00:00:06 | 0 | 0
28/11/2011 00:00:10 | 2 | 0
14/01/2012 10:48:47 | 151 | 156
14/01/2012 10:48:48 | 147 | 151The top trio is one file's opening seconds: 28 November 2011, just past midnight. Three ticks land within ten seconds and all read zero - then a lone 2 W flicker on the one-second channel, gone by the next tick. That is standby electronics breathing, and it is exactly the kind of floor-level signal that minute-level aggregates average into invisibility.
The January 2012 pair comes from a different trace nine weeks later. Two consecutive seconds hold a steady 151 W falling to 147 W, and the eight-second column reads slightly above the one-second figure (156 versus 151) because it averages over a window that includes earlier, higher draw. Read together they sketch a motor holding a running cycle - the flat, humming plateau a refrigerator compressor or a washing machine drum leaves behind - rather than the sharp spike of a kettle element.
Notice the tick spacing up top: five seconds between the first pair of rows, four between the next. The nominal cadence is once per second and real clocks drift around it. Nothing is interpolated to hide that; what was recorded is what ships.
What fields does the dataset include?
Three columns, three jobs. One timekeeper: timestamp, in day/month/year hour:minute:second with a 24-hour clock. Two measurement channels: power_1s, watts averaged over a one-second window, and power_8s, the same quantity smoothed across eight seconds. Every definition below is verified against the source card, and every example is lifted from the shipped rows.
The brevity is deliberate, not thin. Identity does not live in the columns - it rides on the file's placement in the appliance-type organisation, so a trace arrives already knowing it is a Refrigerator rather than a LaundryDryer. Pull the measurements together with that grouping and the three numeric columns turn into named appliance behaviour; pull them without it and you have beautifully labelled anonymity.
What does coverage look like across geography, time and granularity?
Geography: two collection sites. Darmstadt, Germany carries the main corpus; Sydney, Australia contributes a smaller secondary set kept in its own branch. Two countries, two voltage environments, one schema - convenient for checking that a classifier learned signatures rather than mains characteristics.
Temporal: the Darmstadt campaign ran through 2012 and the Sydney set through 2013, with shipped day-files reaching back to November 2011. The recording effort has closed and the collection has been unchanged since February 2020, which makes it a fixed ruler: calibrate against it this year and next year and the ruler has not moved.
Granularity: a nominal 1 Hz sample per metered wall outlet, reported twice - raw through the one-second window and smoothed through the eight-second window. Each CSV covers one outlet over one day, with complete-day traces running midnight to midnight including switched-off hours. Day-files weigh in between roughly 0.4 and 1.5 MB apiece, small enough that the whole corpus loads on commodity hardware without ceremony.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Who uses this data, and for what?
- Appliance identification and NILM benchmarking - each trace arrives pre-labelled with its appliance type, which is precisely the supervised ground truth disaggregation models starve for; train on some types, test on others, and report numbers comparable with a decade of published work built on the same corpus.
- Smart-home and energy-app engineering - reference signatures per appliance category give anomaly detection and device-presence features an empirical floor, so "the dryer is running" becomes a matched filter instead of a guess.
- Standby and idle-load analysis - complete/ traces log straight through switched-off hours by design, making always-on draw measurable rather than inferred from subtracting estimates.
- Synthetic training-data design - the synthetic/ split is a worked example of stretching scarce real usage into balanced day-length samples, a recipe teams reuse when real labelled hours are expensive.
- Edge and embedded classifier work - a three-column schema and a ~315 MB total footprint keep feature pipelines light enough to prototype on modest hardware before anything ships to a plug.
- Reproducible academic work - a frozen, citable benchmark with a stable identity, so methods sections survive reviewers asking for a rerun.
Which personas get the most value?
ML engineers and data scientists get per-outlet labelled truth at a resolution that still resolves compressor cycling and heating elements, without needing a partner utility to grant access to anything. Academics and students inherit a benchmark the NILM literature has cited since 2012, which makes their comparison tables legible to reviewers. Appliance product managers and engineers see how categories from coffeemakers to consoles actually behave at the wall, duty cycle by duty cycle. Smart-home and IoT builders get signature libraries for presence detection and load monitoring features. Energy consultants and efficiency analysts get measured standby floors instead of rule-of-thumb constants. Hardware and metering teams get a published corpus to sanity-check their own plug meters against - if your device reads a refrigerator differently than these traces do, you want to know why before your customers do.
What should I know before requesting a sample?
Three things worth knowing upfront. First, this corpus is device-keyed, not home-keyed: every trace belongs to an appliance, and no dwelling, tariff or occupancy context rides along with it. Treat it as appliance physics, and reach for home-keyed panels such as REFIT or Pecan Street when the research question needs a household around the machine.
Second, mind the vintage. The recordings capture the appliance fleets of 2011-2013 - excellent for method development and for signatures that change slowly, less useful as a portrait of today's inverter-driven, always-connected devices. Scope claims accordingly.
Third, the schema is intentionally tiny, so the appliance-type organisation is doing the labelling. Request the grouping alongside the measurements and your features arrive named; skip it and you will be reconstructing identities from folder listings yourself. And budget for the cadence quirk - ticks drift a few seconds around the nominal second - by resampling deliberately rather than assuming a metronome.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
timestamp | datetime | Date and time of the reading in day/month/year hour:minute:second, 24-hour notation. | 28/11/2011 00:00:01 |
power_1s | number | Power consumption averaged over a one-second window, in Watts. | 0 |
power_8s | number | Power consumption averaged over an eight-second window, in Watts. | 0 |
Questions buyers ask
What makes a power trace "appliance-level"?
Each trace monitors one appliance through one metered plug at one wall outlet, so the watts belong to a single device rather than a circuit shared by several loads or a whole-dwelling aggregate. The appliance type travels with the file as part of its organisation, which is what turns raw watts into labelled examples.
Why are there two power columns, one-second and eight-second?
They are the same measurement through two averaging windows. power_1s keeps fast transitions - element spin-up, compressor start - while power_8s smooths jitter into a steadier line. Because the windows differ, the two columns legitimately disagree on transient-heavy stretches, and that disagreement itself flags eventful seconds.
What separates the complete, incomplete, synthetic and Australia splits?
complete/ holds day-long midnight-to-midnight recordings that continue even while the appliance is switched off. incomplete/ keeps traces with missing measurement points, deliberately left un-interpolated. synthetic/ builds full day-length traces by padding real usage fragments with zero-consumption readings. australia/ isolates the smaller 2013 Sydney collection from the main Darmstadt corpus.
How many appliances does Tracebase cover?
No global count is published in a single index, and inventing one would be worse than admitting the gap. What is documented: the complete/ branch alone organises into roughly 43 appliance-type folders, spanning categories from Refrigerator, Freezer and Dishwasher to MicrowaveOven, LaundryDryer, Coffeemaker, Iron, PC-Desktop, Printer and Playstation3, with the Refrigerator folder carrying 206 day-files. Folder-level counts resolve cleanly whenever a scope is drawn.
Is one reading per second fast enough for load disaggregation?
For most residential appliance identification, yes: compressors, drums, fans and resistive elements all leave multi-second signatures that 1 Hz captures comfortably. Sub-second transients - the switching edge of a triac dimmer, say - fall outside it, and the slight drift of real tick spacing around the nominal interval means downstream resampling should be deliberate rather than assumed.
Why benchmark on recordings from 2011-2013?
Because the corpus is frozen, everyone compares against the same ruler: results stay comparable across papers and years, and nothing under a trained model silently shifts. The trade-off is vintage - the appliance fleet captured predates today's inverter-driven devices - so teams typically pair it with newer collections when current-market behaviour matters.
Datasets that pair with this one
- Pecan Street Dataport — High-Resolution Household Energy & Water The live, home-keyed counterpart: ongoing circuit feeds from 1,000+ US homes versus a frozen device-keyed benchmark.
- REFIT Electrical Load Measurements (Cleaned) House-keyed UK recordings at 8-second cadence - the middle ground between whole-home panels and single-appliance traces.
- EU EPREL — European Product Registry for Energy Labelling Rated efficiency on the regulatory sticker versus measured watts at the wall - specification against behaviour.
- best household appliances datasets Where this corpus ranks among the industry's primary sources, with the criteria spelled out.
- household-appliances data hub All primary datasets in this industry, ranked and cross-linked.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.