Datadory notebook

CBECS microdata CSV download: the building file, delivered whole

Datadory delivers CBECS microdata covering every published record of the Commercial Buildings Energy Consumption Survey: 6,436 responding US commercial buildings carrying about 1,249 columns each - roughly 620 questionnaire variables beside imputation flags and replicate weights - weighted up to an estimated 5.9 million buildings across all 50 states and the District of Columbia, released at nine census divisions across six cycles reaching back to 1992, delivered daily, weekly, or hourly.

1,744 datasets. Pick your catch.

What does a delivered CBECS building row look like?

Two real records from the 2018 file - one row per building, coded characteristics beside measured consumption:

PUBID   : 00001                REGION : 3      CENDIV : 5
PBA     : 2 (office)           SQFT   : 210000 NFLOOR: 994
YRCONC  : 2                    OWNTYPE: 8 (federal government)
MONUSE  : 12                   NWKER  : 350
HDD65   : 4463                 CDD65  : 1759
ELBTU   : 18708970             MFBTU  : 29727152
MFEXP   : 1072600              FINALWT: 2.17

PUBID   : 00002                REGION : 4      CENDIV : 9
PBA     : 2 (office)           SQFT   : 28000  NFLOOR: 5
YRCONC  : 6                    OWNTYPE: 2 (corporation/partnership/LLC)
MONUSE  : 12                   NWKER  : 12
HDD65   : 2424                 CDD65  : 189
ELBTU   : 1528667              NGBTU  : 201988
MFBTU   : 1730655              MFEXP  : 82030
FINALWT : 312.77

The anatomy matters more than any single value. Both rows carry PBA 2 - office - yet everything downstream diverges: a 210,000-square-foot federal building in a heavy-heating climate against a 28,000-square-foot corporate building in a mild one. And FINALWT spans two orders of magnitude between them, which is the survey design speaking: the rarer a building's shape in the stock, the bigger the weight each respondent of that kind carries. Treat the weight as part of the observation, not an appendix.

Get a sample cut to your building types and divisions - real rows come back with the complete field dictionary attached.

Which fields carry the analytical work?

The nineteen-field load-bearing core stacks into three jobs. Identification: PUBID, REGION, CENDIV. Classification: principal building activity (PBA), gross square footage (SQFT) and its size band (SQFTC), floors (NFLOOR), construction era (YRCONC), owner type (OWNTYPE), months in use (MONUSE) and main-shift workers (NWKER). Measurement: degree days base 65 F (HDD65/CDD65) attached from weather data for climate normalization, per-fuel annual consumption in thousand Btu (ELBTU, NGBTU, FKBTU), the derived major-fuels totals MFBTU and dollar-side MFEXP, and FINALWT.

Beyond the core the file opens out into about 1,249 columns, and the honest way to read them is in three groups:

  1. Roughly 620 questionnaire variables describing characteristics, activity, equipment inventories and the consumption-and-expenditure ledger.
  2. Imputation flags marking every value the survey filled in statistically rather than observed - the difference between an honest model feature and a quietly biased one.
  3. Replicate weights (the FINALWT1... series), existing precisely so standard errors reflect the survey design instead of a naive formula.

The classic mistake is averaging a flag column as though it were a measurement. The second classic mistake is ignoring FINALWT entirely and reporting unweighted averages - which answers a question nobody asked.

How do you turn 6,436 rows into defensible national estimates?

A fixed routine separates citable numbers from bar-chart filler, and none of it is optional:

  1. Sum FINALWT within whatever segment you cut. That single habit scales results from 6,436 respondents to the 5.9-million-building universe - stock shares, market sizes, intensity baselines.
  2. Take uncertainty from the replicate weights. They encode the survey's design; a naive variance formula will flatter your precision.
  3. Respect the division floor. Nine census divisions are the finest geography ever released at building level - no state identifiers survive disclosure processing. Regional cuts are built on CENDIV or they are not built.
  4. Expect small gaps against published tables. The building-level records are masked while the agency's pretabulated output is produced from unmasked data, so your tabulations will differ slightly - by design, not defect.
  5. Date-stamp every figure with its cycle year. Waves are periodic rather than annual, so anything you publish is a snapshot claim, and reviewers will ask which snapshot.

Run honestly, the whole sequence is an afternoon of analysis rather than an afternoon of plumbing - which is the point of taking rows that arrive already shaped.

How does CBECS compare with its neighbors?

Three other delivered records circle the same built environment, and they answer different questions. Knowing which one you are asking saves a month of misfitted work.

Practical rule: CBECS when you need every variable side by side for custom segments nobody has published; the Building Performance Database when you need volume and can work in peer aggregates; Portfolio Manager for the buildings you already operate; GHGRP when the unit of analysis is emissions rather than real estate. Analysts usually end up needing two of the four - which is why they arrive with compatible keys rather than four incompatible formats.

What do teams build on commercial building microdata?

  • National benchmarking denominators. Weighted CBECS distributions define what "typical" consumption, expenditure or equipment looks like across 5.9 million buildings - the reference frame peer-group tools and rating scores sit on top of.
  • Market sizing for building services. Summing weights by segment turns the file into a stock estimate, the denominator a janitorial or energy-services go-to-market plan needs before any directory is counted. Certified-provider registers then supply the demand side's counterpart: the ISSA CIMS certified cleaning organization directory lists about 390 certified organizations and the Green Seal certified products directory indexes roughly 360 certified companies.
  • Model features and validation sets. Observed and imputed values ride side by side with flags distinguishing them - unusually honest training material for intensity, benchmarking and anomaly models. The fuller workflow lives on the data scientists use cases page.
  • Program evaluation and policy claims. Before-and-after snapshots on consistent survey footing support efficiency-program impact statements without proprietary data.
  • Cross-country context. When the question goes abroad, Eurostat waste statistics (municipal series from 1995) and the OECD municipal waste statistics (58 reference areas) extend the same discipline - weighted figures, cited vintages - to European and partner economies.

Who works with this record, and for what?

Market researchers and consultants get the definitive segmentation of the US commercial stock - activity, size, era, ownership and workforce on one survey footing - for bottom-up TAMs and category maps.

Data scientists and ML engineers get a labeled, weighted, weather-normalized building panel of 6,436 records by roughly 620 variables, ready for intensity and anomaly work.

Sales and growth teams at building-equipment and energy-services firms get territory sizing from the government's own denominators - buildings, floorspace and workers by segment and division.

Competitive-intel and product teams map how many targets exist, of what type, where, before committing to a wedge.

Journalists, academics and students get the number everyone else cites - sourced to a national probability sample rather than a vendor white paper.

Why get CBECS microdata through Datadory?

Because the record is famous and still awkward. Coded categories arrive undecoded; flag columns masquerade as measurements; the weight apparatus intimidates anyone who has not read the codebook; and the file changes shape between cycles. Datadory ships the rows already shaped: categories decoded or preserved as your spec requires, imputation flags traveling beside the values they qualify, weights intact and documented, join keys stable across cycles.

Files, feeds, or straight into your warehouse. Daily, weekly, or hourly - your call. The sample comes first either way: name the building types and divisions, get real rows, then decide.

Where should you start?

Start with the anchor record, EIA Commercial Buildings Energy Consumption Survey (CBECS), sampled to your building types and divisions before anything is committed. Product-level detail - sample rows, the full field dictionary, coverage chips - lives on its dataset page.

This page is one thread of a wider map. The Environmental & Facilities Services data guide walks the sixteen-record slice, the best environmental & facilities services datasets ranking scores all ten leaders side by side, and the environmental-facilities-services data hub indexes every record with its coverage statement.

Datasets covering the US commercial building question, compared (as of August 2026)
DatasetUnit of observationGeography in outputTemporal depthWhat it adds
EIA Commercial Buildings Energy Consumption Survey (CBECS)One anonymized responding building - 6,436 records, weighted to an estimated 5.9 million buildingsAll 50 states plus DC in the sample, released at nine census divisions; no state identifiersSix cycles, 2018 back to 1992Full row-level file of about 1,249 columns - the only national probability sample pairing per-building characteristics with measured consumption and expenditure
DOE Building Performance Database (BPD)Anonymized building-year records - 1.1 million-plus buildings, roughly 1.56 million building-years, 33 fields eachResolved to 5-digit ZIP code, city, state and ASHRAE climate zoneMeasured data years from 1990 through the most recent cyclePeer-group aggregates or random samples of up to 1,000 buildings - volume, never full extracts; groups under ten buildings suppressed
ENERGY STAR Portfolio ManagerProperty and meter inside a benchmarked portfolioUnited States and Canada; individual records cover any benchmarked property worldwideRolling monthly and annual reporting periods entered per propertyOperational tooling: 1-100 scores across roughly a quarter of US commercial building space, built on self-reported inputs
EPA Greenhouse Gas Reporting Program (GHGRP) dataFacility-level annual emissions rows, with unit-, fuel- and gas-level breakdowns beneathUS reporters located to state, county and coordinatesReporting years 2010-2023, roughly 8,000 reporters a yearLarge-emitter greenhouse gas detail across 32 industry types - industry's emissions footprint, not the building stock

Pick up where this leaves off

Every one of these ships with sample rows before you commit to anything.

Environmental & Facilities Services United States - all 50 states and DC in the national sample…

EIA Commercial Buildings Energy Consumption Survey (CBECS)

PUBID · REGION · CENDIV …+13 more

Environmental & Facilities Services United States - resolved to 5-digit ZIP code, city, state and…

DOE Building Performance Database (BPD)

Labeled

Environmental & Facilities Services United States and Canada - the tool serves as Canada's…

ENERGY STAR Portfolio Manager Data

Environmental Facilities Services United States - facility-level locations with state, county…

EPA Greenhouse Gas Reporting Program (GHGRP) data

Environmental Facilities Services Primarily United States and Canada, with listings across Latin…

ISSA CIMS Certified Cleaning Organization Directory Data

Environmental & Facilities Services Products from manufacturers worldwide

Green Seal Certified Products Directory

Want rows instead of a pitch? Name the datasets.

API, files, or your warehouse. Daily, weekly, or hourly.

Get a sample

Questions worth asking

What does one CBECS record represent?

One responding, in-scope US commercial building - 6,436 of them in the 2018 cycle, arranged across about 1,249 columns. Because the sample is drawn rather than enumerated, each record carries a survey weight, and inflating the respondents by their weights yields an estimated 5.9 million buildings spread across all 50 states and the District of Columbia.

Which CBECS survey years are covered?

Six published cycles: 2018 as the latest completed wave, then 2012, 2003, 1999, 1995 and 1992. Each is a full-stock snapshot rather than a point on a continuous series, so trend work stitches cycles on consistent definitions and cites the vintage alongside every figure.

Can you identify individual buildings or states in the microdata?

No. The records are disclosure-avoided, so no individual building can be identified, and the finest geography released is the census division - nine categories with no state identifiers anywhere in the file. That masking is what makes publishing per-building data possible at all.

Why do computed totals differ from published CBECS tables?

Because published tables are produced from unmasked data while the building-level records are masked. Small gaps are expected and structural, not errors. Read the imputation flags before building arithmetic on filled-in values, and take standard errors from the replicate-weight series rather than a naive formula.

How is CBECS data delivered?

As typed rows - files, feeds or straight into your warehouse - daily, weekly, or hourly, your call. Categories decoded or preserved as specified, imputation flags beside their values, weights intact, join keys stable across cycles. Request a sample cut to your building types and divisions first.