US Census Characteristics of New Housing

Datadory delivers us census characteristics of new housing data: annual characteristics for every new privately-owned US home - square footage, bedrooms, bathrooms, wall material, heating, financing and sales price - as roughly ninety national and regional summary tables a year plus record-level Survey of Construction microdata reaching back to 1999.

What is US Census Characteristics of New Housing?

Most housing statistics count how many homes got built. Characteristics of New Housing answers the better question - what kind of homes? Produced by the US Census Bureau and built on the Survey of Construction, a program partially funded by HUD, it describes the attributes of every new privately-owned residential structure: square footage, bedrooms, bathrooms, exterior wall material, heating system and fuel, fireplaces, air conditioning, stories, foundation, parking, financing type, contract price and price per square foot.

The summary side arrives as roughly ninety tables a year, each one a house count and percent distribution for a single characteristic. The topic list reads like a builder's spec sheet: age-restricted communities, air conditioning, bedrooms broken out against bathrooms, construction method, exterior wall materials by type, financing, fireplaces, floors, foyer, framing, foundation, stories, heating system and fuel. Every characteristic comes in cuts for all completed units, for sold units, and for multifamily units and whole multifamily buildings, so the townhome pipeline never gets buried inside the single-family numbers.

Underneath the tables sits the record level. Each year's Survey of Construction microdata file carries one observation per home with about forty-five documented variables - start, completion and sale dates, square footage, bedrooms, bathrooms, sales price, contract price, permit value, lot value, financing, wall materials, and division and metro location identifiers - plus a survey weight for rolling records up to national estimates. The series runs from 1999 through the 2025 vintage, which landed July 1, 2026.

What do sample rows look like?

Two cuts from the research pass, rendered exactly as they land in your warehouse.

Number of Bedrooms in New Single-Family Houses Completed - the longest summary series.

year | total_houses_thousands | two_bedrooms_or_less | three_bedrooms | four_bedrooms_or_more
1973 | 1197                   | 148                  | 769            | 280
1980 | 957                    | 163                  | 603            | 192

SOC microdata variable guide - the record-level dictionary.

item | variable | description
41   | STRT     | Start Date, YYYYMM such as 200901; 0 = not started
43   | SLPR     | Sales price at first contract signing, incl. improved lot; nearest $100
44   | SQFS     | Square foot area of house; top and bottom 1% set to limit values

Read the bedroom rows as a decade of churn in two lines. Between 1973 and 1980, completions fell from 1,197 thousand to 957 thousand while homes with two bedrooms or less climbed from 148 thousand to 163 thousand - their share rising from about 12 percent to about 17 percent even as the three-bedroom majority held. That is the kind of shift product planners and land buyers can price, and the same tables carry it forward year by year to the present.

What fields does the dataset include?

The dictionary below covers the core summary-table columns and the most-used microdata variables, all verified against the source documentation rather than inferred. Expect two shapes: summary tables give one row per year per characteristic category, while the microdata give one row per home with dates, prices, physical attributes and location identifiers on the same record.

Fields marked in the folded note - financing type, foundation type, parking facility type, division and metro identifiers, survey weight and the remainder of the roughly forty-five-variable microdata guide - are documented in full with examples when you request a sample.

What does coverage look like across geography, time and granularity?

Geography: United States national totals, with most characteristics also broken out at Census region level. The microdata go finer still, carrying Census division and metropolitan area identifiers on each record, so a regional cut you prototype on the summary tables can be reproduced and extended at record level.

Temporal: the summary series reach back decades - the bedrooms table starts in 1973 - and a dedicated historical tables collection covers 2003 through 2017 in one place, alongside a Summary of Changes documenting every file change since 2005. Record-level microdata run 1999 through 2025.

Granularity: annual throughout, in two forms. Summary tables publish counts and percent distributions per characteristic category at national and regional level; the microdata publish individual housing records with a survey weight so national estimates can be rebuilt from the record level. One row per home versus one distribution per characteristic - most teams take both and meet in the middle.

How is the data delivered?

API, files, or your warehouse. Daily, weekly, or hourly.

Who uses this data, and for what?

  • Floor-plan and feature strategy - builders and product planners track what the market actually ships: bedroom counts against bathroom counts, ceiling heights, fireplace penetration, garage size, and where vinyl overtakes brick or stucco region by region.
  • Building-products demand forecasting - wall material, heating system and fuel, air conditioning and foundation splits convert directly into addressable-unit forecasts for siding, HVAC and slab-work suppliers, with regional cuts to route the forecast.
  • Land acquisition and site planning - price per square foot and square-footage distributions by region tell acquisition teams whether a submarket's new product is moving up-market or densifying.
  • Appraisal, lending and valuation benchmarks - contract price, sales price and permit-value fields on record-level observations give valuations desks a defensible basis for new-construction comps.
  • Housing-policy and affordability research - the shrinking share of smaller homes is measurable here across five decades, which makes it the citation of record for starter-home scarcity arguments.
  • Model training - record-level microdata with dates, prices, physical attributes and geography is clean, consistently-coded training material for hedonic price models.

Which personas get the most value?

Market researchers and consultants get five decades of citable federal statistics behind every housing presentation, down to wall material and foyer presence. Investors and quants get characteristics data that explains which builders are exposed to which floor plans and price points - mix matters when the cycle turns. Data scientists and ML engineers get record-level observations with a documented variable guide, consistent coding and survey weights, ideal for hedonic pricing and demand models. Homebuilding executives get the spec-sheet benchmark for their own product lineup. Journalists and academics get the definitive answer to 'are new homes getting bigger?' with the history to prove it. Product teams building real-estate tools get authoritative characteristics to anchor comparisons and enrich listings.

What should I know before requesting a sample?

Three things worth knowing upfront. First, mind the two shapes. The summary tables are wide distributions - one row per year per characteristic category - while the microdata are narrow and deep, one row per home. Joining them takes a deliberate decision about which level your analysis lives at; most production pipelines ingest both and keep them in separate tables keyed on year and geography.

Second, the record-level values are deliberately confidentialized. Prices and square footages are top-and-bottom coded - the top and bottom 1 percent of reported values become the values at those limits - and records carry flags. That means record-level averages will not match published medians exactly, and any national figure rebuilt from microdata needs the survey weight applied.

Third, pick the variant before you model. Completed-unit tables, sold-unit tables, multifamily-unit tables and multifamily-building tables each tell a different story about the same underlying program, and region-level detail is standard on most but not every summary table. Say which cut you want and the sample arrives matched to it.

Field dictionary

Every field below is documented against real records. The full dictionary ships with the sample.

Field dictionary - summary-table columns and core SOC microdata variables
fieldtypedefinitionexample
YearintegerCalendar year of construction or sale in the summary tables; the bedrooms series reaches back to 1973.1973
total_houses_thousandsnumberTotal new single-family houses completed in the characteristic category, in thousands.1197
percent_distributionnumberPercent share of houses falling in each category of the characteristic.12.4
STRTdateMicrodata start date of the home, YYYYMM; 0 means not started.200901
COMPdateMicrodata completion date of construction.200910
SALEdateMicrodata date of sale.200911
SLPRnumberSales price at the time the first sales contract was signed, including the improved lot; whole dollars rounded to the nearest $100; top and bottom values within division set to limits.250000
CONPRnumberContract price on the original general contractor contract; applies to contractor-built houses.180000
SQFSintegerSquare foot area of the house; the top and bottom 1 percent of reported values are changed to the values at those limits.2400
BEDRintegerNumber of bedrooms.4
FULB / HAFBintegerNumber of full bathrooms and half bathrooms.2
WAL1 / WAL2stringPrimary and secondary exterior wall material.Vinyl

Questions buyers ask

What is the difference between the summary tables and the SOC microdata?

The summary tables are annual distributions: one row per year per characteristic category, giving counts and percent shares nationally and by Census region. The microdata are record-level: one row per surveyed home with dates, prices, square footage, bedrooms, bathrooms, materials and location identifiers. Analysts use the tables for trends and the microdata for modeling.

How far back does the data go?

The longest summary series, bedrooms in new single-family houses, starts in 1973. A dedicated historical tables collection consolidates 2003 through 2017, and a Summary of Changes documents every file change since 2005. Record-level Survey of Construction microdata runs from 1999 through the 2025 vintage.

Are individual homes identifiable in the microdata?

No. Records carry no information identifying specific addresses or builders, and reported values are protected by top-and-bottom coding, meaning extreme price and square-footage values are replaced with the values at set limits, plus explicit flags. Files also carry disclosure-authorization numbers documenting the review applied.

Which price fields exist and how do they differ?

SLPR is the sales price at the time the first sales contract was signed, including the improved lot, rounded to the nearest $100. CONPR is the contract price on the original general contractor contract and applies to contractor-built houses. Permit value and lot value ride alongside them on each record-level observation.

Does the data cover multifamily buildings, not just single-family homes?

Yes. Alongside the single-family completed-unit tables there are dedicated cuts for multifamily units in buildings and for whole multifamily buildings, so apartments, condominiums and townhome product can be analyzed on their own terms rather than averaged into a single-family story that would misread both segments.

Can I rebuild national estimates from the record-level data?

Yes. Each microdata record carries a survey weight designed for producing national estimates, and the documented variable guide explains its use. Apply the weights before aggregating; unweighted record counts will understate totals because the survey samples the construction universe rather than enumerating it.

Datasets that pair with this one

See the rows before you pay anything.

Name this dataset and we send real records from it — scoped to the fields you asked for.

See pricing