Skip to content

Modeling Land Values in Baltimore

Implementing a land value tax (LVT) would require accurate and efficient assessments of land value across all properties in a city. Colleagues from the National Center for Smart Growth, Center for Open Geographical Science, and I investigated the potential to estimate land value using a statistical method called a hedonic price model and detailed property characteristics across Baltimore, Maryland. We used data on sales of residential properties over a ten year period from 2013 to 2022 to estimate how variables related to both land and improvements contributed to sale prices. We then broke those contributions into components estimating the separate values of land and improvements. The result was a model for estimating land values within Baltimore based on commonly-available data inputs. We tested the model on non-apartment residential properties throughout the city in 2022, developing a land value surface based on factors such as zoning, proximity to the city center, access to jobs, nearby crime rates, and neighborhood tree canopy. A surface like this might eventually be used to set land value taxation rates based on the location characteristics, without the need for on-the-ground assessment.

Our process generated a few lessons for land value modeling in practice:

  • In theory, publicly available data has strong potential to be used for land valuation because key variables that impact land value—proximity to amenities, neighborhood characteristics, buildable potential—are public in nature. Improvements should be more difficult to value because they are characteristics of private property and tend to be poorly represented in public datasets. Even the well-developed database we used, from Maryland’s State Department of Assessments and Taxation (SDAT), had substantial incomplete data about property characteristics that limited our ability to account for improvements. In practice, of course, estimating land values with the modeling approach we used involves first estimating improvement values as an intermediary. Unless policymakers can agree on how to value land characteristics more directly, land valuation will be limited by access to detailed data about improvements.
  • Census data tabulated in areal units, such as block groups, were too coarse for effectively predicting land values. When we added census variables to our models, they tended to dramatically overestimate land values, forcing improvement values to be very small, often negative, to compensate. While improvement values can sometimes be negative, such as with derelict buildings, these are a fringe case, even in a city experiencing substantial vacancy levels such as Baltimore. Instead, the issue with census data likely stemmed from the inability of areal statistics, such as percentages within block groups, to adequately represent the variability of sociodemographic characteristics between nearby places. We recommend not directly using census data tabulated in areal units for modeling land values. However, because census data are an important resource for representing neighborhood characteristics, determining how to summarize them with more locational specificity may be an important way to improve land valuation.
  • Spatial models have theoretical potential to improve land value estimates by accounting for how properties are influenced by their neighbors. Geostatisticians call these neighborly influences spatial lags. In our tests, however, spatial models posed practical challenges for separating land and improvement values. Most concerningly, adding spatial lags tended to inflate land values, forcing many decomposed improvement values to be negative, similar to the effects of adding census data, and potentially for similar reasons. Because spatial lags were calculated as averages within a radius of each property, they may have inappropriately represented highly variable neighborhoods as homogenous. Harnessing spatial lags for land valuation may require more customized definitions of neighborhood structure.

While our final model was relatively simple—it used four key variables to represent improvement value and eight to represent land value—it captured nearly 85% of variability in sales prices of non-apartment residential properties and effectively revealed how proximity to downtown and major transportation corridors impacted land values in Baltimore. In doing so, it provided proof of concept for how land values may be estimated without on-the-ground assessments and generated practical considerations for how data and modeling techniques can be applied to land valuation.