Urban Development

Identification of Proxy Datasets for Urban Boundaries: A New Perspective on Infrastructure Planning

Introduction

The starting point of infrastructure planning is often the definition of a "city." However, city boundaries are not naturally occurring—they are human constructs, depending on the dataset and assumptions used. A recent perspective article published in *Nature Cities* (Van Migerode et al., 2026) systematically reviews proxy datasets used to identify city boundaries, including population distribution, built-up area, and nighttime lights, and proposes a purpose-driven selection framework. For infrastructure analysts, project financiers, and regional planners, this framework has profound practical implications: different boundary definitions directly alter assessments of infrastructure demand, capital allocation, and regional connectivity.

The Spectrum of Proxy Datasets

Researchers typically use three types of proxy datasets to delineate city boundaries:

  • Population datasets: Such as global gridded population data (e.g., WorldPop, GPW), emphasizing population density and agglomeration. Their advantage is being directly related to the social dimension of urbanization, but they are susceptible to administrative boundaries and census quality.
  • Built-up area datasets: Based on remote sensing imagery (e.g., MODIS, Landsat) to extract impervious surfaces, reflecting physical expansion. Suitable for analyzing land-use change, but may overlook low-density or informal settlements.
  • Nighttime light datasets: Such as DMSP/OLS, VIIRS, capturing human activity intensity. Their advantage is being related to economic activity, but they suffer from oversaturation and low spatial resolution.

Each dataset implies a different concept of "city": population data emphasizes "people," built-up area emphasizes "things," and nighttime lights emphasize "activity." These differences can cause the same area to be included or excluded from city boundaries, thereby affecting subsequent analyses.

A Purpose-Driven Selection Framework

  • The authors propose that the selection of proxy datasets should be based on the application purpose, rather than pursuing some "correct" boundary. For example:
  • To assess the urban heat island effect, built-up area data may be more appropriate (Yang et al., 2023);
  • To study economic agglomeration, nighttime light or mobile phone signaling data are more effective (Dingel et al., 2021; Dong et al., 2024);
  • To plan disaster resource allocation, population data are more critical (MacManus et al., 2021).For infrastructure investment, this framework implies:

1. Project Financing: When evaluating urban infrastructure projects, lending institutions should clearly define how the "urban area" served by the project is delineated. Different boundaries lead to different risk exposures and return expectations.

2. Regional Development Corridors: When planning transportation corridors or energy networks, the continuity of boundaries determines the corridor layout. Nighttime light data may reveal actual economic connections across administrative boundaries, while population data may reinforce administrative divisions.

3. Construction in the Global South: Many developing cities lack high-quality administrative data, making proxy datasets the only option. However, uncritically applying methods from high-income countries may underestimate population density in informal settlements, leading to inadequate infrastructure planning.

From Data to Decision: Case Considerations

Suppose an international development agency plans to invest in a railway corridor in a rapidly urbanizing region of Africa. Using population data may confine the city boundary to the high-density core area, suggesting stations be concentrated in the city center; while using nighttime light data may reveal a low-density but economically active suburban industrial belt, suggesting an extended route. These two decisions will significantly affect the project's capital allocation, passenger demand forecasts, and financial models.

Similarly, in the expansion of a port city, built-up area data can accurately reflect physical occupancy but cannot capture commuter flows. Relying solely on built-up area data to plan the port access road may overlook cross-district traffic pressure.

Implications for Infrastructure Researchers

This study reminds us: urban boundaries are not objective facts but analytical tools. When citing any data on "urban area," "urbanization rate," or "urban population," infrastructure analysts must question the proxy dataset behind it. Otherwise, seemingly precise statistics may mask fundamental conceptual biases.

Future directions include: integrating multi-source data (e.g., mobile phone signaling + high-resolution imagery) to improve accuracy, and developing standardized boundary definitions for specific infrastructure categories (e.g., transit-oriented urban boundaries).

Conclusion

The identification of urban boundaries is a seemingly basic yet critical issue in infrastructure planning. The work by Van Migerode et al. provides a practical selection logic that enables researchers to make well-informed trade-offs based on specific needs. For the global infrastructure community, embracing this uncertainty and actively managing it will help more efficiently guide capital flows, reduce misjudgments, and support sustainable urban expansion.

Reference trail · globalinfrareview

globalinfrareview frames this note through Projects / Investment / Energy & Utilities. Projects / Investment / Energy & Utilities explains the local editorial angle; Source links should be opened before the summary is reused (dates, names and status changes still need checking).

Source links

  1. https://www.nature.com/articles/s44284-026-00464-6Primary

Related articles

Back to channel