
==== Front
Sci Data
Sci Data
Scientific Data
2052-4463
Nature Publishing Group UK London

3829
10.1038/s41597-024-03829-5
Data Descriptor
New Commuting Zone delineation for the U.S. based on 2020 data
http://orcid.org/0000-0001-8415-0441
Fowler Christopher S. csfowler@psu.edu

grid.29857.31 0000 0001 2097 4281 Department of Geography Penn State University, University Park, USA
6 9 2024
6 9 2024
2024
11 97529 1 2024
27 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
This paper documents the creation of new ‘commuting zones’ for the United States based on 2020 geographies and data. Commuting zones originated in the late 1980’s as the result of an effort by the Economic Research Service of the U.S. Department of Agriculture to provide a county-based delineation of functional regions that covered the entire U.S. and linked rural areas to their nearest economic center. The commuting zones presented here update the 2010 era definitions. Additionally, this new delineation incorporates measures of quality to facilitate comparison with earlier decades and to account for the fact that the quality of commuting-based delineations may have been substantially affected by changing patterns of work and residence during and after the covid-19 pandemic. The methodology used here seeks to replicate as nearly as possible the method employed in earlier versions of this delineation using hierarchical clustering on a proportional flows matrix generated from county to county commuting flows. The data and all scripts used to generate them and this paper are available at https://github.com/csfowler/CommutingZones2020.git.

Subject terms

Geography
Economics
Sociology
issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcBackground & Summary

Commuting zones (CZ’s) originated in the late 1980’s as a collaboration between the U.S. Department of Agriculture’s Economic Research Service (USDA ERS) and researchers focused on understanding the economics of rural places in the United States1,2. The delineations responded to several critical needs in the data infrastructure of the U.S. including:Economic data is primarily released at the county-level, but county geographies are heterogeneous in terms of population and size and do not represent coherent economic units, so a larger unit built from combinations of counties was desirable, and

Metropolitan areas, themselves composed of counties and better able to represent coherent economic units, exclude most rural counties.

In addition to their original purpose, CZ’s proved useful for both longitudinal and spatial analysis because they were composed of counties whose geographies change more slowly than other administrative units and because they cover the entirety of the U.S. without gaps. Since their original publication in 1987 CZ’s have seen wide use in research on topics ranging from health3,4, poverty5, racism6, and energy7, to more traditional economic analyses tied to labor markets8,9. Use of the delineations has accelerated in recent years with over forty papers citing the 2016 paper describing the 2010 delineation and steadily rising downloads of the delineation files even as they become more and more dated. A need for an update to the 2010 delineations to incorporate new data from the 2020 census is clearly warranted.

Since the establishment of CZ’s in 1987, research on regionalization, the problem of grouping units together to create meaningful super-units has seen significant growth. Major advances in cluster detection of various sorts have led to a range of proposed alternatives to the relatively simple method (hierarchical cluster analysis on a proportional flow matrix) used to delineate CZ’s and the broader class of substantially similar delineations known as ‘functional regions’10–19. These alternative proposals include different measures of fit such as ‘modularity’11, maximization of interaction measures17, graph-theoretical measures of dominant flows19, and fuzzy set theory16.

While these alternative methods have much to offer, detailed reviews of alternative methods combined with a range of metrics to assess their quality found significant uncertainty present in both the extant delineations and their alternatives20,21. While the CZ delineations can be fairly critiqued, they benefit from a consistent and simple methodology covering almost half a century of research and give up little to the more complicated alternatives in terms of capturing the kinds of regional effects for which they were designed21.

A larger issue with respect to the quality of the CZ delineations is a changing geography of work in the U.S., hastened by the covid-19 pandemic, that calls into question the use of county to county commuting data as a basis for understanding economic regions22. The rise of remote work and the continued decline of small and medium sized cities23,24 have redrawn the connections between places that are at the heart of most functional region definitions and suggest that the basis for understanding these regions will need to shift in the near future. Additionally, there has long been an acknowledgement that the geography of regions is not consistent across the U.S. with major differences between the overlapping labor markets of the Northeast and the centralized regions found around Western cities like Denver and Las Vegas25. Despite beginning conversations about alternatives, Federal agencies have decided to retain commuting flows as the basis for generating metropolitan delineations for at least another decade26,27.

In light of the identified uncertainty in both the data and delineation methods employed to generate CZ’s it is critical that delineations be accompanied by metrics describing the quality of the delineation. Fowler21 documented a detailed set of metrics that allow researchers to distinguish between delineations but also among regions within a single delineation. These metrics allow researchers to address variations in quality as well as fine tune their analyses to differences in whether they are studying labor markets, commuting regions, or some other slight variation on the functional region concept.

This data descriptor documents the generation of CZ’s based on recently released data from the 2020 decennial census28 and incorporating the quality metrics developed by Fowler and Jensen21. The delineation method from prior decades, which incorporates a proportional flow matrix and hierarchical cluster analysis is retained to preserve the continuity but the addition of further metrics for understanding commuting zone quality and county fit extends the work begun by Fowler and Jensen21 to address the uncertainty inherent in these delineations. The paper documents some of this uncertainty uncovered in the process of technical validation and proposes a small number of modifications that users of the data might benefit from. The data is made freely available along with delineations from previous decades at https://sites.psu.edu/psucz and the complete replication code for generating the data and evaluation metrics as well as this paper with its accompanying visualizations is available at https://github.com/csfowler/CommutingZones2020.git.

Methods

The data on which the delineation is based comes from a variety of sources consistent with delineations from earlier decades. The raw data on journey to work counts and core-based statistical area (CBSA) delineations were accessed as tables directly from the census27–29, while county boundaries were accessed from the Census api using tidycensus30. Quarterly wage data for calculating evaluation metrics was acquired directly from the Bureau of Labor Statistics31.

The methodology for delineating 2020 CZ’s is meant to replicate the original method by Tolbert and Sizer32 with the adaptations to current data sources documented in Fowler et al.20. The evaluation metrics included with the data are those described in Fowler and Jensen21. The core concept underlying the delineation is a measure of connection between counties such that counties where a higher proportion of commuters travel between the counties represents a stronger connection. The metric, called ‘proportional flow,’ is:Cij+CjiminWi,Wj

where Cij and Cji represent the counts of commuters leaving county i for county j and leaving county j for county i respectively, and Wi and Wj are the total workforces of those two counties. Consistent with the original methodology any connection that achieved a value greater than 1 was reduced to 0.999; indicating maximum connectivity. The proportional flow matrix was then converted to a dissimilarity matrix by subtracting the proportional flow value from 1 and setting the diagonal of the resulting matrix equal to zero.

The delineation of commuting zones then uses hierarchical cluster analysis33 on this dissimilarity matrix with a cutoff value of 0.977 utilized in the 2010 delineations as described in Fowler et al.20. The result is a delineation with 593 commuting zones. Compared with 625 in 2010. The delineation includes six commuting zones for Puerto Rico, which was not included in the 2010 delineation. As noted in the technical validation below, alternatives to this cutoff value are worth exploring as it is somewhat arbitrary, but the value does have utility for maintaining consistency across delineations and there is no compelling evidence for the choice of an alternative in the range of descriptive and fit metrics covered below.

We can examine the degree to which cluster delineations are similar using the Jacand similarity for each county c where ten_c is the list of all counties in the same commuting zone as c in the 2010 delineation and twenty_c is the list of all counties in the same commuting zone as c in the 2020 delineation. The Jacand similarity is defined as:Similarityc=tenc∩twentyctenc∪twentyc

with a maximum of 1 and a minimum of 0. For this comparison to work a small allowance has to be made for counties and county equivalents that changed between 2010 and 2020. The most significant change is the move from counties to planning regions in Connecticut. For simplicity, county and county equivalent land-based centroids were joined to 2010-era commuting zones so that each observation in the 2020 data set has exactly one assigned 2010 commuting zone. The Jacand similarity is then based on the assignment of the 2020 observation to its 2010 commuting zone as compared to its assignment to the 2020 delineation.

On the whole the comparison of the two delineations shows a relatively high level of similarity. The mean similarity score comparing the two delineations is 0.659, or 0.676 if the scores for Puerto Rico (not included in the 2010 delineation) are omitted. This value jumps to 0.752 if we weight scores based on population; confirming that the most populous places are more stable in terms of their commuting zone membership. Figure 1 shows the distribution of similarity scores for each county and shows both a high overall fit and the lack of any particularly strong regional effect when it comes to low similarity scores.Fig. 1 Map of Jacand Similarity of 2020 Commuting Zones and 2010 Commuting Zones.

Overall, the commuting zones for this delineation are quite similar to earlier delineations as further shown in Table 1. The only meaningful differences between the two delineations in terms of general characteristics are the smaller number of CZ’s identified in 2020 (particularly notable given the addition of 10 new CZ’s for Puerto Rico), the increase in the number of non- contiguous commuting zones (from 3 to 6), and the existence of a very small commuting zone (31 sq. km) comprised solely of Falls Church City, VA whose outlying counties shift to other commuting zones in the 2020 delineation. While there is no ‘correct’ number of commuting zones, the consolidation of CZ’s between 2010 and 2020 is consistent with the expected pattern of consolidation associated with continued growth of the largest urban areas and expanded commuting distances bringing smaller centers into the orbit of larger ones. Even the increased presence of non-contiguous counties is an expected, albeit sometimes conceptually problematic, outcome of increased commuting distances and remote-work opportunities. The existence of single county CZ’s has always been a problem for this delineation method, and the addition of Falls Church City to the list of single county CZ’s is a reminder that the original methodology included a step where expert opinion was allowed to modify the results-a practice abandoned for the delineations in 2000 and 2010 for its lack of transparency even if it likely increased the quality and usability of the delineations.Table 1 Diagnostic Statistics comparing 2010 and 2020 Commuting Zones.

Measure	2010	2020	
No. of Counties in Delineation	3, 222	3, 143	
Number of Commuting Zones	593	625	
No. of Single County CZs	47	40	
No. of Counties in Largest CZ	19	20	
No. of Non-Contiguous CZs	6	3	
Min Population in CZ	416	997	
Average Population in CZ	563, 862	493, 993	
Max Population in CZ	18, 794, 886	17, 877, 006	
Min Area of CZ (sq.km)	31	316	
Average Area of CZ (sq.km)	15, 768	14, 954	
Max Area of CZ (sq.km)	500, 112	500, 076	
Compactness of CZ’s	0.4	0.4	

We can further examine the characteristics of the new CZ delineation by examining the fit characteristics of commuting zones and individual counties as described in Fowler and Jensen21. The purpose of this exercise is to better understand where and how commuting zones differ in terms of their fit with our theoretical understanding of how commuting zones should function. Fowler and Jensen21 distinguish between Core, Connection, and Containment as three theoretical frames for understanding functional regions. Here we examine these characteristics summarized for the complete 2020 and 2010 delineations. While a detailed examination exceeds the scope of this article, these values are also calculated for individual counties and as averages for individual CZ’s and made available with the provided delineation to aid researchers in understanding gradations in the way counties and CZ’s fit the theoretical frame we have established for them.

Core

Core refers to whether a commuting zone has an important economic center and how critical a role that center plays in the commuting zone. This concept is most salient in delineations such as those for (aptly named) core-based statistical areas where the role of this economic center is the driving point of interest. Core-centered delineations do tend to omit many rural places and make difficult and subjective decisions about how big an economic center has to be in order to count as a core. In advance of the release of 2020 metropolitan area delineations the Office of Management and Budget hosted a rather contentious discussion about raising the threshold size for an urban core from 50,000 to 100,000 persons34. While OMB ultimately decided to retain the 50,000 person threshold, the decision raises important questions about the role of economic centers in defining functional regions and so Table 2 reports information on how many CZ’s contain an OMB defined core county (77.6% in 2020) as well as what the average number of residents in a CZ who work in a core county (41%) and what the average share of the CZ workforce that lives in the core is (37%). The relatively low numbers for the latter two statistics are a reflection of the fact that many CZ’s are conceptually designed around inclusion of exurban and rural places. The results in Table 2 are encouraging for their relative stability. The number of metropolitan areas that get split does go up substantially from 36 to 44 but the Core measures are otherwise quite stable across delineations, but see the more detailed analysis of this phenomenon in the technical validation section below. Visual inspection of where and how much these values changed between delineations (omitted for brevity) does not raise any red flags about the new delineation.Table 2 Fit Statistics comparing 2010 and 2020 Commuting Zones.

	Measure	2020	2010	
Core	No. Metros Split	44	36	
Share w/ Core	77.6%	77.9%	
Share who work in Core	41%	40%	
Share who live in Core	37%	36%	
Connection	Min. Wage Correlation	0.06	0.1	
Mean Wage Correlation	0.91	0.84	
Mean Wage Corr no singles	0.9	0.82	
Containment	Min. Contained	58%	65%	
Mean Contained	88%	88%	
Share of Pop Contained	93%	93%	

Connection

Connection refers to the degree to which counties within a CZ are connected economically. Based on an examination of prior use of CZ delineations this has been most heavily referenced in the economics literature where the idea of a ‘labor-shed’ assumes that within a CZ wages should move together because the workforce can presumably choose to work for any employer within the CZ leading to an equalizing effect on wages within the labor-shed. Following Fowler and Jensen21 I implement a measure of pairwise wage correlation based on five years of BLS wage data so that the average pairwise wage correlation for county i is:pwci=12N∑i∈C∑j∈Cwijpi,j,t

for each county i in commuting zone C and the other counties j in that commuting zone for year t. N is the total count of counties in C. Pairwise correlations are weighted by the share of the resident labor force reslf in counties i and j compared with the resident labor force in all counties k within C such that:wij=reslfi+reslfj2*∑k∈Creslfk

High values for this correlation indicate that ups and downs in wages over a five-year period are highly correlated within CZ’s at an average of 0.91 up from 0.84 in the 2010 delineation. A weakness with this methodology is that CZs with just a single county have, by definition, a perfect correlation so I also provide the correlation with single-county CZ’s removed, which is still high at 0.9. The minimum value across all CZ’s is also provided (“Min. Wage Correlation”). Initial inspection of the counties with low (negative) wage correlation does not reveal any systematic flaws, only counties that are quite different from their neighbors but no more similar to other nearby counties. An example is shown in Fig. 2 where CZ 154 on the Illinois/Missouri/Kentucky border includes Pope County, IL, a county dominated by the Shawnee National Forest and negatively correlated with the rest of the weakly correlated counties in the CZ. There is not a better location for Pope County in the neighboring CZ’s and so this is not a flaw in the delineation but rather a reflection of the fact that CZ’s are designed to be inclusive of rural places and that some places will better fit our conceptual model than others.Fig. 2 Detailed map of CZ 154 in Illinois which has the lowest mean wage correlation in the country. Pope County, Illinois is mostly Shawnee National Forest and is negatively correlated with the rest of the weakly correlated counties in the CZ.

Containment

Finally, containment is the measure most closely associated with the original intent of CZ’s as a unit of analysis for studying rural places and their connections to economic centers. In this conceptual model CZ’s function like watersheds where every drop of rain that falls stays within the watershed. In total 93% of the U.S. population lives and works in the same CZ, a clear fit with this conceptual model. When we look at averages across counties within a CZ this number falls to 88% of the population as it does not account for differences in county population size or the comparatively lower containment rates in lightly populated rural counties. This value is unchanged from the 2010 delineation. The CZ with the lowest mean containment sits at 58%, down substantially from the 65% minimum in 2010. The culprit CZ is 552 in Virginia, which sits South and West of the D.C. metro area and North and West of the Richmond metro area. 552 contains Fauquier County, considered part of the DC metro area and exhibiting a strong negative pairwise wage correlation with the other counties in 552 (−0.28). Fauquier would seemingly benefit from being transferred into CZ 91 just to the North and East, but it gets pulled a little too hard towards the Richmond-centered CZ 543 to the South, and ends up serving as the ‘core’ for its own CZ instead of being attached to either of the larger metro areas. A careful examination of Fig. 3 demonstrates the way this pattern is repeated across the region with smaller CZ’s delineated between larger metro areas rather than being distributed to the major metros. A higher cutoff value in the hierarchical clustering algorithm would tend to reduce this problem, but would over-aggregate counties in other parts of the country where commuting patterns overlap less. There is not a clear misallocation here, only another example of the compromises entailed in any delineation seeking to cover a geography as diverse as is found in the U.S.Fig. 3 Detailed map of CZ 552 in Virginia which has the lowest mean containment in the country. Fauquier County, Virginia is part of the DC metropolitan area but also connected to Richmond, VA.

A comparison of the delineations for 2010 and 2020 confirms the general suitability of the new delineation in terms of the suggested fit metrics. The 2020 delineation appears largely similar to the delineation that preceded it suggesting that it should function reasonably well for longitudinal analysis or for work that seeks to replicate the observational conditions generated under earlier delineations. To this end, the generalized fit statistics should give researchers some degree of comfort that the underlying changes in where and how people work have not fundamentally reshaped the nation’s functional regions.

Data Records

The data associated with this paper are available at 10.17605/OSF.IO/J256U35. Data are provided at the county scale (with 3222 counties assigned to a commuting zone) and at the commuting zone scale (593 commuting zones each assigned an ID). The files are available in two formats; comma separated for ease of use and ESRI shapefile with geometry attached for use in GIS applications. The GIS files are delivered with the data projected to EPSG 5070, U.S. Geological Survey Albers Equal Area Conic projection. Two additional files are made available in csv format containing fit statistics for the county and commuting zones. The fit statistics are those presented in Table 2 as well as others described in Fowler and Jensen21 and are meant to allow users to examine differences within and between commuting zones in greater detail. All files are available at https://sites.psu.edu/psucz/ and can be generated from scratch with the code provided at https://github.com/csfowler/CommutingZones2020.git.

Technical Validation

Three possible issues in the CZ delineations are raised by the similarity measure shown in Fig. 1 and the statistics presented in Tables 1, 2; the decision to retain the method- ology, especially the cutoff value of 0.977, used in prior delineations, the continued presence of non-contiguous commuting zones, and the significant increase in the number of metropolitan areas split by the new delineation. In this section I briefly explore the characteristics of these potential areas of concern.

Optimizing jacand similarity

In the original methodology for delineating CZ’s Tolbert et al.1 used a cutoff value of 0.98 in the hierarchical cluster analysis that defined their delineation. Fowler et al.20 revised this number to 0.977 as a way of approximating the 1990 delineation (which could not be precisely replicated based on the provided methodology) and carried this value forward to the 2010 delineation. For consistency I use that cutoff here to define the 2020 delineation. Since the number itself is arbitrary and intended to create an approximation of the earlier delineation it is worth considering whether alternative values for the cutoff might improve the Jacand similarity across delineations. To that end Fig. 4 shows the results of a similarity test on cutoffs across a range of values from 0.8 to 0.999 and shows the maximum value for Jacand similarity was achieved at a cutoff of 0.972. At this value the mean Jacand similarity across all CZ’s was 0.679 as compared to 0.677 at the 0.977 cutoff. The difference is small but suggests that the cutoff value could be optimized to improve the overall similarity between delineations. However, the difference is small enough that it is unlikely to have a significant impact on the results of most studies and consistency with prior delineations is likely to be more important than the small increase in similarity.Fig. 4 Mean and Standard Deviation of Jacand Similarity for 2010 and 2020 Commuting Zones employing cutoff values from 0.925 to 0.999.

Non-contiguous commuting zones

The initial delineation of commuting zones in 1987 relied on post-hoc expert guidance that moved counties between commuting zones to achieve a result that the analysts felt was more suitable than the raw output of the hierarchical cluster analysis described in the methodology. A key component of this work was to insure that commuting zones were all contiguous. Later revisions in 2000 and 2010 retained the published methodology but dropped the quality control step and allowed for the existence of non-contiguous counties to be part of commuting zones. This is arguably a reasonable decision given the heterogeneity of county sizes and commuting patterns across the United States. It also supports replicability and simplicity in describing the delineation. Nevertheless, the idea of having to drive through a different commuting zone (or multiple zones) to get back to the commuting zone in which you began the trip conflicts with the conceptual model of what a commuting zone is meant to represent. Looking to other kinds of districting problems there is no clear solution; most political districting requires contiguity as a criteria for drawing districts36 while school districts regularly create non-contiguous attendance zones to balance school populations and to promote (or constrain) equitable access to education. This is to say that we cannot a priori reject a delineation because of the presence of non-contiguous districts. An analysis of the non-contiguous counties in the 2020 delineation finds a mixture of anomalies based on relatively small commuting flows, particularly in Alaska, and three commuting zone assignments that appear to be real failures of the method to correctly assign counties.

Figure 5 shows the non-contiguous counties and the commuting zones to which they are assigned. The most significant issue involves San Diego County, California which is inexplicably assigned to a commuting zone to the east of Sacramento that includes South Lake Tahoe and part of Western Nevada. A closer inspection of the commuting data for 2020 shows San Diego containing 1.588 million residents who work within the county but only twelve-thousand who work in neighboring Orange County, seven-thousand in neighboring Riverside County, and six-thousand in nearby Los Angeles County. This allocation seems unlikely in the extreme, but it is what the survey on commuting and the relatively simplistic proportional flow methodology gives us. Of additional concern as far as the methodology goes, there are only sixty eight people making the commute between the northern counties in the CZ and San Diego. Similar phenomena (though impacting a much smaller number of people) are visible in Marion County, Mississippi and Madison County, Florida. In both cases the counties are assigned to commuting zones that are geographically distant with credible alternatives to which they could be assigned immediately adjacent to them. Figure 5 shows the non-contiguous counties and their assigned CZ’s.Fig. 5 Non-Contiguous Counties 2020 Commuting Zones. Non-Contiguous Counties (maroon) with arrows pointing to the commuting zone (in blue) to which they are assigned.

A key goal of this paper is to make the commuting zones entirely consistent with previous delineations AND completely replicable. This means that the non-contiguous counties should be retained, but Table 3 offers a short list of commuting zone reassignments that fix all but one of the non-contiguity problems and result in a more compelling final delineation.Table 3 Suggested alterations to commuting zone delineation to reduce the impact of non-contiguous assignments.

FIPS	County Name	State Name	CZ ID	New CZ ID	Notes	
02060	Bristol Bay	AK	15	25	Add to adjacent CZ	
02185	North Slope	AK	15	15	Retain non-adjacent CZ	
06073	San Diego	CA	59	594	New single county CZ	
08111	San Juan	CO	80	79	Add to adjacent CZ	
12079	Madison	FL	99	92	Add to adjacent CZ	
20067	Grant	KS	77	211	Add to adjacent CZ	
28091	Marion	MS	306	304	Add to adjacent CZ	

Splitting of metropolitan areas

A third area of concern with the revised delineations for 2020 is that they have an increased tendency to split CBSA’s when compared to the 2010 delineation (44 as opposed to 36). This increase is somewhat misleading as five additional CBSA splits occur in Puerto Rico, which was not covered by the earlier delineation. CBSA’s are delineated with broadly similar criteria; including reliance on Core counties (as defined by population) and strength of connection between counties as measured by commuting data so we should expect CZ’s to generally retain these connections between counties. When CZ’s split CBSA’s it is potentially a sign that we have too many CZ’s (they are subdividing functional regions). Conversely, if we programmatically preserve CBSA’s we will obtain a delineation of CZ’s where the functional regions are far larger than intended.

To explore the significance of this issue within the 2020 CZ delineation I examine the CBSA’s that are split by the clustering algorithm. As Figure 6 demonstrates the places where CZ’s split CBSA’s are concentrated in parts of the country with overlapping commuting patterns where the clustering algorithm tends to create smaller CZ’s between metropolitan areas. These smaller CZ’s end up splitting off metropolitan counties with some regularity. The detail in Fig. 3 offers an example that is repeated in multiple places where CBSA’s are split. Splits tend to be in areas where the functional region concept breaks down somewhat with weak ties spread across multiple cores as it does between the DC and Richmond metro areas. If we reduce the number of commuting zones to eliminate these ‘between’ CZs we will end up with commuting zones that are too large and capture multiple major economic cores in one unit. While the smaller size of the CZs is a compromise, this compromise is consistent with past practice and an examination of the split CBSAs does not reveal a clear alternative or improvement.Fig. 6 CBSA’s Split by 2020 Delineation. The 2020 delineation splits 44 CBSA’s, 8 more than the 2010 delineation.

Acknowledgements

The author would like to thank John Cromartie for helpful comments on an earlier draft of this paper.

Author contributions

Christopher Fowler was the sole author of this publication and the code used to generate it.

Code availability

All code used to generate the data from publicly available sources as well as the code for generating this paper with its accompanying figures is available at https://github.com/csfowler/CommutingZones2020.git.

Competing interests

The author declares no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Tolbert, C. M. & Killian, M. S. Labor Market Areas for the United States. Washington D.C. 10.22004/ag.econ.277959 (1987).
2. Killian, M. S. & Tolbert, C. M. Mapping social and economic space: The delineation of local labor market areas in the United States. In: Singelmann, J., Deseran, F. A. (eds) Inequality in Local Labor Markets. Westview Press, Boulder, CO, pp 69–79 (1993).
3. Ferro S Serra C The complex interplay between weather, social activity, and COVID-19 in the US SSM - Population Health 2023 23 101431 10.1016/j.ssmph.2023.101431 37287717
Ferro, S. & Serra, C. The complex interplay between weather, social activity, and COVID-19 in the US. SSM - Population Health 23, 101431, 10.1016/j.ssmph.2023.101431 (2023).37287717 10.1016/j.ssmph.2023.101431
4. Mullachery PH Bilal U Urban scaling of opioid analgesic sales in the United States PLOS ONE 2021 16 e0258526 10.1371/journal.pone.0258526 34637453
Mullachery, P. H. & Bilal, U. Urban scaling of opioid analgesic sales in the United States. PLOS ONE 16, e0258526, 10.1371/journal.pone.0258526 (2021).34637453 10.1371/journal.pone.0258526
5. Fowler CS Kleit RG The Effects of Industrial Clusters on the Poverty Rate Economic Geography 2014 90 129 154 10.1111/ecge.12038
Fowler, C. S. & Kleit, R. G. The Effects of Industrial Clusters on the Poverty Rate. Economic Geography 90, 129–154, 10.1111/ecge.12038 (2014).10.1111/ecge.12038
6. Chantarat T., Mentzer K. M., Van Riper D. C., Hardeman R. R. Where are the labor markets?: Examining the association between structural racism in labor markets and infant birth weight. Health & Place 74,102742, 10.1016/j.healthplace.2022.102742 (2022).
7. Owen AL The Fracking Boom, Labor Structure, and Adolescent Fertility Population Research and Policy Review 2022 41 2211 2231 10.1007/s1111322-09722-6
Owen, A. L. The Fracking Boom, Labor Structure, and Adolescent Fertility. Population Research and Policy Review 41, 2211–2231, 10.1007/s1111322-09722-6 (2022).10.1007/s1111322-09722-6
8. Autor D Dorn D This Job is “Getting Old”: Measuring Changes in Job Opportunities using Occupational Age Structure American Economic Review 2009 99 45 51 10.1257/aer.99.2.45
Autor, D. & Dorn, D. This Job is “Getting Old”: Measuring Changes in Job Opportunities using Occupational Age Structure. American Economic Review 99, 45–51, 10.1257/aer.99.2.45 (2009).10.1257/aer.99.2.45
9. Van Sandt A Carpenter CW Tolbert CM Decomposing local bank impacts with demand thresholds The Annals of Regional Science 2023 70 333 352 10.1007/s00168-022-01148-4
Van Sandt, A., Carpenter, C. W. & Tolbert, C. M. Decomposing local bank impacts with demand thresholds. The Annals of Regional Science 70, 333–352, 10.1007/s00168-022-01148-4 (2023).10.1007/s00168-022-01148-4
10. Foote A Kutzbach MJ Vilhuber L Recalculating…: How Uncertainty in Local Labour Market Definitions Affects Empirical Findings Applied Economics 2021 53 1598 1612 10.1080/00036846.2020.1841083
Foote, A., Kutzbach, M. J. & Vilhuber, L. Recalculating…: How Uncertainty in Local Labour Market Definitions Affects Empirical Findings. Applied Economics 53, 1598–1612, 10.1080/00036846.2020.1841083 (2021).10.1080/00036846.2020.1841083
11. Dash Nelson G Rae A An Economic Geography of the United States: From Commutes to Megaregions PLOS ONE 2016 11 e0166083 10.1371/journal.pone.0166083 27902707
Dash Nelson, G. & Rae, A. An Economic Geography of the United States: From Commutes to Megaregions. PLOS ONE 11, e0166083, 10.1371/journal.pone.0166083 (2016).27902707 10.1371/journal.pone.0166083
12. Karlsson C Olsson M The identification of functional regions: Theory, methods, and applications The Annals of Regional Science 2006 40 1 18 10.1007/s00168-005-0019-5
Karlsson, C. & Olsson, M. The identification of functional regions: Theory, methods, and applications. The Annals of Regional Science 40, 1–18, 10.1007/s00168-005-0019-5 (2006).10.1007/s00168-005-0019-5
13. Coombes M From City-region Concept to Boundaries for Governance: The English Case Urban Studies 2014 51 2426 2443 10.1177/0042098013493482
Coombes, M. From City-region Concept to Boundaries for Governance: The English Case. Urban Studies 51, 2426–2443, 10.1177/0042098013493482 (2014).10.1177/0042098013493482
14. Halás M Klapka P Tonev P The use of migration data to define functional regions: The case of the Czech Republic Applied Geography 2016 76 98 105 10.1016/J.APGEOG.2016.09.010
Halás, M., Klapka, P. & Tonev, P. The use of migration data to define functional regions: The case of the Czech Republic. Applied Geography 76, 98–105, 10.1016/J.APGEOG.2016.09.010 (2016).10.1016/J.APGEOG.2016.09.010
15. Halás, M. et al. A definition of relevant functional regions for international comparisons: The case of Central Europe. Area. 10.1111/area.12487 (2018).
16. Halás M Klapka P Erlebach M Unveiling spatial uncertainty: A method to evaluate the fuzzy nature of functional regions Regional Studies 2019 53 1029 1041 10.1080/00343404.2018.1537483
Halás, M., Klapka, P. & Erlebach, M. Unveiling spatial uncertainty: A method to evaluate the fuzzy nature of functional regions. Regional Studies 53, 1029–1041, 10.1080/00343404.2018.1537483 (2019).10.1080/00343404.2018.1537483
17. Flórez-Revuelta F Casado-Díaz JM Martínez-Bernabeu L An evolutionary approach to the delineation of functional areas based on travel-to-work flows International Journal of Automation and Computing 2008 5 10 21 10.1007/s11633-008-0010-6
Flórez-Revuelta, F., Casado-Díaz, J. M. & Martínez-Bernabeu, L. An evolutionary approach to the delineation of functional areas based on travel-to-work flows. International Journal of Automation and Computing 5, 10–21, 10.1007/s11633-008-0010-6 (2008).10.1007/s11633-008-0010-6
18. Casado-Díaz JM Martínez-Bernabéu L Flórez-Revuelta F Automatic parameter tuning for functional regionalization methods Papers in Regional Science 2016 96 859 880 10.1111/pirs.12199
Casado-Díaz, J. M., Martínez-Bernabéu, L. & Flórez-Revuelta, F. Automatic parameter tuning for functional regionalization methods. Papers in Regional Science 96, 859–880, 10.1111/pirs.12199 (2016).10.1111/pirs.12199
19. Kropp P Schwengler B Three-Step Method for Delineating Functional Labour Market Regions Regional Studies 2016 50 429 445 10.1080/00343404.2014.923093
Kropp, P. & Schwengler, B. Three-Step Method for Delineating Functional Labour Market Regions. Regional Studies 50, 429–445, 10.1080/00343404.2014.923093 (2016).10.1080/00343404.2014.923093
20. Fowler CS Rhubart DC Jensen L Reassessing and Revising Commuting Zones for 2010: History, Assessment, and Updates for U.S. Labor-Sheds 1990-2010 Population Research and Policy Review 2016 35 263 286 10.1007/s11113-016-9386-0
Fowler, C. S., Rhubart, D. C. & Jensen, L. Reassessing and Revising Commuting Zones for 2010: History, Assessment, and Updates for U.S. Labor-Sheds 1990-2010. Population Research and Policy Review 35, 263–286, 10.1007/s11113-016-9386-0 (2016).10.1007/s11113-016-9386-0
21. Fowler CS Jensen L Bridging the gap between geographic concept and the data we have: The case of labor markets in the USA Environment and Planning A: Economy and Space 2020 52 1395 1414 10.1177/0308518X20906154
Fowler, C. S. & Jensen, L. Bridging the gap between geographic concept and the data we have: The case of labor markets in the USA. Environment and Planning A: Economy and Space 52, 1395–1414, 10.1177/0308518X20906154 (2020).10.1177/0308518X20906154
22. Fowler CS Cromartie J The Role of Data Sample Uncertainty in Delineations of Core Based Statistical Areas and Rural Urban Commuting Areas Spatial Demography 2023 11 6 10.1007/s40980-023-00118-4
Fowler, C. S. & Cromartie, J. The Role of Data Sample Uncertainty in Delineations of Core Based Statistical Areas and Rural Urban Commuting Areas. Spatial Demography 11, 6, 10.1007/s40980-023-00118-4 (2023).10.1007/s40980-023-00118-4
23. Franklin RS The demographic burden of population loss in US cities, 2000– 2010 Journal of Geographical Systems 2021 23 209 230 10.1007/s10109-019-00303-4
Franklin, R. S. The demographic burden of population loss in US cities, 2000– 2010. Journal of Geographical Systems 23, 209–230, 10.1007/s10109-019-00303-4 (2021).10.1007/s10109-019-00303-4
24. Bagchi-Sen S Franklin RS Rogerson P Seymour E Urban inequality and the demographic transformation of shrinking cities: The role of the foreign born Applied Geography 2020 116 102168 10.1016/j.apgeog.2020.102168
Bagchi-Sen, S., Franklin, R. S., Rogerson, P. & Seymour, E. Urban inequality and the demographic transformation of shrinking cities: The role of the foreign born. Applied Geography 116, 102168, 10.1016/j.apgeog.2020.102168 (2020).10.1016/j.apgeog.2020.102168
25. Plane DA The geography of urban commuting fields: Some empirical evidence from new England Professional Geographer 1981 33 182 188 10.1111/j.0033-0124.1981.00182.x
Plane, D. A. The geography of urban commuting fields: Some empirical evidence from new England. Professional Geographer 33, 182–188, 10.1111/j.0033-0124.1981.00182.x (1981).10.1111/j.0033-0124.1981.00182.x
26. Office of Management and Budget 2020 Standards for Delineating Core Based Statistical Areas. Federal Register https://www.federalregister.gov/documents/2021/07/16/2021-15159/2020-standards-for-delineating-core-based-statistical-areas Accessed 8/7/2024 (2021).
27. U.S. Census Bureau 2020 Census Qualifying Urban Areas and Final Criteria Clarifications. Federal Register https://www.federalregister.gov/documents/2022/12/29/2022-28286/2020-census-qualifying-urban-areas-and-final-criteria-clarifications Accessed 8/7/2024 (2022).
28. U.S. Census Bureau Commuting (Journey to Work) https://www2.census.gov/programs-surveys/demo/tables/metro-micro/2020/commuting-flows-2020/ Accessed 8/7/2024 (2020).
29. U. S. Census Bureau Micropolitan and Metropolitan Delineation Files https://www.census.gov/geographies/reference-files/time-series/demo/metro-micro/delineation-files.html Accessed 8/7/2024 (2023).
30. Walker, K. & Herman, M. Tidycensus: Load US Census Boundary and Attribute Data https://CRAN.R-project.org/package=tidycensus Accessed 8/7/2024 (2022).
31. Bureau of Labor Statistics Quarterly Census of Employment and Wages https://www.bls.gov/cew/ Accessed 8/7/2024 (2020).
32. Tolbert, C. & Sizer, M. U.S. Commuting Zones and Labor Market Areas: A 1990 Update. Washington D.C. 10.22004/ag.econ.278812 (1990).
33. Kaufman, L. & Rousseeuw, P. J. Finding groups in data: An introduction to cluster analysis. John Wiley & Sons., New York (2009).
34. Pipa, A. F. & Geismar, N. The new “rural”? The implications of OMB’s proposal to redefine nonmetro America. Brookings https://www.brookings.edu/articles/the-new-rural-the-implications-of-ombs-proposal-to-redefine-nonmetro-america/ Accessed 8/7/2024 (2021).
35. Fowler CS Commuting Zones for 2020 2024 10.17605/OSF.IO/J256U
Fowler, C. S. Commuting Zones for 202010.17605/OSF.IO/J256U (2024).10.17605/OSF.IO/J256U
36. Webster GR Reflections on current criteria to evaluate redistricting plans Political Geography 2013 32 3 14 10.1016/j.polgeo.2012.10.004
Webster, G. R. Reflections on current criteria to evaluate redistricting plans. Political Geography 32, 3–14, 10.1016/j.polgeo.2012.10.004 (2013).10.1016/j.polgeo.2012.10.004
