
==== Front
Commun Biol
Commun Biol
Communications Biology
2399-3642
Nature Publishing Group UK London

38374434
5708
10.1038/s42003-023-05708-y
Article
The frequency of pathogenic variation in the All of Us cohort reveals ancestry-driven disparities
http://orcid.org/0000-0002-3319-6690
Venner Eric venner@bcm.edu

1
http://orcid.org/0000-0002-0853-1231
Patterson Karynne 2
Kalra Divya 1
Wheeler Marsha M. 2
Chen Yi-Ju 1
http://orcid.org/0000-0003-2437-5352
Kalla Sara E. 1
Yuan Bo 1
http://orcid.org/0000-0001-5001-3334
Karnes Jason H. 34
http://orcid.org/0000-0001-9686-5217
Walker Kimberly 1
Smith Joshua D. 2
McGee Sean 2
Radhakrishnan Aparna 2
Haddad Andrew 5
http://orcid.org/0000-0001-7474-2339
Empey Philip E. 6
Wang Qiaoyan 1
Lichtenstein Lee 7
Toledo Diana 7
Jarvik Gail 89
http://orcid.org/0000-0001-7770-299X
Musick Anjene 10
Gibbs Richard A. 1
on behalf of the All of Us Research Program InvestigatorsAhmedani Brian 11
Johnson Christine D. Cole 11
Ahsan Habib 12
Anton-Culver Hoda 13
Topol Eric 14
Baca-Motes Katie 14
Moore-Vogel Julia 14
Jain Praduman 15
Begale Mark 15
Jain Neeta 15
Klein David 15
Sutherland Scott 15
Korf Bruce 16
Lewis Beth 16
Gharavi Ali G. 17
Hripcsak George 17
Boerwinkle Eric 18
Hebbring Scott Joseph 19
Burnside Elizabeth 20
Farrar-Edwards Dorothy 20
Taylor Amy 21
Desa Liliana Lombardi 22
Thibodeau Steve 23
Cicek Mine 23
Schlueter Eric 24
Holmes Beverly Wilson 24
Daviglus Martha 25
Harris Paul 26
Wilkins Consuelo 26
Dan Roden 26
Doheny Kim 27
Eichler Evan 28
Jarvik Gail 28
Funk Gretchen 29
Philippakis Anthony 30
Rehm Heidi 30
Gabriel Stacey 30
Gibbs Richard 31
Rico Edgar M. Gil 32
Glazer David 33
Burke Jessica 34
Greenland Philip 35
Shenkman Elizabeth 36
Hogan William R. 36
Igho-Pemu Priscilla 37
Karlson Elizabeth W. 38
Smoller Jordan 38
Murphy Shawn N. 38
Ross Margaret Elizabeth 39
Kaushal Rainu 39
Winford Eboni 40
Kheterpal Vik 41
Moreno Francisco A. 42
Thomas Cheryl 43
Lunn Mitchell 44
Obedin-Maliver Juno 44
Marroquin Oscar 45
Visweswaran Shyam 45
Reis Steven 45
McGovern Patrick 46
Talavera Gregory 47
O’Connor George T. 48
Ohno-Machado Lucila 49
Randal Fornessa 50
Theodorou Andreas A. 51
Reiman Eric 51
Roxas-Murray Mercedita 52
Stark Louisa 53
Tepp Ronnie 54
Zhou Alicia 55
Topper Scott 55
Trousdale Rhonda 56
Tsao Phil 57
Weiss Scott T. 58
Whittle Jeffrey 59
Zuchner Stephan 60
Carrasquillo Olveen 60
Lewis Megan 61
Uhrig Jen 61
Okihiro May 62
Argos Maria 25
Aschebook-Kilfoy Brisa 25
Bartlett Laura 63
Carlin Roberta 64
Cohn Elizabeth 65
Colon-Lopez Vivian 66
Cooper Karl 64
Cottler Linda 67
Crook Errol 68
Culler Elizabeth 69
Drum Charles 64
Eder Milton 67
Edmunds Mark 70
Everhart Rachel 71
Falcon Adolph 32
Fein Becky 72
Frano Zeno 59
Garrett Michael 73
Halverson Sandra 74
Handberg Eileen 36
Ho Joyce 35
Horne Laura 72
Isasi Rosario 60
Isom Jessica 75
Jarmin Jessica 76
Jula Megan 77
Kamyar Royan 78
Kleiman Frida 65
Kohane Isaac 79
Lamarca Babbette 73
Lee Brendan 31
Lennon Niall 30
Levy Dessie 80
Mahr Todd 81
Makahi Emily 62
Marshall Vivienne 82
Mayer-Davis Elizabeth 83
McCauley Jacob 60
McKinney Jeffrey 84
McPherson David 18
Meller Robert 37
Melo Jose 66
Lin David Ming-Hung 85
Minor Michael 80
Muse Evan 14
Parakh Kapil 86
Peltz-Rauchman Cathryn 11
Laras Linda Perez 87
Raveendran Subhara 88
Reilly Gail 40
Reilly Jody 89
Rivera Nelida 87
Rosales Laura 31
Rosser Tracie 90
Salgin Linda 47
Sawyer Sherilyn 91
Simonson William 92
Sitapati Amy 49
So-Armah Cynthia 75
Stegeman Gene 93
Suver Christin 94
Taitel Michael 95
Taylor Kyla 40
Tinoco Daniel Hernandez 40
Vassy Jason 91
Walz Jamie 84
Watkins Preston 96
Wilkerson Blaker 97
Yamazaki Katrina 21
Basford Melissa 26
Boschetti Amaryllis Silva 49
Breeden Matthew 98
Chandrasekaran Suchitra 90
Clark Cheryl 58
Enard Kim 98
Fresko Yuri 89
Grucza Richard 98
Kelley Robert 90
Keogh Kathleen 22
Kraft Monica 99
Lough Christopher 100
Malmstrom Ted 98
Nemeskal Paul 75
Pagel Matt 90
Scherrer Jeffrey 98
Skukla Sanjay 19
Smith Debra 101
Turner Bryce 102
Vos Miriam 90

1 https://ror.org/02pttbw34 grid.39382.33 0000 0001 2160 926X Human Genome Sequencing Center, Baylor College of Medicine, Houston, TX USA
2 https://ror.org/00cvxb145 grid.34477.33 0000 0001 2298 6657 Department of Genome Sciences, University of Washington, Seattle, WA USA
3 https://ror.org/03m2x1q45 grid.134563.6 0000 0001 2168 186X University of Arizona, R Ken Coit College of Pharmacy, Department of Pharmacy Practice and Science, Tucson, AZ USA
4 https://ror.org/05dq2gs74 grid.412807.8 0000 0004 1936 9916 Vanderbilt University Medical Center, Department of Biomedical Informatics, Boston, MA USA
5 https://ror.org/01an3r305 grid.21925.3d 0000 0004 1936 9000 Department of Pharmaceutical Sciences, University of Pittsburgh School of Pharmacy, Pittsburgh, PA USA
6 https://ror.org/01an3r305 grid.21925.3d 0000 0004 1936 9000 Department of Pharmacy and Therapeutics, University of Pittsburgh School of Pharmacy, Pittsburgh, PA USA
7 https://ror.org/05a0ya142 grid.66859.34 0000 0004 0546 1623 Broad Institute of MIT and Harvard, Cambridge, MA USA
8 grid.34477.33 0000000122986657 Department of Medicine (Medical Genetics), University of Washington School of Medicine, Seattle, WA USA
9 grid.34477.33 0000000122986657 Department of Genome Sciences, University of Washington School of Medicine, Seattle, WA USA
10 grid.453125.4 0000 0004 0533 8641 NIH All of Us Research Program, National Institutes of Health Office of the Director, Bethesda, MD USA
11 https://ror.org/02kwnkm68 grid.239864.2 0000 0000 8523 7701 Henry Ford Health System, Detroit, MI USA
12 https://ror.org/0076kfe04 grid.412578.d 0000 0000 8736 9513 University of Chicago Medical Center, Chicago, IL USA
13 https://ror.org/05t99sp05 grid.468726.9 0000 0004 0486 2046 University of California, Irvine - Irvine, California, CA USA
14 grid.214007.0 0000000122199231 Scripps Research Translational Institute - La Jolla, California, CA USA
15 Vibrent Health, Fairfax, VA USA
16 https://ror.org/008s83205 grid.265892.2 0000 0001 0634 4187 University of Alabama at Birmingham, Birmingham, AL USA
17 https://ror.org/00hj8s172 grid.21729.3f 0000 0004 1936 8729 Columbia University - New York City, New York, NY USA
18 https://ror.org/03gds6c39 grid.267308.8 0000 0000 9206 2401 University of Texas Health Science Center at Houston, Houston, TX USA
19 grid.280718.4 0000 0000 9274 7048 Marshfield Clinic Research Institute, Marshfield, WI USA
20 https://ror.org/01y2jtd41 grid.14003.36 0000 0001 2167 3675 University of Wisconsin at Madison, Madison, WI USA
21 https://ror.org/05ewbqm54 grid.428181.6 Community Health Center, Inc., Middletown, CT USA
22 Sun River Health - Beacon, New York, NY USA
23 grid.66875.3a 0000 0004 0459 167X Mayo Clinic and Foundation, Rochester, MN USA
24 Cooperative Health, South Carolina, SC USA
25 grid.185648.6 0000 0001 2175 0319 University of Illinois at Chicago, Evanston, IL USA
26 https://ror.org/05dq2gs74 grid.412807.8 0000 0004 1936 9916 Vanderbilt University Medical Center, Nashville, TN USA
27 grid.21107.35 0000 0001 2171 9311 Johns Hopkins University School of Medicine, Baltimore, MD USA
28 https://ror.org/00cvxb145 grid.34477.33 0000 0001 2298 6657 University of Washington, Seattle, Washington, WA USA
29 FiftyForward, Nashville, TN USA
30 grid.66859.34 0000 0004 0546 1623 Broad Institute - Boston Massachusetts, Boston, USA
31 https://ror.org/02pttbw34 grid.39382.33 0000 0001 2160 926X Baylor College of Medicine, Houston, TX USA
32 https://ror.org/050dg4y92 grid.422141.0 0000 0004 0615 1598 National Alliance for Hispanic Health, Washington, DC USA
33 grid.497059.6 Verily Life Sciences - South San Francisco, California, CA USA
34 grid.420015.2 0000 0004 0493 5049 MITRE Corporation, McLean, VA USA
35 https://ror.org/000e0be47 grid.16753.36 0000 0001 2299 3507 Northwestern University, Evanston, IL USA
36 https://ror.org/02y3ad647 grid.15276.37 0000 0004 1936 8091 University of Florida, Gainesville, FL USA
37 https://ror.org/01pbhra64 grid.9001.8 0000 0001 2228 775X Morehouse School of Medicine, Atlanta, GA USA
38 grid.417182.9 0000 0004 5899 4861 Partners Health Care - Boston, Massachusetts, MA USA
39 https://ror.org/05bnh6r87 grid.5386.8 0000 0004 1936 877X Cornell University, Weill Medical College - New York City, New York, NY USA
40 Cherokee Health Systems, Knoxville, TN USA
41 https://ror.org/04h5v2n16 grid.511652.4 CareEvolution, Inc, Ann Arbor, MI USA
42 https://ror.org/03m2x1q45 grid.134563.6 0000 0001 2168 186X University of Arizona, Tucson, Tucson, AZ USA
43 https://ror.org/0081myy06 grid.475121.4 0000 0000 9496 947X Delta Research and Educational Foundation, Washington, DC USA
44 https://ror.org/00f54p054 grid.168010.e 0000 0004 1936 8956 Stanford University - Stanford, California, CA USA
45 https://ror.org/01an3r305 grid.21925.3d 0000 0004 1936 9000 University of Pittsburgh, Pittsburgh, PA USA
46 Wondros - Los Angeles, California, CA USA
47 grid.428482.0 0000 0004 0616 2975 San Ysidro Health Center - San Ysidro, California, CA USA
48 grid.239424.a 0000 0001 2183 6745 Boston Medical Center - Boston, Massachusetts, MA USA
49 grid.266100.3 0000 0001 2107 4242 University of California, San Diego – San Diego, California, CA USA
50 https://ror.org/00yt7g523 grid.432281.e Asian Health Coalition, Chicago, IL USA
51 https://ror.org/039wwwz66 grid.418204.b 0000 0004 0406 4925 Banner Health, Phoenix, AZ USA
52 Montage Marketing Group, Rockville, MD USA
53 https://ror.org/03r0ha626 grid.223827.e 0000 0001 2193 0096 University of Utah, Salt Lake City, Utah USA
54 https://ror.org/03t8z4688 grid.479349.4 HCM Strategists, Austin, TX USA
55 https://ror.org/031ghq268 grid.512147.4 Color Genomics, Inc. – Burlingame, California, CA USA
56 grid.422616.5 0000 0004 0443 7226 NYC Health + Hospitals - New York City, New York, NY USA
57 VA AoU Coordinating Center - Palo Alto, California, CA USA
58 https://ror.org/04b6nzv94 grid.62560.37 0000 0004 0378 8294 Brigham and Women’s Hospital – Boston, Massachusetts, MA USA
59 https://ror.org/00qqv6244 grid.30760.32 0000 0001 2111 8460 Medical College of Wisconsin, Milwaukee, WI USA
60 https://ror.org/02dgjyy92 grid.26790.3a 0000 0004 1936 8606 University of Miami School of Medicine, Miami, FL USA
61 grid.62562.35 0000000100301493 Research Triangle Institute - Research Triangle Park, Raleigh, NC USA
62 Waianae Coast CHC, Waianae, HI USA
63 grid.280285.5 0000 0004 0507 7840 National Library of Medicine (NLM), Bethesda, MD USA
64 https://ror.org/00x584494 grid.427592.f American Association of Health and Disability, Sedona, AZ USA
65 grid.257167.0 0000 0001 2183 6649 Hunter College - New York City, New York, NY USA
66 https://ror.org/05asdy483 0000 0004 0611 0614 University of Puerto Rico Comprehensive Cancer Center, San Juan, WA USA
67 CTSA Community Engagement Programs, Gainesville, FL USA
68 https://ror.org/01s7b5y08 grid.267153.4 0000 0000 9552 1255 University of South Alabama, Mobile, AL USA
69 TPC: Blood Assurance - Signal Mountain, Tennessee, TN USA
70 San Diego Blood Bank – San Diego, California, CA USA
71 grid.239638.5 0000 0001 0369 638X TPC: Denver Health - Denver, Colorado, CO USA
72 TPC: Active Minds -, Washington, DC USA
73 https://ror.org/044pcn091 grid.410721.1 0000 0004 1937 0407 University of Mississippi Medical Center, Jackson, MS USA
74 TPC: DLH Corp, Atlanta, GA USA
75 grid.32224.35 0000 0004 0386 9924 Mass General Hospital - Boston, Massachusetts, MA USA
76 Tactis -, Washington, DC USA
77 https://ror.org/05pjy6894 grid.429448.7 0000 0004 0558 7293 TPC: Mary’s Center -, Washington, DC USA
78 https://ror.org/020tbqw59 grid.455925.a TPC: Owaves -, Washington, DC USA
79 grid.38142.3c 000000041936754X Harvard Medical School - Boston, Massachusetts, MA USA
80 National Baptist Convention, Nashville, TN USA
81 https://ror.org/01p3c3c27 grid.413464.0 0000 0000 9478 5072 Gundersen Health System, La Crosse, WI USA
82 https://ror.org/02jx0eg57 grid.477992.6 South Texas Blood and Tissue Center, San Antonio, TX USA
83 grid.10698.36 0000000122483208 University of North Carolina at Chapel Hill - Chapel Hill, Chapel Hill, NC USA
84 Sensis - Glendale, California, CA USA
85 https://ror.org/00ek9j693 grid.280646.e 0000 0004 6016 0057 TPC: Bloodworks Northwest – Seattle, Washington, WA USA
86 TPC: Fitbit - San Francisco, California, CA USA
87 COSSMA, Aibonito, PR Puerto Rico
88 Patients Like Me, Houston, TX USA
89 https://ror.org/010g9bb70 grid.418124.a 0000 0004 0462 1752 Quest Diagnostics Incorporated – Secaucus, Secaucus, NJ USA
90 https://ror.org/03czfpz43 grid.189967.8 0000 0004 1936 7398 Emory University, Atlanta, GA USA
91 VA AoU Coordinating Center - Boston, Massachusetts, MA USA
92 Cascade Regional Blood Services - Tacoma, Washington, WA USA
93 ExamOne, Mission, KS USA
94 https://ror.org/049ncjx51 grid.430406.5 0000 0004 6023 5303 Sage Bionetworks – Seattle, Washington, WA USA
95 grid.497251.c 0000 0004 4684 2094 Walgreen Co., Deerfield, IL USA
96 WebMD Health Corp – New York City, New York, NY USA
97 grid.469680.5 0000 0004 0388 194X Blue Cross Blue Shield, Chicago, IL USA
98 https://ror.org/01p7jjy08 grid.262962.b 0000 0004 1936 9342 Saint Louis University, Saint Louis, MO USA
99 grid.425214.4 0000 0000 9963 6690 Mount Sinai Health System - New York City, New York, NY USA
100 LifeSouth, Gainesville, FL USA
101 SunCoast Blood Center, Bradenton, FL USA
102 https://ror.org/0294hxs80 grid.253561.6 0000 0001 0806 2909 University Southern California - Los Angeles, California, CA USA
19 2 2024
19 2 2024
2024
7 1748 8 2022
13 12 2023
© The Author(s) 2024
https://creativecommons.org/licenses/by/4.0/ Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons license, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons license and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/.
Disparities in data underlying clinical genomic interpretation is an acknowledged problem, but there is a paucity of data demonstrating it. The All of Us Research Program is collecting data including whole-genome sequences, health records, and surveys for at least a million participants with diverse ancestry and access to healthcare, representing one of the largest biomedical research repositories of its kind. Here, we examine pathogenic and likely pathogenic variants that were identified in the All of Us cohort. The European ancestry subgroup showed the highest overall rate of pathogenic variation, with 2.26% of participants having a pathogenic variant. Other ancestry groups had lower rates of pathogenic variation, including 1.62% for the African ancestry group and 1.32% in the Latino/Admixed American ancestry group. Pathogenic variants were most frequently observed in genes related to Breast/Ovarian Cancer or Hypercholesterolemia. Variant frequencies in many genes were consistent with the data from the public gnomAD database, with some notable exceptions resolved using gnomAD subsets. Differences in pathogenic variant frequency observed between ancestral groups generally indicate biases of ascertainment of knowledge about those variants, but some deviations may be indicative of differences in disease prevalence. This work will allow targeted precision medicine efforts at revealed disparities.

A comparison of the frequency of pathogenic mutations in 73 genes in the All of Us cohort highlights the differences in pathogenic variation attributed to ancestry.

Subject terms

Disease genetics
Molecular medicine
Genome informatics
issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Implementing genomic medicine will require interpreting genomic variation in real-world clinical populations. At present, lack of diversity in large genomics cohorts is a widely recognized problem1–4. Most sequencing studies have focused on European ancestry populations5 and it is predicted that much of the pathogenic variation present in the general population is specific to an ancestral population6,7. Overcoming this gap in diagnostic yield will necessitate collecting diverse genomic data paired with electronic health record data8,9.

To advance precision medicine, the All of Us Research Program from the National Institutes of Health (NIH) is generating a unique dataset which includes genetic, electronic health record and survey data from a diverse participant cohort10. All of Us is targeting a cohort size of more than one million, with a focus on individuals who have been traditionally underserved by biomedical research10. Whole genome sequence data is generated at one of three All of Us Genome Centers located at the Baylor College of Medicine, the Broad Institute and the Northwest Genomics Center at the University of Washington. Data are transferred via the All of Us Data and Research Center at Vanderbilt University to Clinical Validation Laboratories at the Baylor College of Medicine, Northwest Genomics Center and Color Genomics, where sequence data are interpreted for health-related reporting. Data are also deposited at the Data Resource Center for further processing and sharing with qualified researchers via the All of Us Researcher Workbench, a cloud-computing platform. The data generation and results workflow are conducted under an investigational device exemption through the FDA11.

A strength of the All of Us Research Program is the diversity of the participants, relative to those in previously studied large cohorts. The All of Us Researcher Workbench, a cloud computing platform, contains whole-genome sequencing data from 98,590 participants (10.1038/s41586-023-06957-x) in a data release called alpha3. Based on genetically predicted ancestry (see Methods), this dataset included 49,668 (50.4%) participants with predominantly European ancestry, 22,897 (23.2%) participants with predominantly African ancestry, 15,893 (16.1%) Latino/Admixed American ancestry participants, 2,113 (2.1%) East Asian ancestry participants, 940 (1.0%) South Asian ancestry participants, 193 (0.2%) Middle Eastern participants and 6,886 (7%) participants which do not group unambiguously and were designated as Other. In contrast, the UK Biobank project contains 94% European ancestry individuals, the Million Veterans project contains 77% and eMERGE contains 73%.5

However, it is currently unknown whether the frequency of pathogenic variants in genes conferring appreciable health risks in the All of Us cohort will differ from those in previously ascertained healthy populations due to the unprecedented diversity of participants enrolled in the program. Identification of such differences will be a powerful reinforcement of the importance of the program’s strategy for recruitment and engagement of participants from underrepresented groups.

To examine the frequency of pathogenic genomic variation in participants we examined data from a set of 73 genes that harbor actionable secondary findings12. These genes are associated with diseases, including hereditary breast cancer, hemochromatosis, dislipidemias and cardiomyopathies and represent some of the most well-studied targets for genomic medicine12. We annotated each participant with their calculated genetic ancestry and searched for known pathogenic variants with established criteria for pathogenicity, aided by databases of curated variants from previous clinical genomic projects13,14, as well as an early de-identified review of All of Us data. The preliminary results from analysis of data from more than 98,000 All of Us participants showed variability in the rates of pathogenic variants between ancestry groups, prompting further analysis of the source of the differences.

Results

Rates of pathogenic variation in the all of Us dataset

To understand the rates of previously-known pathogenic variants broken down by predicted ancestry groups in the All of Us data, we used the ‘VIP’ database to annotate P/LP variants15 present in whole genome sequencing data in All of Us participants across 73 genes with actionable secondary findings12. This database contains variants curated by the HGSC-CL variant interpretation group during projects such as eMERGE III13 and HeartCare14 as well as an initial assessment of de-identified variants from the All of Us cohort itself. Figure 1a shows that the European ancestry group has the highest rate of previously known pathogenic variants (2.13%), followed by the “Other” group at 1.82% and the African ancestry group at 1.52%. Using a Chi-square test for independence, the rates of pathogenic variants are significantly different between the African ancestry, European and Admixed American/Latino groups (p < 0.00001, n = 1757 total observed P/LP variants, Supplementary Table 1), which have sufficient data to perform this test. These differences are significant even when leaving out the gene HFE, which has a large known difference in pathogenic rates between ancestry groups. These significant differences in the rates of pathogenic variation could be explained either by a bias in the ascertainment of pathogenic variation in the variant database or by differences in the underlying disease prevalence between ancestry groups.Fig. 1 Pathogenic variants by ancestry.

Using a database of known pathogenic mutations and annotations for rare, pLoF variants, we searched the beta release of the All of Us cohort for pathogenic variants, on the Researcher Workbench. Figure 1a shows the rates of pathogenic variation, broken down by predicted genetic ancestry groups. Error bars show 95% the confidence intervals for the total set of pathogenic variants (including both VIP P/LP variants and rare pLoF). Figure 1b shows the breakdown of pathogenic variants by disease area. The blue line and bar depict the rate of Pathogenic and Likely pathogenic variants, the gray bar the rate of novel, predicted loss of function variants and the orange bar depicts predicted loss of function variants that were known pathogenic variants at the time of analysis. The yellow line shows the total variants in each ancestry group.

To gain a more comprehensive understanding of pathogenic and likely pathogenic variants in participants, we incorporated rare (GnomAD popmax allele frequency below 0.001), predicted loss-of-function (pLoF) variants into the prior analysis of known pathogenic variants (Supplementary Table 2). We focused on 38 specific genes where loss-of-function is a recognized cause of disease, and we also looked for any overlap with known pathogenic variation (refer to Fig. 1a, orange and gray bars). Our findings revealed 1114 variants, comprising 562 frameshifts, 112 splice acceptor variants, 100 splice donor variants, and 340 stop gain variants (see Supplementary Table 3 for details). Many of these pLoF variants overlap the pathogenic/likely pathogenic variants in the VIP database (Fig. 1a, b, orange bars). For example, in the group with European ancestry, the overlap was 0.46%. Including pLoF variants in our analysis increases the overall rates of positive variants, ranging from 2.26% in the European ancestry group to 1.32% in the Latino / Admixed American group, with smaller ancestry groups having wider confidence intervals.

When looking purely at rates of rare pLoF variants, the European ancestry group had a lower rate of findings than both the South Asian ancestry group (0.85%) and the East Asian ancestry group (0.62%), though this may be due to using GnomAD, which contains samples from predominantly European ancestry, as a filter. The variation of pLoF variants between different ancestry groups was less than that of the pathogenic/likely pathogenic variants (pLoF standard deviation of variant rates was 0.003, versus a standard deviation of 0.004 for P/LP variant rates). Given that interpreting novel LoF variants can be simplified and automated, these findings suggest that the detection of pathogenic variants in studies with a large number of participants with European ancestry may be contributing to some of the differences observed between groups.

To evaluate whether known variants with ancestral divergence are replicated in the All of Us cohort, we examined the frequency of the rs334 mutation in HBB, known to be associated with sickle cell disease, and the APOL1 G1 and G2 alleles (Table 1). The results confirm the expected ancestral divergences, with non-reference alleles appearing 17,969 times within the 22,897 participants of African ancestry, and 68 times among the 49,668 participants of European ancestry.Table 1 Non-reference sample counts in known ancestrally divergent genes.

	African	Latino / Admixed American	East Asian	European	Middle Eastern	Other	South Asian	
APOL1 (G1 & G2 alleles)	17,969/22,897 (78.48%)	1041/15,893 (6.55%)	1/2113 (0.05%)	60/49,668 (0.12%)	1/193 (0.52%)	1166/6886 (16.93%)	2/940 (0.21%)	
HBB (rs334)	2059/22,897 (8.99%)	229/15,893 (1.44%)	0/2113 (0.00%)	6/49,668 (0.01%)	0/193 (0.00%)	202/6886 (2.93%)	2/940 (0.21%)	
To confirm the presence of known alleles with ancestral divergence, we examined the APOL1 G1 and G2 alleles, as well as the rs334 SNP in the HBB gene. The counts in the table above include all instances of the rs73885319, rs60910145 SNPS (APOL1 G1 allele) as well as the G2 deletion (rs71785313) and the rs334 allele in the HBB gene, associated with Sickle-cell disease. Participants may carry more than one allele. These variants are present at a higher rate in participants with African ancestry, as expected.

In order to understand which genes have divergent rates of pathogenic variants between ancestry groups, we normalized pathogenic rates of all genes against the European ancestry group’s pathogenic rate and checked for outliers in population proportions using a Bonferoni-corrected z-test. Several genes show significant differences from the European ancestry group’s rates (Table 2). Interestingly, the genes HFE and PALB2 differences are known16,17 but the differences in the gene PKP2 have not been reported previously to our knowledge. APOB shows differences that may indicate altered sources of genetic disease prevalence in some ancestry groups. In the African ancestry subgroup, the PALB2 and PKP2 findings were replicated using ClinVar as a source of pathogenic variants instead of the VIP database (Supplementary Table 4). These gene-level differences present targets for future investigations of health disparities.Table 2 Genes having rates of pathogenicity that differ from the European pathogenic rate.

Gene	Ancestry group	Ancestry group path. variants	European group path. variants	p value	
APOB	African	4/22,897 (0.02%)	57/49,668 (0.11%)	0.0001	
PKP2	African	33/22,897 (0.14%)	20/49,668 (0.04%)	0.00001	
PALB2	African	33/22,897 (0.14%)	30/49,668 (0.06%)	0.0014	
Pathogenic rates in each gene were compared to the European rate, with significant deviations in population proportions detected with a Bonferroni-adjusted z-test. All of Us allele frequencies are based on biologically independent samples.

Comparison with GnomAD

In order to understand how these findings compare to other large cohorts, we compared these rates of pathogenic findings to gnomAD. To do so, we first identified pathogenic variation using the VIP database and ClinVar separately. We then annotated these variants with the allele frequencies from the All of Us cohort and gnomAD, and summed up the variants in genes to provide gene-level summaries. Under the hypothesis that gnomAD may contain affected individuals, we also made selected comparisons to either the gnomAD non-cancer subset for cancer genes or the non-TopMed subset for genes related to cardiovascular disease. The results (Fig. 2) show the relative difference between All of Us and gnomAD positive rate, broken down by gene. Overall, we observe high-level concordance between these data sources (Pearson correlation 0.99 between All of Us and gnomAD positive rates). However, when comparing gene-level pathogenic rates in ancestry groups other than the European ancestry group, some differences are apparent. For example, in the African ancestry group, the incidence of P/LP BRCA2 variants differs between All of Us cohort and gnomAD (0.093% vs 0.161%, Fig. 3a). However, gnomAD is available in different subsets18, and when the non-cancer subset is used, the pathogenic rate is very similar to that in the All of Us cohort (0.093% vs 0.095%). A similar story is found with the LDLR gene in the Admixed American/Latino ancestry group: using a broad comparison between gnomAD and All of Us, the rates of pathogenicity are divergent (Fig. 3b). However, when using the gnomAD non-TopMed subset, the rates of pathogenicity are highly similar, possibly due to dyslipidemia studies in TopMed19. Other examples are not as clear. For example, we note divergent pathogenicity rates in the RYR1 gene between the gnomAD and All of Us African ancestry groups, but could not account for them using a gnomAD subset. Overall, these results show that the rate of pathogenic variant findings in the All of Us are similar to those seen in gnomAD, and the rates become even closer when we account for disease populations within the gnomAD resource.Fig. 2 Relative positive rates for All of Us vs gnomAD.

This figure shows relative frequencies of previously-curated pathogenic or likely pathogenic variants between the All of Us cohort and gnomAD, broken down by gene and ancestry group. Overall, there is a high level of concordance between variant frequencies of pathogenic variants; most genes show very small differences relative to gnomAD. Ancestries are shown as Dark blue for African, Orange for Latino / Admixed American, Gray for East Asian, Yellow for European and light blue for Other.

Fig. 3 Comparisons to gnomAD subsets.

Though there is high level concordance between the rates of pathogenic variants in All of Us and gnomAD, in some cases there are differences specific to a gene and ancestry group. For example, in participants with African ancestry, the rate of pathogenic variants in BRCA2 diverges from gnomAD (a). However, when the non-cancer subgroup of gnomAD is used, the rates are much more similar. A similar situation is seen in the Admixed American / Latino ancestry group with LDLR (b). Using the non-TopMed portion of gnomAD brings the rates much closer.

Comparing with eMERGE III

In further analysis, we examined the rate of pathogenic variants as seen in the eMERGE III project. This program involved eleven clinical sites providing samples to two clinical laboratories for analysis and reporting. The network utilized a custom gene panel of 109 genes, incorporating the American College of Medical Genetics 59 secondary findings list along with clinical site-specific genes. Notably for this comparison, a substantial fraction of the samples from this program were not chosen based on specific diseases. The observations in the All of Us cohort align with expectations drawn from pathogenic variant frequency in the Association of Molecular Pathologists/American College of Medical Genetics secondary findings genes (See Supplementary Table 5, correlation coefficient for the rate of findings is 0.78). Overall, the eMERGE III cohort exhibited a higher incidence of pathogenic findings. This can possibly be attributed to the different genes present in the gene panel and the fact that the set of genes that were returned varied among clinical sites. Moreover, the eMERGE data underwent complete review, unlike the All of Us data, where we are citing known pathogenic variants. There is a remarkably close match for hemochromatosis, which had the same reporting criteria for the homozygous pathogenic allele C282Y. The substantial discrepancy in clotting disorders is likely due to the decision to report the F5 Leiden variant in eMERGE III.

Assessing the potential for participant selection effects

As an additional approach to understand whether All of Us participants with known genetic diseases are self-selecting for participation in the program, which could impact the counts of pathogenic variants impacting Mendelian disease. We selected four rare and four more common variants associated with disease and examined their allele frequencies relative to that in GnomAD (Table 3). To best match the All of Us ancestry groups, the GnomAD Eur group was taken as a combination of Finnish, non-Finnish and Ashkenazi populations. For these variants, we observe a close match between their GnomAD and All of Us allele frequencies. Only the common HFE rs1800562 alleles, in which a 6% difference was found, was significantly different (Bonferoni corrected critical value of 0.00625, n = 82,106 biologically independent participant samples). Based on this analysis, we do not find an appreciable amount of ascertainment bias in this population.Table 3 Comparison of allele frequencies for specific variants with GnomAD.

Variant	GnomAD AF	All of Us AF	p value (Two-tailed z test)	Population	
BRCA2 c.5946delT (p.Ser1982Argfs*22)	0.0364% (26/71,468)	0.0342% (34/99,336)	0.815	Eur	
BRCA2 c.2808_2811delAAAG (p.Lys938Ilefs*7)	0.0014% (2/145,080)	0.0020% (2/99336)	0.703	Eur	
BRCA2 c.8537_8538delAG (p.Ser2846Glufs*2)	0.0021% (3/145,182)	0.0010% (1/99,336)	0.525	Eur	
LDLR c.682 G > T (p.Glu228X)	0.0024% (2/82,108)	0.0050% (5/99,336)	0.375	Eur	
APOL1 G1 p.S342G - rs73885319 (GRCh38:chr22:36265860:A > G)	22.27% (9208/41,338)	22.36% (10,239/45,794)	0.753	Afr	
APOL1 G1 p.I384M - rs60910145 GRCh38:chr22:36265988:T > C/G	22.43% (9047/40,338)	21.94% (10,045/45,794)	0.081	Afr	
HBB rs334	4.34% (1799/41,432)	4.50% (2,59/45,794)	0.269	Afr	
HFE rs1800562	6.05% (4964/82,106)	6.41% (6371/99,336)	0.001	Eur	
To evaluate whether self-selection by participants with known genetic diseases impacts our study, we examined the allele frequencies of eight disease-associated variants (four rare, four common) in comparison with GnomAD. Except for a 6% difference in the common HFE rs1800562 alleles, the frequencies in our study closely match those in GnomAD. GnomAD and All of Us allele frequencies are based on biologically independent samples.

Enriching for specific genetic factors

When broken down by disease area, the results show that the predominant health-related findings will be in breast cancer, familial hypercholesterolemia, dilated cardiomyopathy and hereditary hemochromatosis (Fig. 1b). To further understand how these findings relate to the participant’s available health information, we made use of additional data resources provided in the AoU Researcher Workbench. The All of Us Research Workbench provides participant health information in the forms of electronic health record condition codes as well as survey questions answered by the participants. To begin to understand how this phenotypic information available in the workbench matches with genetic findings, we looked at breast cancer as an example. Samples were selected if participants had either “Malignant neoplasm of female breast” or “Malignant tumor of breast” condition codes (SNOMED Codes 254837009 and 372064008) or if the participant answered “breast cancer” to the question “Has a doctor or health care provider ever told you that you have or had any following cancers?” 8603 participants fell into this cohort, with 1653 of those having whole-genome sequencing data thus far. This represents an enrichment for P/LP variants in breast cancer patients (32/1653, 1.94% vs 414/98,590, 0.42%, p < 0.00001), demonstrating the ability to use All of Us participant level data to select cohorts enriched for specific genetic factors.

Discussion

The diverse cohort collected as part of the All of Us will be a rich resource for advancing precision medicine. Here, we have examined the pathogenic variant rate in the beta release of the All of Us Researcher Workbench data, finding significant variability between groups of participants with differing ancestry. This variability is likely the result of multiple factors, but ascertainment of pathogenic variants in databases is likely to contribute substantially. Future work will show whether variant interpretation of the All of Us diverse cohort will have an impact on this ascertainment bias of pathogenic variants, but future targeted efforts that aim to perform clinical interpretation of non-European participants could also be necessary.

Ancestry-linked differences in variant pathogenicity found in different genes highlight the importance of the All of Us’s diverse cohort. Different ancestry backgrounds likely carry different burdens of risk depending on the frequency of pathogenic haplotypes. Implementing precision medicine will require understanding these varied risk profiles and gathering detailed information on the haplotype structure of the population. Clinical testing is guided to some extent today by ancestry, for example, BRCA2 in the Ashkenazi Jewish population20 and HLA-B testing for SJS/TEN in some Asian populations9,21,22. As we better understand genetic risk burdens in population groupings we can more precisely target genetic testing and aid interpretation.

Another benefit of this data has been to allow the program to project forward to the “return of health-related results” phase of the program. Based on the variants examined in this work, we estimate that 17% of variant interpretations will require manual assessment of literature, which has strong implications for the time required to complete the review of that variant. This in turn has allowed us to better project the resources required to return health-related results.

We primarily used an internal database of pathogenic genomic variants in this work (the ‘VIP’ database) with ClinVar23 used as a confirmatory resource. Although ClinVar is an essential and widely-used resource, its heterogeneity poses a challenge. For example, the well-known HFE NM_000410.4:c.845 G > A variant, which we previously reported clinically for the eMERGE III13 project when in the homozygous state, is associated with hemochromatosis. However, at the time of publication, this variant is listed as “conflicting” in ClinVar and so was excluded under our simple ClinVar filtering scheme (see Methods). Many other variants likely fall into a similar scenario. It would be possible to fine-tune filters for ClinVar variants and reanalyze this data, which may provide a more complete picture of known pathogenic variation, but there are large effects on sensitivity and specificity that arise due to this tuning that would need to be understood. This level of curation was beyond the scope of the current study.

Ancestry estimation approaches that assign a single continental ancestry to an individual’s entire genome face many issues, including failing to accurately portray the make-up of admixed individuals and potentially grouping individuals that may not share recent ancestry. In future work we expect to adopt more nuanced approaches to ancestry prediction. One such approach may be to use ‘local ancestry’, in which ancestry information is tracked at the variant level, with each variant assigned proportionally to one of many known ancestry groups.

It is paramount that the field address ascertainment biases in knowledge of pathogenic variants. Case-control data are a powerful tool towards achieving this, and the field would benefit greatly from creating more diverse cohorts of patients with accompanying variant interpretations and deep phenotyping, and from sharing that data widely. As new variants of unknown significance are identified, functional studies are also beneficial in their interpretation. The advent of high-throughput functional screening techniques is accelerating24 this data collection and could be targeted at variants from underrepresented populations.

An alternative, although unlikely, explanation for the elevated rate of pathogenic variants seen in this study in the European ancestry population, is that the selected genes carry a burden of pathogenic variants that is in fact specific to individuals of European ancestry. The American College of Medical Genetics took an evidence-based approach12 to selecting genes, prioritizing genes which, at the time, had sufficient evidence showing that patient morbidity and mortality could be reduced through genetic testing while limiting the burden on patients and clinical laboratories. Given known disparities in genetics knowledge, this process may have preferentially identified genes whose impacts on individuals with European ancestry is well studied. Remedying this will require future case-control studies that include diverse populations, which expert panels can then review as they decide on future secondary finding lists.

As a baseline for our expectations of the frequency of pathogenic variants we have compared to gnomAD25. At a high-level, we found strong concordance between these cohorts, although there are some outliers when individual gene-level frequencies are examined. Some differences are likely due to the cohorts applying slightly different ancestry definitions. For example, the precise definition of the “Other” group is likely to be different, and the gnomAD counts do not include the Middle Eastern and South Asian ancestry groups. Using local ancestry (i.e., assigning ancestry not at the sample-level but at the haplotype level) would help resolve this issue.

The current study faces a number of limitations. First, our variant knowledge relies on variant interpretations that were primarily carried out for other genomic reporting projects and during the lead-up to the return of health-related results for the program. As we complete the health-related return of results and the accompanying variant reviews, we will also curate new variants, which is likely to benefit the underserved populations. We are also limited by the cohort creation process adhered to by the All of Us research program. This process recruits participants through partnerships with universities, research centers, and community health centers, direct volunteers, and through community engagement. This process will undoubtedly shape the cohort; for example, we know that participants with a family history of disease are more likely to sign up for genetics studies26. At this time, it is not clear what effect this selection process would have on the rates of rare, pathogenic variants. As evidenced by our examination of selected known pathogenic variants and of the high-level comparisons with GnomAD, there does not appear to be a large selection bias. Further study is needed in order to understand whether there are localized or more subtle selection effects at work in this cohort.

A further limitation is in the high-throughput nature of the pLoF analysis. Mirroring the Association of Molecular Pathologists/American College of Medical Genetics PVS1 guidelines, we filtered pLoF variants by their allele frequency, keeping only those that are rare in the GnomAD database. However, this potentially introduces a bias, as GnomAD is overrepresented with European ancestry individuals. There may remain pLoF variants that would be removed from other ancestry groups with more data. Though it is currently a limitation, as diverse datasets become available, these high-throughput estimates will improve. Although we have not been able to detect selection bias in the cohort in this analysis, the cohort likely does contain both enrichments and depletions of pathogenic variants in specific disease genes due to self-selection of the participants. These effects may be local to specific disease areas or specific subpopulations within the cohort. Researchers making use of this resource should be aware of this potential bias. Finally, although the full cohort for the All of Us project is anticipated to be the most diverse genetic resource available, the current release still features a large proportion of participants of European ancestry. Future releases will enable further exploration beyond what is possible in the current release.

The All of Us Researcher Workbench enabled this first assessment of pathogenic variants within the All of Us cohort, and the diversity of that population has allowed us to begin to detect different frequencies in pathogenic variation between several of the ancestry groups. More work in this area will further reveal groups of participants who carry under-studied pathogenic variation and allow us to target precision medicine efforts at those disparities.

Methods

All of Us demonstration Projects

The All of Us research progra recruits participants that have been underrepresented in biomedical research through a network of affiliated HPOs and direct volunteers27. Demonstration projects were designed to describe the cohort, replicate previous findings for validation, and avoid novel discovery in line with the program values to ensure equal access by researchers to the data28.The work described here was proposed by Consortium members, reviewed and overseen by the program’s Science Committee, and confirmed as meeting criteria for non-human subjects research by the All of Us Institutional Review Board. The initial release of data and tools used in this work was published recently28.

All of Us research Hub

This work was performed on data collected by the All of Us Research Program using the All of Us Researcher Workbench, a cloud-based platform where approved researchers can access and analyze All of Us data. The All of Us data currently includes surveys, electronic health records, and physical measurements. The details of the surveys are available in the Survey Explorer found in the Research Hub, a website designed to support researchers3. Each survey includes branching logic and all questions are optional and may be skipped by the participant. PM recorded at enrollment include systolic and diastolic blood pressure, height, weight, heart rate, waist and hip measurement, wheelchair use, and current pregnancy status. EHR data was linked for those consented participants. All three datatypes are mapped to the Observational Medicines Outcomes Partnership (OMOP) common data model v 5.2 maintained by the Observational Health Data Sciences and Informatics collaborative. To protect participant privacy, a series of data transformations were applied. These included data suppression of codes with a high risk of identification such as military status; generalization of categories, including age, sex at birth, gender identity, sexual orientation, and race; and date shifting by a random (less than one year) number of days, implemented consistently across each participant record. Documentation on privacy implementation and creation of the Curated Data Repository is available in the All of Us Registered Tier Data Dictionary4. The Researcher Workbench currently offers tools with a user interface built for selecting groups of participants (Cohort Builder), creating datasets for analysis (Dataset Builder), and Workspaces with Jupyter Notebooks to analyze data. The notebooks enable use of saved datasets and direct query using R and Python 3 programming languages.

Annotation of known pathogenic variants

We used the ‘VIP’ database to annotate known pathogenic or likely pathogenic (P/LP) variants. The VIP database is a collection of variants compiled during clinical reporting activities carried out at the Human Genome Sequencing Center-Clinical Laboratory (Supplementary Data 1). These variants were manually assessed by a team of variant curation experts led by board-certified clinical geneticists for previous Human Genome Sequencing Center-Clinical Laboratory projects, such as the NIH’s eMERGE III program13 and HeartCare14, a local cardiovascular risk assessment project. All the variants were interpreted based on the guidelines provided by the Association for Molecular Pathology/American College of Medical Genetics and Genomics29, as well as the most recent ClinGen recommendations. The VIP database includes 59,405 genomic variants. These variants are classified into different categories such as pathogenic, likely pathogenic, variants of uncertain significance, benign, likely benign, and risk alleles. They are distributed across 8,653 genes (See the supplemental file vip_gene_count.xslx for details). One of the key advantages of the VIP database is the consistency in variant curation, as all the variants have been curated by the Clinical Variant Interpretation team at HGSC-CL using a uniform set of criteria. This contrasts with resources like ClinVar, which may include contradictory assessments of variant pathogenicity.

Samples/dataset

Aggregate data for this study was generated from All of Us participant data (N = 98,590) using the All of Us Researcher Workbench cloud computing platform. We accessed variant data (single nucleotide variants and indels) from whole-genome sequencing in the alpha3 data release provided by the All of Us Data Resource Center. Details regarding All of Us Data Resource Center’s genomic pipelines can be found here (10.1038/s41586-023-06957-x). All of Us variant data were generated using the GRCh38 human reference build and made available on the pre-production version of the All of Us Researcher Workbench using the Hail framework30. All relevant ethical regulations were followed. All participants in this study provided written informed consent. This work was approved by the Institutional Review Board (IRB) of the All of Us Research Program.

For our analyses, we subsetted the whole genome variants Hail matrix table to coding regions for the 73 genes listed in the American College of Medical Genetics Supplementary Findings v3.0 list12. In these 73 genes, we also included regions 2000 bp upstream and 1000 bp downstream of the coding regions to ensure we do not exclude any pathogenic variants outside of the coding regions. Variants were filtered on genotype quality (GQ > 20) and variant pathogenicity (using ‘Path’ or ‘LPath’ values in the Vip_variant_interpretation field from the VIP database). These variants were then grouped by predicted genetic ancestry and genes to build a contingency table, with heterozygous variant counts across most genes. For the autosomal recessive genes MUTYH, ATP7B and KCNQ1, only homozygous variants were counted. In HFE, only the NM_000410.4:c.845 G > A variant in the homozygous state was considered. We annotated All of Us participants with predicted genetic ancestry provided by the All of Us DRC.

For predicting genetic ancestry, the All of Us Data Resource Center extracted variants from the exon regions of all autosomal, basic, protein-coding transcripts in GENCODE v42, for a training dataset consisting of samples from the Human Genome Diversity Project and 1000 Genomes. These were then used to build principle components (PCs) using Hail. These PCs were then used as the features for a random forest classifier. Samples were assigned an ancestry group if the classification probability exceeded 90%. The ancestry super-populations correspond to the ancestry definitions used within gnomAD25, the Human Genome Diversity Project31,32, and 1000 Genomes6. These include African/African American (afr), American Admixed/Latino (amr), East Asian (eas), European (eur), Middle Eastern (mid), South Asian (sas), and Other (oth; not unambiguously clustering with super-population in the principal component analysis)31. Concordance between self-reported race/ethnicity and these ancestry predictions is 0.898. For full information see https://support.researchallofus.org/hc/en-us/article_attachments/14969477805460/All_Of_Us_Q2_2022_Release_Genomic_Quality_Report__1_.pdf.

To detect predicted loss of function (pLoF) variants, we additionally annotated aggregate variant data using Variant Effect Predictor with the LOFTEE plugin25. pLoF variants were filtered, retaining only those with ‘high confidence’. The following variant effects were treated as loss of function: “frameshift_variant”, “stop_gained”, “stop_lost”, “splice_acceptor_variant, “splice_donor_variant”, and “start_lost” when they were seen in the “vep.most_severe_consequence” field. Other Variant Effect Predictor “HIGH” impact effects such as “transcript_ablation”33 were not present in the dataset.

Statistics and reproducibility

The comparison of pathogenic variant counts between predicted genetic ancestry groups, in the All of Us dataset, was done with a Chi-Square test for independence. To meet the Chi-square requirement that all entries have at least a count of 5, the less-represented ancestries (East Asian, Middle Eastern, South Asian and Other) and genes were aggregated into a single column and row in the contingency table, respectively (Supplementary Table 1). To detect genes with outlying rates of P/LP variants, we compared proportions to the European ancestry rate under the null hypothesis that the proportions are equal. We used a Z-test statistic:1 Z=p^1−p^2p^(1−p^)(1n1+1n2)

Where p^ is the overall proportion:2 p^=Y1+Y2n1+n2

Y1 and Y2 are the group (i.e., participants with primarily European ancestry vs participants with primarily African ancestry) variant counts, n1 and n2 are the group totals and p^1 and p^2 are the group proportions. To translate the Z-score to a p value we assumed a two-tailed normal distribution.

ClinVar comparisons used variant_summary.txt from Jan, 22, 2022, downloaded from the ClinVar downloads site34. We preprocessed this file by (1) filtering out GRCh37 data from the downloaded file and using GRCh38 entries only and (2) filtering the variants 2 stars and above where the ClinicalSignificance field is “Likely pathogenic” or “Likely pathogenic, risk factor” or “Pathogenic” or “Pathogenic/Likely pathogenic” or “Pathogenic/Likely pathogenic, risk factor” or “Pathogenic, risk factor” and the ReviewStatus field is “criteria provided, multiple submitters, no conflicts”, and the LastEvaluated date is on or after Jan 1, 2016. We included three and four star pathogenic variants (“reviewed by expert panel” and “practice guideline” respectively) regardless of the last evaluated time.

GnomAD data used gnomAD v2.1.1 liftover data set from the download site35. For ClinVar, we chose entries from 2016 or later and having two or more stars. We joined the gnomAD and ClinVar dataset by using a variant’s chromosome-position-ref-alt combination as a primary key. We then calculated the pathogenic and likely pathogenic (P/LP) ratio in each gene by adding up the alternate allele count in each gene as the numerator and used the maximum of total alleles in each gene as the denominator. We calculated P/LP ratio/frequency in each gene in the dataset, building a contingency table that mirrored the format derived from that from the All of Us variant data.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

Peer Review File

Supplementary Information

Description of Additional Supplementary Files

Supplementary Data 1

Supplementary Data 2

Reporting Summary

Supplementary information

The online version contains supplementary material available at 10.1038/s42003-023-05708-y.

Acknowledgements

The All of Us Research Program is supported by the National Institutes of Health, Office of the Director: Regional Medical Centers: 1 OT2 OD026549; 1 OT2 OD026554; 1 OT2 OD026557; 1 OT2 OD026556; 1 OT2 OD026550; 1 OT2 OD 026552; 1 OT2 OD026553; 1 OT2 OD026548; 1 OT2 OD026551; 1 OT2 OD026555; IAA #: AOD 16037; Federally Qualified Health Centers: HHSN 263201600085U; Data and Research Center: 5 U2C OD023196; Biobank: 1 U24 OD023121; The Participant Center: U24 OD023176; Participant Technology Systems Center: 1 U24 OD023163; Communications and Engagement: 3 OT2 OD023205; 3 OT2 OD023206; and Community Partners: 1 OT2 OD025277; 3 OT2 OD025315; 1 OT2 OD025337; 1 OT2 OD025276. In addition, the All of Us Research Program would not be possible without the partnership of its participants. We thank our colleagues, Kelsey Mayo, Ashley Able, Ashley Green, Andrea Ramirez, and Sokny Lim for providing their support and input throughout the demonstration project lifecycle. We thank Dr. Jun Qian and Dr. Lina Sulieman for providing input on the project’s code review. We thank Jennifer Zhang and the DRC Genomics Curation Teams for providing the data artifacts used for the project. We would also like to acknowledge the work of the late Professor Deborah A. Nickerson, who was a key member in the early phases of this project. We thank the DRC’s Research Support team for their help during implementation. We also thank the All of Us Science Committee and All of Us Steering Committee for their efforts evaluating and finalizing the approved demonstration projects. The All of Us Research Program would not be possible without the partnership of contributions made by its participants. See below for a roster of past and present All of Us principle investigators. To learn more about the All of Us Research Program’s research data repository, please visit https://www.researchallofus.org/.

Author contributions

Conceived and designed the analysis: E.V., A.M., R.G., G.J., P.E. Collected the data: E.V., D.K., K.P., M.W., Y.C., L.L. Performed the analysis: E.V., K.P., D.K., M.W., Y.C. Wrote the paper: E.V., K.P., D.K., M.W., Y.C., S.K., B.Y., J.K., K.W., J.S., S.M., A.R., A.H., P.E., Q.W., L.L., D.T., G.J., A.M., R.G.

Peer review

Peer review information

Communications Biology thanks Ambroise Wonkam, Hsin-Chou Yang and Karoline Kuchenbaecker for their contribution to the peer review of this work. Primary Handling Editor: George Inglis. A peer review file is available.

Data availability

All sequencing data used in this study are available on the All of Us Researcher Workbench in the v7 release. Researchers can register to access this resource at: https://www.researchallofus.org/. The VIP database of curated variants is available on gitlab: https://gitlab.com/bcm-hgsc/neptune. Source data underlying Figs. 1–3 are provided in Supplementary Data 2.

Code availability

All code used to carry out this project reside in the All of Us Researcher Workbench in the ‘Demo - Assessment of pathogenic variants across the All of Us Research Program’: https://www.researchallofus.org/.

Competing interests

E.V. owns shares in Codified Genomics, a provider of genetic interpretation software. All BCM-affiliated authors declare that Baylor Genetics is a BCM affiliate that derives revenue from genetic testing. All other authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

A list of authors and their affiliations appears at the end of the paper.
==== Refs
References

1. Miga KH Wang T The need for a human pangenome reference sequence Annu. Rev. Genom. Hum. Genet. 2021 22 81 102 10.1146/annurev-genom-120120-081921
2. Sirugo, G., Williams, S. M. & Tishkoff, S. A. The missing diversity in human genetic studies. Cell 177 1080 (2019).
3. Popejoy AB Fullerton SM Genomics is failing on diversity Nature 2016 538 161 164 10.1038/538161a 27734877
4. Carlson CS Diversity is future for genetic analysis Nature 2016 540 341 341 10.1038/540341d 27974770
5. Abul-Husn NS Kenny EE Personalized medicine and the power of electronic health records Cell 2019 177 58 69 10.1016/j.cell.2019.02.039 30901549
6. 1000 Genomes Project Consortium. A global reference for human genetic variation Nature 2015 526 68 74 10.1038/nature15393 26432245
7. Lupski JR Belmont JW Boerwinkle E Gibbs RA Clan genomics and the complex architecture of human disease Cell 2011 147 32 43 10.1016/j.cell.2011.09.008 21962505
8. Stark Z Integrating genomics into healthcare: a global responsibility Am. J. Hum. Genet. 2019 104 13 20 10.1016/j.ajhg.2018.11.014 30609404
9. Manolio TA Global implementation of genomic medicine: we are not alone Sci. Transl. Med. 2015 7 290ps13 10.1126/scitranslmed.aab0194 26041702
10. The “All of Us” research program The “All of Us” research program N. Engl. J. Med. 2019 381 668 676 10.1056/NEJMsr1809937 31412182
11. Venner E Whole-genome sequencing as an investigational device for return of hereditary disease risk and pharmacogenomic results as part of the All of Us Research Program Genome Med. 2022 14 34 10.1186/s13073-022-01031-z 35346344
12. Miller DT ACMG SF v3.0 list for reporting of secondary findings in clinical exome and genome sequencing: a policy statement of the American College of Medical Genetics and Genomics (ACMG) Genet. Med. 2021 23 1381 1390 10.1038/s41436-021-01172-3 34012068
13. eMERGE Consortium. Electronic address: agibbs@bcm.edu & eMERGE Consortium. Harmonizing Clinical Sequencing and Interpretation for the eMERGE III Network Am. J. Hum. Genet. 2019 105 588 605 10.1016/j.ajhg.2019.07.018 31447099
14. Murdock DR Genetic testing in ambulatory cardiology clinics reveals high rate of findings with clinical management implications Genet. Med. 2021 23 2404 2414 10.1038/s41436-021-01294-8 34363016
15. Eric V Neptune: an environment for the delivery of genomic medicine Genet. Med. 2021 23 1838 1846 10.1038/s41436-021-01230-w 34257418
16. Evans MK Longo DL PALB2 mutations and breast-cancer risk N. Engl. J. Med. 2014 371 566 568 10.1056/NEJMe1405784 25099582
17. Alexander J Kowdley KV HFE-associated hereditary hemochromatosis Genet. Med. 2009 11 307 313 10.1097/GIM.0b013e31819d30f2 19444013
18. Gudmundsson S Variant interpretation using population databases: Lessons from gnomAD Hum. Mutat. 2022 43 1012 1030 10.1002/humu.24309 34859531
19. Natarajan P Deep-coverage whole genome sequences and blood lipids among 16,324 individuals Nat. Commun. 2018 9 3391 10.1038/s41467-018-05747-8 30140000
20. D’Andrea E Which BRCA genetic testing programs are ready for implementation in health care? A systematic review of economic evaluations Genet. Med. 2016 18 1171 1180 10.1038/gim.2016.29 27906166
21. Rattanavipapong W Koopitakkajorn T Praditsitthikorn N Mahasirimongkol S Teerawattananon Y Economic evaluation of HLA-B*15:02 screening for carbamazepine-induced severe adverse drug reactions in Thailand Epilepsia 2013 54 1628 1638 10.1111/epi.12325 23895569
22. Towse A Should NICE’s threshold range for cost per QALY be raised? Yes BMJ 2009 338 b181 b181 10.1136/bmj.b181 19171561
23. Landrum MJ ClinVar: improving access to variant interpretations and supporting evidence Nucleic Acids Res. 2018 46 D1062 D1067 10.1093/nar/gkx1153 29165669
24. Bock, C. et al. High-content CRISPR screening. Nat. Rev. Methods Primers 2, (2022).
25. Karczewski KJ The mutational constraint spectrum quantified from variation in 141,456 humans Nature 2020 581 434 443 10.1038/s41586-020-2308-7 32461654
26. Makhnoon S Garrett LT Burke W Bowen DJ Shirts BH Experiences of patients seeking to participate in variant of uncertain significance reclassification research J. Community Genet. 2019 10 189 196 10.1007/s12687-018-0375-3 30027524
27. All of Us Research Program Investigators. The ‘All of Us’ research program N. Engl. J. Med. 2019 381 668 676 10.1056/NEJMsr1809937 31412182
28. Ramirez AH The Research Program: Data quality, utility, and diversity Patterns (N Y) 2022 3 100570 10.1016/j.patter.2022.100570 36033590
29. Richards S Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology Genet. Med. 2015 17 405 424 10.1038/gim.2015.30 25741868
30. Hail genomics toolkit. https://hail.is/. Accessed 25 July 2022.
31. How the All of Us genomic data are organized. https://aousupporthelp.zendesk.com/hc/en-us/articles/4614687617556-How-the-All-of-Us-Genomic-data-are-organized. Accessed 25 July 2022.
32. Cavalli-Sforza LL The Human Genome Diversity Project: past, present and future Nat. Rev. Genet. 2005 6 333 340 10.1038/nrg1579 15803201
33. Genomic variant consequences. https://useast.ensembl.org/info/genome/variation/prediction/predicted_data.html. Accessed 25 July 2022.
34. ClinVar downloads. https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/. Accessed 25 July 2022.
35. gnomAD. https://gnomad.broadinstitute.org/downloads. Accessed 25 July 2022.
