
==== Front
HGG Adv
HGG Adv
Human Genetics and Genomics Advances
2666-2477
Elsevier

S2666-2477(24)00083-6
10.1016/j.xhgg.2024.100343
100343
Report
Mid-pass whole-genome sequencing in a Malagasy cohort uncovers body composition associations
Hamid Iman 14
Raveloson Séverine Nantenaina Stéphie 24
Spiral Germain Jules 2
Ravelonjanahary Soanorolalao 2
Raharivololona Brigitte Marie 2
Randria José Mahenina 3
Zafimaro Mosa 3
Randriambola Tsiorimanitra Aimée 2
Andriantsoa Rota Mamimbahiny 2
Andriamahefa Tojo Julio 2
Rafidison Bodonomena Fitahiana Laza 2
Mughal Mehreen 1
Emde Anne-Katrin 1
Hendershott Melissa 1
LeBaron von Baeyer Sarah 1
Wasik Kaja A. 1
Ranaivoarisoa Jean Freddy 2
Yerges-Armstrong Laura 1
Castel Stephane E. stephane@variantbio.com
1∗
Rakotoarivony Rindra dirranrak@gmail.com
25∗∗
1 Variant Bio, Inc., Seattle, WA 98109, USA
2 University of Antananarivo, Faculty of Sciences, Mention Anthropobiologie et Développement Durable, Antananarivo 101, Madagascar
3 University of Antananarivo, Faculty of Medicine, Ministry of Public Health, Antananarivo 101, Madagascar
∗ Corresponding author stephane@variantbio.com
∗∗ Corresponding author dirranrak@gmail.com
4 These authors contributed equally

5 Lead contact

22 8 2024
10 10 2024
22 8 2024
5 4 10034328 11 2023
19 8 2024
© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Summary

The majority of human genomic research studies have been conducted in European-ancestry cohorts, reducing the likelihood of detecting potentially novel and globally impactful findings. Here, we present mid-pass whole-genome sequencing data and a genome-wide association study in a cohort of 264 self-reported Malagasy individuals from three locations on the island of Madagascar. We describe genetic variation in this Malagasy cohort, providing insight into the shared and unique patterns of genetic variation across the island. We observe phenotypic variation by location and find high rates of hypertension, particularly in the Southern Highlands sampling site, as well as elevated self-reported malaria prevalence in the West Coast site relative to other sites. After filtering to a subset of 214 minimally related individuals, we find a number of genetic associations with body composition traits, including many variants that are only observed in African populations or populations with admixed African ancestry from the 1000 Genomes Project. This study highlights the importance of including diverse populations in genomic research for the potential to gain novel insights, even with small cohort sizes. This project was conducted in partnership and consultation with local stakeholders in Madagascar and serves as an example of genomic research that prioritizes community engagement and potentially impacts our understanding of human health and disease.

We present WGS data from three locales in Madagascar. We describe genetic variation in this cohort. We find population-enriched functional variants and several associations with anthropometric traits, contributing to a broader understanding of human health. This study was conducted in partnership with stakeholders in Madagascar, where communities remain underrepresented in genomic research.

Keywords

human genetics
population genetics
genetic diversity
whole genome sequencing
genome-wide association study
body composition
runs of homozygosity
admixture
human health
community engagement
==== Body
pmcMain Text

Humans have inhabited Madagascar for at least 2,000 years.1,2,3 Malagasy communities can primarily trace their genetic ancestry to eastern African Bantu-speaking groups most likely from southeastern Africa and Austronesian-speaking groups most likely from modern-day Borneo, Indonesia.4,5 For simplicity, we refer to these ancestral sources as “African” and “Austronesian,” respectively, throughout this study. Austronesian founders arrived in eastern Madagascar and African founders on the northwest coast of Madagascar.6,7,8,9 These ancestral groups likely settled the island at different times and in waves over the course of 1,000–2,000 years, with subsequent admixture beginning approximately 28 generations in the past. Following a period of population structuring by geography, more recent internal migrations occurred as a result of increased travel across regions within Madagascar.6,7,8,9

There have been few genome-wide studies including Malagasy cohorts from Madagascar. Of these existing studies, most have primarily focused on selective and demographic history, including the timing of settlement, evidence of past bottlenecks, admixture, and positive selection.4,5,8,10,11 The unique genetic architecture of Madagascar is of particular interest from a population genetics perspective. For one, there is a known signal of strong post-admixture selection at the malaria-protective Duffy-null variant on chromosome 1.10,12 An additional study identified selection signals in three different groups from the southwestern and southeastern coasts of Madagascar.4 Beyond these studies, two of which focused on the single locus signal at Duffy-null, other signals of selection in Malagasy cohorts have not been investigated. Further, these past studies either focused on a single or few regions in Madagascar (e.g., the central highlands12 or specific groups in the south4) or else grouped individuals from multiple regions across the island into one population for analysis purposes.10 However, due to the settlement history of the island, patterns of genetic diversity vary across Madagascar, including relative contributions of Austronesian and African sources.8 Accordingly, there may be differences in selective pressures and genomic regions under selection for groups residing in different parts of Madagascar that previous studies have not focused on. Therefore, adding to the depth and diversity of human genetic data available from Madagascar will be beneficial for future selection studies in the region. Moreover, the recent admixture and selective pressures unique to the island may have led novel or globally rare variants with functional impacts to increase in frequency. This provides a distinct opportunity to detect genetic associations either at novel loci or with novel variants at known loci, which can add to our understanding of the underlying biology driving an association.

Of note, the genome-wide studies carried out in Madagascar to date have all relied on genotyping array data. Here, we describe whole-genome sequencing (WGS) data from a Malagasy cohort in Madagascar. We also identify several genetic associations with anthropometric phenotypes, providing representation of a Malagasy cohort in publicly available genome-wide association studies (GWASs). This study expands the diversity of groups included in genomic research, allowing us to gain novel insight into human health both in Madagascar and globally.

This study was conducted in collaboration with scientists at Variant Bio (VB) and the University of Antananarivo Department of Anthropobiology and Sustainable Development (UA). First, representatives of VB and UA conducted an initial community consultation in order to determine study sites and assess the feasibility and interest in genetic research within the communities. We also established a benefit-sharing agreement based on the needs of the communities, with access to clean water and improved school infrastructure being the most common concerns cited across locations. Our group’s recent publication on VB’s approach to community engagement and benefit sharing provides further details on this project in Madagascar as a case study for the community-driven distribution of financial benefits.13

Ultimately, the study included community-based recruitment of 266 volunteers over the age of 18 from three locales in order to gain population-representative samples: Tsianaloka Village, Tsimafana (on the west coast); Ampandrialaza Village, Ampangabe (in the central highlands); and Tsiandatsiana Village, Ankerana (in the southern highlands) (Figure 1A). This study was approved by the ethics committee Comité Malgache d'Ethique Pour les Sciences et les Technologies (CMEST) (permit number: N° 01-20/05/2021/CMEST/AM). Participants were required to be at least 18 years old, not pregnant, and willing and able to consent. Participants signed an informed consent form approved by CMEST.Figure 1 Cohort overview

(A) Map of three sampling sites in Madagascar corresponding to the West Coast (Tsianaloka Village, Tsimafana), Central Highlands (Ampandrialaza Village, Ampangabe), and Southern Highlands (Tsiandatsiana Village, Ankerana).

(B) PCs 2–4 versus PC1 (percentage of variance explained by PCs 1–4, respectively: 38.66%, 13.13%, 12.55%, and 0.05%).

(C) ADMIXTURE plot for K = 3.

(B) and (C) include the Malagasy cohort (West Coast, Central Highlands, Southern Highlands) and relevant reference populations from 1000 Genomes (GWD, YRI, LUH, CHB, GBR) and additional reference samples from IGDP and PGDP. Points in the scatterplot are colored by group legend, and global ancestry bars are colored by proportion of African-related (pink), Asian-related (purple), or European-related (yellow) ancestry.

For simplicity, we refer to these locales throughout this study as “West Coast,” “Central Highlands,” and “Southern Highlands,” respectively; however, we note that the results described here are not necessarily representative of large geographic regions on the island. We collected history of malaria infection in addition to 15 anthropometric and spirometric phenotypes across the three study sites. See the supplemental materials and methods for more details on participant inclusion criteria and phenotyping.

The majority of participants self-reported excellent health, though approximately 20% of participants reported having been to a doctor in the year preceding this study. The sample cohort was relatively young, with an average age across the three sites ranging from 29 to 35 years old (West Coast = 35 years, Southern Highlands = 33 years, Central Highlands = 29.5 years) (Figure 2A; Table S1). Only 14 individuals (∼5%) reported having malaria in the past year, and 20% of participants reported ever having malaria. The proportion of both male and female participants reporting ever having malaria is noticeably higher on the West Coast (male = 56%, female = 41%) compared to the other two study sites (0%–11%) (Table S2). It is possible that this difference in self-reported malaria infection is reflective of an increased presence of malaria parasites along the coast.14Figure 2 Phenotype distributions by site

(A) Distribution of ages stratified by site (color) and sex (x axis).

(B) Distribution of traits values stratified by site and sex for six of the collected phenotypes (FEV1 max [percentage], systolic blood pressure [mmHg], weight [kg], body fat percentage [percentage], hip circumference [cm], waist circumference [cm]).

We measured both systolic blood pressure (SBP) and diastolic blood pressure (DBP) and categorized individuals who had stage 1 or stage 2 hypertension versus normal blood pressure based on American College of Cardiology (ACC) guidelines.15 We find that approximately 61% of participants either self-reported being hypertensive or had elevated blood pressure during the visit. Of the 161 participants in this category, 73 showed SBP greater than or equal to 140 mmHg or DBP greater than or equal to 90 mmHg at the time of measurement. This makes up 27.4% of all participants in this study, with the Southern Highlands showing, on average, higher blood pressure than the other two sites (Figure 2B; Table S1). We note that some participants may have been on medications to reduce blood pressure at the time of measurement. Several previous studies have similarly reported high rates of hypertension (27%–49% depending on the community),16,17,18 and the Malagasy Ministry of Health has reported that 40% of adults in Madagascar are hypertensive according to the criteria of SBP ≥ 140 and DBP ≥ 90.19 In Figure 2B, we show the distributions of six phenotypes stratified by site and sex: maximum forced expiratory volume in 1 s (FEV1 max) measured in percentage, SBP measured in mmHg, weight measured in kg, body fat percentage, and hip and waist circumferences measured in cm. Descriptive statistics for all collected phenotype distributions, stratified by site and sex, are provided in Tables S1 and S2.

WGS, joint genotyping, and imputation were performed for all participants, with 264 passing quality control (QC) at a mean coverage of 4.6× (supplemental materials and methods).20 We refer to this as mid-pass WGS, as we and other groups have shown that sequencing at a mean depth > 4× produces accurate imputed genotype calls, although we acknowledge that there is no standard definition.20,21 Reference genomes were included from the Indonesian Genome Diversity Project (IGDP) (EGA: EGAS00001003054)22 and the Papuan Genome Diversity Project (PGDP) (EGA: EGAS00001005393)23 and 1000 Genomes references from eastern Asia (Han Chinese [CHB]), Europe (British [GBR]), and central, eastern, and western Africa (Yoruba [YRI], Luhya [LWK], Gambian [GWD]).24 Including the Malagasy cohort, the final call set included 1,077 individuals. A summary table of population-level allele frequencies for the ∼29 million alleles segregating in the Malagasy cohort is provided (see data and code availability).

We performed separate quality checks and filtering at the individual and variant levels for GWASs and population genetics analyses, respectively, due to the different assumptions and input requirements of each analysis (supplemental materials and methods). We ran a principal-component analysis (PCA) in Hail (v.0.2) and ADMIXTURE for a subset of 977 individuals passing QC, including 208 of the Malagasy participants.25,26 Individuals within the Malagasy cohort cluster by locale along the African to Asian ancestry cline in PC space (Figure 1B). The clustering by site is likely indicative of different proportions of African and Austronesian ancestry resulting from the complex settlement history of the island. Indeed, we see from ADMIXTURE results at K = 3 that the average proportion of African-related global ancestry is higher on the West Coast (66% African) compared to a higher average proportion of Austronesian-related global ancestry in the Central Highlands (41% African), with the global ancestry proportions in the Southern Highlands being closer to equal contributions from the two ancestral sources (53% African) (Figure 1C). We include ADMIXTURE results from K = 2 to K = 5 in Figure S1, though overall estimates of global African-related ancestry do not differ significantly across values of K. While our study is not an in-depth analysis of the demographic history of Madagascar or the relationships between global references, our PCA and ADMIXTURE results confirm, at a high level, our expectations from previous studies that have more deeply investigated these questions.5,8 Namely, we confirm higher proportions of African-related ancestry along the coast and relatively higher proportions of Austronesian-related ancestry in the inland communities.

We identified genetic variants that are unique or enriched in the Malagasy cohort relative to other reference populations. Specifically, we defined a variant as “enriched” if the minor-allele frequency (MAF) was less than 0.01 in gnomAD v.3,27,28 less than 0.01 in all reference populations included in joint calling and imputation (LUH, GAM, YRI, CHB, GBR, IGDP, PGDP), and more than 0.05 in the Malagasy cohort. We further annotated these enriched variants as having likely functional impacts if they had a Combined Annotation Dependent Depletion (CADD) score greater than 30 or else were annotated by Variant Effect Predictor (VEP) with a consequence of “high” or “moderate” impact according to the Ensembl IMPACT rating (https://ensembl.org/info/genome/variation/prediction/predicted_data.html) (e.g., stop gain, stop lost, frameshift, missense, etc.). Specifically, high-impact variants are defined as those that are “assumed to have a high (disruptive) impact in the protein, probably causing protein truncation or loss of function or triggering nonsense-mediated decay”, and moderate-impact variants are defined as those that are “non-disruptive and might change protein effectiveness.” Out of 44,731,585 total variants passing QC, we identified 20,656 variants that were enriched in the Malagasy cohort (MAF > 0.05) and either absent or low frequency (<0.01) in other global reference populations as described above. Of these, 116 had likely functional impact according to our criteria (Figure 3A; Table S3).Figure 3 Genetic variation overview

(A) Count of variants enriched in the Malagasy cohort relative to global populations, grouped by functional category (VEP consequence, x axis) and Ensembl IMPACT classification (high, moderate, low, modifier).

(B) Sum of the genome contained in runs of homozygosity (ROHs) by Madagascar study site or reference population.

Next, runs of homozygosity (ROHs) were called to identify signatures of recent or older bottlenecks. We detected ROHs for each individual using bcftools/RoH.29 We called ROHs for each individual and calculated the total amount of the genome contained within ROHs. Larger amounts of the genome contained in ROHs, and in particular ROHs spanning long genomic distances, may be indicative of recent population contraction or bottleneck events and/or certain categories of non-random mating.30 In Figure 3B, we see that the overall patterns of ROHs align with expectations. That is, bottlenecked out-of-Africa (OOA) populations show higher sum-total amounts of the genome contained in ROHs, and continental African populations show lower values due to the higher effective population sizes. Among the three study sites in Madagascar, the Central Highlands have higher values of sum total ROHs, reflective of a historical bottleneck in the region. This is consistent with a past study that found evidence of a bottleneck based on patterns of identity by descent (IBD) sharing from SNP array data in a subgroup located in the central highlands region.8,31

We performed GWASs for the 16 phenotypes collected across 12,528,464 million variants passing GWAS QC (supplemental materials and methods), combining all individuals from across the three sampling sites into one cohort due to the low sample size (213–214 individuals, varying by phenotype). We identified 22 variants reaching genome-wide significance (p < 5e−8) for at least one phenotype. Summary statistics for the lead variant associations are shown in Table 1, variant-level summary statistics for all significant associations can be found in Table S4, and an additional summary of all GWASs is provided in Table S5. Full summary statistics are provided for all GWAS (data and code availability).Table 1 Summary statistics for the lead variants from 15 significant associations

Trait	Gene	Lead variant key	rsID	p	Beta	MG	AFR	EAS	EUR	SAS	
FEV1 max	CCL28	chr5_43372087_A_G	rs11950040	4.11e−08	−0.54	0.38	0.58	0.15	0.53	0.38	
FEV1 max	intergenic	chr4_178078017_G_A	rs115631167	2.41e−08	−1.74	0.03	0.03	0.00	0.00	0.00	
Blood pressure: Systolic	GAS6	chr13_113845979_T_C	rs148486428	4.09e−08	2.17	0.02	0.01	0.00	0.00	0.00	
Blood pressure: Systolic	regulatory region (near FSTL5)	chr4_162485593_A_G	rs11936889	7.54e−09	1.49	0.03	0.09	0.00	0.00	0.00	
Weight	intergenic	chr13_89846914_G_A	rs9588779	1.94e−08	1.94	0.01	0.02	0.00	9.94e−04	0.00	
Body fat percentage	RPTOR	chr17_80547159_G_T	rs115215854	1.18e−08	−1.16	0.04	0.03	0.00	0.00	0.00	
Weight	RPTOR	chr17_80547159_G_T	rs115215854	5.33e−09	1.21	0.04	0.03	0.00	0.00	0.00	
Body fat percentage	ENSG00000228876 (near MYCN)	chr2_16256096_C_T	rs113251741	1.58e−08	−1.39	0.02	0.02	0.00	0.00	0.00	
Hip circumference	ENSG00000228876 (near MYCN)	chr2_16256096_C_T	rs113251741	2.72e−08	−1.72	0.02	0.02	0.00	0.00	0.00	
Weight	ENSG00000228876 (near MYCN)	chr2_16256096_C_T	rs113251741	2.38e−09	1.49	0.02	0.02	0.00	0.00	0.00	
Body fat percentage	WASHC5	chr8_125034853_C_A	rs16900241	2.82e−08	−1.35	0.02	0.07	0.00	9.94e−04	0.00	
Weight	WASHC5	chr8_125034853_C_A	rs16900241	1.06e−09	1.50	0.02	0.07	0.00	9.94e−04	0.00	
Waist circumference	DOK5	chr20_54554331_T_C	rs115838297	1.76e−09	−1.62	0.03	0.04	0.00	0.00	0.00	
Waist-to-hip ratio	DOK5	chr20_54554331_T_C	rs115838297	1.24e−08	−1.51	0.03	0.04	0.00	0.00	0.00	
Waist circumference	ICA1	chr7_8213081_CA_C	rs575387455	2.75e−08	−2.44	0.01	0.01	0.00	0.00	0.00	
The p values and beta coefficients for each lead variant and trait associated locus are listed. The lead variant key (chr_pos_ref_alt), rsID, VEP annotated gene (or closest coding gene in parentheses), and allele frequencies are included. MG, Malagasy cohort from this study; AFR, 1000 Genomes African populations; EAS, 1000 Genomes East Asian populations; EUR, 1000 Genomes European populations; SAS, 1000 Genomes South Asian populations.

Our results include many low- to moderate-frequency variants observed only in African-ancestry populations that are associated with body composition traits (Tables 1 and S4). For example, a haplotype overlapping with RPTOR (MIM: 607130) on chromosome 17 was associated with both body fat percentage and weight (Figure 4A; Table 1). The lead variant (rs115215854) is an intron variant in RPTOR. While variants in RPTOR have previously been associated with body mass index (BMI) in biobank-scale studies (UK Biobank, FinnGen), these cohorts had sample sizes on the order of 450,000–800,000 individuals.32,33,34,35 With a fraction of that sample size (n = 214), we detect significant associations at RPTOR. Further, the lead variant for these two body composition associations segregates at low frequencies (∼3%) in African populations and populations with admixed African ancestry from the 1000 Genomes and in the Malagasy cohort and is unobserved in other 1000 Genomes populations around the world (Table 1).Figure 4 GWAS Manhattan plots for two anthropometric traits

p values for all autosomal variants passing GWAS QC on a −log10 scale for (A) body fat percentage and (B) waist circumference. The annotated gene for the top variant, or nearest protein-coding gene in parentheses, is listed for each locus reaching genome-wide significance (p < 5e−8, red dashed line).

Similarly, we identified a significant association with waist circumference and waist-to-hip ratio on chromosome 20, with the lead variant (rs115838297) in the intron of DOK5 (MIM: 608334) (Figure 4B). Variants in DOK5 have been weakly associated with BMI in the UK Biobank with large cohort sizes.32,33,36 Again, we calculated the allele frequencies for the lead variant from this study in 1000 Genomes populations and found that this allele is observed only in African populations or populations with admixed African ancestry and in the Malagasy cohort and segregates at a low frequency (∼4%) (Table 1).

Among the significant associations we detect is a variant in WASHC5 (MIM: 610657) on chromosome 20 that is positively associated with weight and negatively associated with body fat percentage (rs16900241) (Figure 4A). We find no past studies associating variants in WASHC5 with body composition traits. WASHC5 encodes the protein strumpellin, which is most highly expressed in skeletal muscle and also highly expressed in the thyroid.37 We also identified a significant association with the FEV1 max on chromosome 5. The peak variant (rs11950040) is globally common, with an allele frequency of 42.8% in 1000 Genomes, and is a downstream gene variant for CCL28 (MIM: 605240) (Tables 1 and S4). Additional summary statistics, including population-level allele frequencies for lead variants and all significant variant associations, are provided in Tables 1 and S4, respectively.

Despite the small cohort size in this study, GWASs for well-powered quantitative body composition traits like weight and waist circumference led to the discovery of genes associated with these traits or else the discovery of ancestry-enriched variants at known trait-associated loci. We included a minor allele count (MAC) threshold in our QC criteria (see supplemental materials and methods); however, we recognize that this cohort is a small sample size and that a number of our GWAS associations are with variants between 1% and 5% MAF. We provide sequencing and imputation quality metrics for all variants reaching genome-wide significance for any trait association in Table S6. We flag that one association has a single low-frequency (∼1%) variant reaching genome-wide significance in an intergenic region with no nearby coding genes within 1 Mb of the variant (chr13_89846914_G_A; associated with weight), and there is only one non-imputed alternate allele call for this variant, so this association should be interpreted with caution. We also note that multiple of these low-frequency variants were associated with related body composition traits. We checked whether these low-frequency variants are all carried by the same few individuals, which might suggest confounding due to stratification. There were 35 total individuals that were carriers of at least one of the variants reaching genome-wide significance for any of the body composition traits. Out of these 35, only 9 individuals were carriers for more than one of these alleles. Therefore, we did not find evidence that all of the body composition associations are driven by the same individuals. Though we urge caution in interpretation of associations with low-frequency variants, we are encouraged that a number of these associations are in genes that have previously been associated with relevant phenotypes in larger cohorts (e.g., DOK5, RPTOR). Additionally, the lead variant associated with FEV1 max, chr5_43372087_A_G/rs11950040, is also common in European-ancestry cohorts from the 1000 Genomes and has a nominally significant association with peak expiratory flow in a multi-ancestry European cohort.38 Although we combined the cohort across sites due to low sample sizes, the slight genetic differentiation and varying environmental or selective pressures across the sites suggest that deeper sampling and investigations in each region could identify group-specific genetic associations with traits of interest.

This study of genetic and phenotypic variation across three villages in Madagascar demonstrates site-specific trait distributions, including malaria incidence and genetic substructure by locale. These results emphasize the need for more granular genomic studies that take into account region-specific genetic and environmental differences. In particular, these results have implications for studies that have scanned the Malagasy genome for signals of positive selection. Studies have identified the strong selection at Duffy-null in a cohort with participants grouped together either from all across the island of Madagascar10 or at one specific site.12 Instead, a study designed to detect selection signals in each site independently may find that the post-admixture Duffy-null signal is not as widespread as assumed, particularly in places where malaria incidence is low. A site-specific study design may also find unique and novel selection signals at each locale that were masked in a study that grouped all cohorts despite the clear substructuring that we observe in our analysis. The WGS data described in this study from three specific sites are a valuable start for identifying such region-specific signals. As better global reference population datasets become available, future studies can benefit from the increased depth of genomic sampling from the three sites presented in this study to further investigate their demographic and evolutionary history, including investigations of local ancestry and fine-scale substructure.

Malagasy communities have been underrepresented in WGS studies and GWASs to date, but here we show that even with a small cohort, we can gain potentially important insights regarding the genetic basis of clinically relevant traits. This study demonstrates the importance of partnering with local communities in order to expand the diversity of groups who participate in and potentially benefit from genomic research.

Data and code availability

Consistent with the informed consent signed by study participants, academic researchers should apply for access to de-identified individual-level genotype data through the European Genome-Phenome Archive. The accession number for the individual-level genotype data in this paper is EGA: EGAD50000000708. Population-level allele frequencies for all alleles segregating in the Malagasy cohort (i.e., variants with alternate allele frequencies between 0 and 1 in Madagascar) are publicly available at https://public.variantbio.com/MGUA. This summary table includes reference population allele frequencies from 1000 Genomes, gnomAD, IGDP, and PGDP and relevant variant annotations, such as CADD, VEP annotated gene, and predicted functional consequence. Annotated variant-level summary statistics are provided for all variants that were included in GWASs. Full GWAS summary statistics were uploaded to the GWAS Catalog for each of the 16 phenotypes. The accession numbers for the GWAS summary statistics are NHGRI-EBI GWAS Catalog: GCST90297828, GCST90297829, GCST90297830, GCST90297831, GCST90297832, GCST90297833, GCST90297834, GCST90297835, GCST90297836, GCST90297837, GCST90297838, GCST90297839, GCST90297840, GCST90297841, GCST90297842, and GCST90297843. A detailed workflow and scripts for mid-pass WGS variant calling and imputation are available on GitHub: https://github.com/variant-bio/mid-pass.

Acknowledgments

We would like to thank all of the individuals in Madagascar who participated in this study. We acknowledge Anthropobiologie et Développement Durable, Ecole Doctorale Sciences de la Vie et de l'Environnement, Université d'Antananarivo, and all participating communities and in-field contributors. We also thank Leslie Hepner and Noah Collins for their assistance with training, translation, and study execution. Finally, we would like to thank Murray P. Cox, Guy S. Jacobs, and Professor Herawati Sudoyo for providing access to the IGDP and PGDP data for these analyses.

Author contributions

I.H. and S.N.S.R. analyzed data and wrote the manuscript. R.R., G.J.S., S.R., J.F.R., B.M.R., J.M.R., K.A.W., S.L.v.B., L.Y.-A. and S.E.C. designed the study. S.N.S.R., R.R., M.H., G.J.S., S.R., J.F.R., B.M.R, J.M.R, M.Z., T.A.R., R.M.A., T.J.A., B.F.L.R., and S.L.v.B. carried out the study. S.N.S.R., R.R., G.J.S., S.R., B.M.R., J.M.R., M.Z., T.A.R., R.M.A., T.J.A., B.F.L.R., and S.E.C. processed and quality controlled all subject phenotype and meta data. A.-K.E. processed sequencing data, performed imputation, and quality control. M.M. carried out population genetic analyses. R.R., G.J.S., S.R., J.F.R., and S.L.v.B. performed community engagement and consultation. S.E.C. carried out GWAS. All authors reviewed and contributed to writing the manuscript.

Declaration of interests

I.H., M.M., A.-K.E., M.H., S.L.v.B., K.A.W., L.Y.-A., and S.E.C. are employees and option or shareholders of Variant Bio, Inc.; K.A.W. and S.E.C. are co-founders of Variant Bio, Inc., and S.E.C. is a member of its board of directors. L.Y.-A. is a shareholder of GSK.

Web resources

Ensembl Calculated Consequences, https://ensembl.org/info/genome/variation/prediction/predicted_data.html

Hail, https://github.com/hail-is/hail

Mid-pass WGS variant calling and imputation workflow, https://github.com/variant-bio/mid-pass

OMIM, http://www.omim.org

Variant summary statistics, https://public.variantbio.com/MGUA

Supplemental information

Document S1. Figures S1–S4, Tables S2, S4, and S5, and supplemental materials and methods

Table S1. Descriptive statistics for continuous phenotypes collected stratified by site and community

For each quantitative trait we provide the range [min, max], mean, standard deviation, 95% confidence interval, median, quartile 1, quartile 3 [Q1,Q3], and total number of samples with a value for each phenotype.

Table S3. Summary of 116 Malagasy enriched and high-/moderate-impact variants

Includes variant & gene information, and allele frequencies for Malagasy and reference populations.

Table S6. Sequencing and imputation quality metrics for significant variants

Includes metrics calculated in the GWAS cohort (214 individuals) for all variants which reach genome-wide significance for any trait association.

Document S2. Article plus supplemental information

Supplemental information can be found online at https://doi.org/10.1016/j.xhgg.2024.100343.
==== Refs
References

1 Burney D.A. Burney L.P. Godfrey L.R. Jungers W.L. Goodman S.M. Wright H.T. Jull A.J.T. A chronology for late prehistoric Madagascar J. Hum. Evol. 47 2004 25 63 15288523
2 Perez V.R. Godfrey L.R. Nowak-Kemp M. Burney D.A. Ratsimbazafy J. Vasey N. Evidence of early butchery of giant lemurs in Madagascar J. Hum. Evol. 49 2005 722 742 16225904
3 Dewar R.E. Wright H.T. The culture history of Madagascar J. World Prehist. 7 1993 417 466
4 Pierron D. Razafindrazaka H. Pagani L. Ricaut F.X. Antao T. Capredon M. Sambo C. Radimilahy C. Rakotoarisoa J.A. Blench R.M. Genome-wide evidence of Austronesian–Bantu admixture and cultural reversion in a hunter-gatherer group of Madagascar Proc. Natl. Acad. Sci. USA 111 2014 936 941 24395773
5 Heiske M. Alva O. Pereda-Loth V. Van Schalkwyk M. Radimilahy C. Letellier T. Rakotarisoa J.A. Pierron D. Genetic evidence and historical theories of the Asian and African origins of the present Malagasy population Hum. Mol. Genet. 30 2021 R72 R78 33481023
6 Deschamps H. Histoire de Madagascar 1961 Berger-Levrault
7 Horridge A. The Austronesian Conquest of the Sea — Upwind Bellwood P. Fox J.J. Tryon D. The Austronesians 2006 ANU Press 143 160
8 Pierron D. Heiske M. Razafindrazaka H. Rakoto I. Rabetokotany N. Ravololomanga B. Rakotozafy L.M.A. Rakotomalala M.M. Razafiarivony M. Rasoarifetra B. Genomic landscape of human diversity across Madagascar Proc. Natl. Acad. Sci. USA 114 2017 E6498 E6506 28716916
9 Adelaar A. Towards an integrated theory about the Indonesian migrations to Madagascar Peregrine P.N. Peiros I. Feldman M.W. Ancient Human Migrations: A Multidisciplinary Approach 2009 Univeristy of Utah Press 149 172
10 Pierron D. Heiske M. Razafindrazaka H. Pereda-Loth V. Sanchez J. Alva O. Arachiche A. Boland A. Olaso R. Deleuze J.F. Strong selection during the last millennium for African ancestry in the admixed population of Madagascar Nat. Commun. 9 2018 932 29500350
11 Brucato N. Fernandes V. Kusuma P. Černý V. Mulligan C.J. Soares P. Rito T. Besse C. Boland A. Deleuze J.F. Evidence of Austronesian Genetic Lineages in East Africa and South Arabia: Complex Dispersal from Madagascar and Southeast Asia Genome Biol. Evol. 11 2019 748 758 30715341
12 Hodgson J.A. Pickrell J.K. Pearson L.N. Quillen E.E. Prista A. Rocha J. Soodyall H. Shriver M.D. Perry G.H. Natural selection for the Duffy-null allele in the recently admixed people of Madagascar Proc. Biol. Sci. 281 2014 20140930
13 LeBaron von Baeyer S. Crocker R. Rakotoarivony R. Ranaivoarisoa J.F. Spiral G.J. Castel S. Farnum A. Vance H. Collins N. Fox K. Wasik K. Why community consultation matters in genomic research benefit-sharing models Genome Res. 34 2024 1 6 38296591
14 Arambepola R. Keddie S.H. Collins E.L. Twohig K.A. Amratia P. Bertozzi-Villa A. Chestnutt E.G. Harris J. Millar J. Rozier J. Spatiotemporal mapping of malaria prevalence in Madagascar using routine surveillance and health survey data Sci. Rep. 10 2020 18129
15 Whelton P.K. Carey R.M. Aronow W.S. Casey D.E. Jr. Collins K.J. Dennison Himmelfarb C. DePalma S.M. Gidding S. Jamerson K.A. Jones D.W. 2017 ACC/AHA/AAPA/ABC/ACPM/AGS/APhA/ASH/ASPC/NMA/PCNA Guideline for the Prevention, Detection, Evaluation, and Management of High Blood Pressure in Adults: A Report of the American College of Cardiology/American Heart Association Task Force on Clinical Practice Guidelines J. Am. Coll. Cardiol. 71 2018 e127 e248 29146535
16 Manus M.B. Bloomfield G.S. Leonard A.S. Guidera L.N. Samson D.R. Nunn C.L. High prevalence of hypertension in an agricultural village in Madagascar PLoS One 13 2018 e0201616
17 Ratovoson R. Rasetarinera O.R. Andrianantenaina I. Rogier C. Piola P. Pacaud P. Hypertension, a Neglected Disease in Rural and Urban Areas in Moramanga, Madagascar PLoS One 10 2015 e0137408
18 Rabarijaona L.M.P.H. Rakotomalala D.P. Rakotonirina E.-C.J. Rakotoarimanana S. Randrianasolo O. Prévalence et sévérité de l’hypertension artérielle de l’adulte en milieu urbain à Antananarivo Rev. D’Anesthésie-Réanimation Médecine D’Urgence 1 2009 24 27
19 Malagasy Ministry of Health High blood pressure (arterial hypertension) 2021 Institut Medical de Madagascar https://imm-mg.com/en/2021/06/29/hypertension-arterielle/
20 Emde A.-K. Phipps-Green A. Cadzow M. Gallagher C.S. Major T.J. Merriman M.E. Topless R.K. Takei R. Dalbeth N. Murphy R. Mid-pass whole genome sequencing enables biomedical genetic studies of diverse populations BMC Genom. 22 2021 666
21 Martin A.R. Atkinson E.G. Chapman S.B. Stevenson A. Stroud R.E. Abebe T. Akena D. Alemayehu M. Ashaba F.K. Atwoli L. Low-coverage sequencing cost-effectively detects known and novel variation in underrepresented populations Am. J. Hum. Genet. 108 2021 656 668 33770507
22 Jacobs G.S. Hudjashov G. Saag L. Kusuma P. Darusallam C.C. Lawson D.J. Mondal M. Pagani L. Ricaut F.X. Stoneking M. Multiple Deeply Divergent Denisovan Ancestries in Papuans Cell 177 2019 1010 1021.e32 30981557
23 Brucato N. André M. Tsang R. Saag L. Kariwiga J. Sesuki K. Beni T. Pomat W. Muke J. Meyer V. Papua New Guinean Genomes Reveal the Complex Settlement of North Sahul Mol. Biol. Evol. 38 2021 5107 5121 34383935
24 1000 Genomes Project ConsortiumAuton A. Brooks L.D. Durbin R.M. Garrison E.P. Kang H.M. Korbel J.O. Marchini J.L. McCarthy S. McVean G.A. Abecasis G.R. A global reference for human genetic variation Nature 526 2015 68 74 26432245
25 Alexander D.H. Novembre J. Lange K. Fast model-based estimation of ancestry in unrelated individuals Genome Res. 19 2009 1655 1664 19648217
26 Hail Team Hail v0.2 https://github.com/hail-is/hail 2023
27 Chen S. Francioli L.C. Goodrich J.K. Collins R.L. Kanai M. Wang Q. Alföldi J. Watts N.A. Vittal C. Gauthier L.D. A genome-wide mutational constraint map quantified from variation in 76,156 human genomes Preprint at bioRxiv 2022 10.1101/2022.03.20.485034
28 Karczewski K.J. Francioli L.C. Tiao G. Cummings B.B. Alföldi J. Wang Q. Collins R.L. Laricchia K.M. Ganna A. Birnbaum D.P. The mutational constraint spectrum quantified from variation in 141,456 humans Nature 581 2020 434 443 32461654
29 Narasimhan V. Danecek P. Scally A. Xue Y. Tyler-Smith C. Durbin R. BCFtools/RoH: a hidden Markov model approach for detecting autozygosity from next-generation sequencing data Bioinformatics 32 2016 1749 1751 26826718
30 Ceballos F.C. Joshi P.K. Clark D.W. Ramsay M. Wilson J.F. Runs of homozygosity: windows into population history and trait architecture Nat. Rev. Genet. 19 2018 220 234 29335644
31 Browning S.R. Browning B.L. Accurate Non-parametric Estimation of Recent Effective Population Size from Segments of Identity by Descent Am. J. Hum. Genet. 97 2015 404 418 26299365
32 Kichaev G. Bhatia G. Loh P.R. Gazal S. Burch K. Freund M.K. Schoech A. Pasaniuc B. Price A.L. Leveraging Polygenic Functional Enrichment to Improve GWAS Power Am. J. Hum. Genet. 104 2019 65 75 30595370
33 Pulit S.L. Stoneman C. Morris A.P. Wood A.R. Glastonbury C.A. Tyrrell J. Yengo L. Ferreira T. Marouli E. Ji Y. Meta-analysis of genome-wide association studies for body fat distribution in 694 649 individuals of European ancestry Hum. Mol. Genet. 28 2019 166 174 30239722
34 Barton A.R. Sherman M.A. Mukamel R.E. Loh P.-R. Whole-exome imputation within UK Biobank powers rare coding variant association and fine-mapping analyses Nat. Genet. 53 2021 1260 1269 34226706
35 Sakaue S. Kanai M. Tanigawa Y. Karjalainen J. Kurki M. Koshiba S. Narita A. Konuma T. Yamamoto K. Akiyama M. A cross-population atlas of genetic associations for 220 human phenotypes Nat. Genet. 53 2021 1415 1424 34594039
36 Zhu Z. Guo Y. Shi H. Liu C.L. Panganiban R.A. Chung W. O'Connor L.J. Himes B.E. Gazal S. Hasegawa K. Shared genetic and experimental links between obesity-related traits and asthma subtypes in UK Biobank J. Allergy Clin. Immunol. 145 2020 537 549 31669095
37 O’Leary N.A. Wright M.W. Brister J.R. Ciufo S. Haddad D. McVeigh R. Rajput B. Robbertse B. Smith-White B. Ako-Adjei D. Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation Nucleic Acids Res. 44 2016 D733 D745 26553804
38 Shrine N. Guyatt A.L. Erzurumluoglu A.M. Jackson V.E. Hobbs B.D. Melbourne C.A. Batini C. Fawcett K.A. Song K. Sakornsakolpat P. New genetic signals for lung function highlight pathways and chronic obstructive pulmonary disease associations across multiple ancestries Nat. Genet. 51 2019 481 493 30804560
