
==== Front
Heliyon
Heliyon
Heliyon
2405-8440
Elsevier

S2405-8440(24)12054-3
10.1016/j.heliyon.2024.e36023
e36023
Research Article
Genome-wide association study based on clustering by obesity-related variables uncovers a genetic architecture of obesity in the Japanese and the UK populations
Takahashi Ippei a
Ohseto Hisashi a
Ueno Fumihiko ab
Oonuma Tomomi b
Narita Akira b
Obara Taku abc
Ishikuro Mami ab
Murakami Keiko ab
Noda Aoi abc
Hozawa Atsushi ab
Sugawara Junichi ab
Tamiya Gen abd
Kuriyama Shinichi shinichi.kuriyama.e6@tohoku.ac.jp
abe⁎
a Graduate School of Medicine, Tohoku University, Sendai, Japan
b Tohoku Medical Megabank Organization, Tohoku University, Sendai, Japan
c Tohoku University Hospital, Sendai, Japan
d RIKEN Center for Advanced Intelligence Project, Tokyo, Japan
e International Research Institute of Disaster Science, Tohoku University, Sendai, Japan
⁎ Corresponding author. Graduate School of Medicine, Tohoku University, Sendai, Japan. shinichi.kuriyama.e6@tohoku.ac.jp
09 8 2024
30 8 2024
09 8 2024
10 16 e360238 6 2023
6 8 2024
8 8 2024
© 2024 The Authors. Published by Elsevier Ltd.
2024

https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
Whether all obesity-related variants contribute to the onset of obesity or one or a few variants cause obesity in genetically heterogeneous populations remains obscure. Here, we investigated the genetic architecture of obesity by clustering the Japanese and British populations with obesity using obesity-related factors. In Step-1, we conducted a genome-wide association study (GWAS) with body mass index (BMI) as the outcome for eligible participants. In Step-2, we assigned participants with obesity (BMI ≥25 kg/m2) to five clusters based on obesity-related factors. Subsequently, participants from each cluster and those with a BMI <25 kg/m2 were combined. A GWAS was conducted for each cluster.

Several previously identified obesity-related genes were verified in Step-1. Of the genes detected in Step-1, unique obesity-related genes were detected separately for each cluster in Step-2. Our novel findings suggest that a smaller sample size with increased homogeneity may provide insights into the genetic architecture of obesity.

Highlights

• We investigated the genetic architecture of obesity by clustering a population with obesity.

• CGWAS (Cluster-based GWAS), a method combining clustering and GWAS, was implemented.

• In our cGWAS, unique obesity-related genes were separately detected in each cluster.

• Homogenizing obese characteristics may shed light on the missing heritability of obesity.
==== Body
pmc1 Introduction

Obesity is a serious global medical and economic issue that represents a major risk factor for many lifestyle-related diseases, such as diabetes, hyperlipidemia, and hypertension [1,2]. The global proportion of individuals with a body mass index (BMI) ≥ 25 kg/m2 is reportedly 36.9 % and 38.0 % for men and women, respectively [3]. The pathogenesis of obesity is complex and includes regulation of calorie utilization, appetite, and physical activity, including health care availability, socioeconomic status, and underlying genetic and environmental factors [4,5].

Heritability of BMI has been extensively reported. For instance, in twin studies, BMI heritability ranged from 30 % to 90 % [[6], [7], [8]], whereas in genome-wide association studies (GWASs), it was estimated to be 20–30 % [[9], [10], [11], [12]]. Only approximately 6 % has been reported based on genome-wide significant loci [[9], [10], [11]]. Although GWASs using BMI as an outcome have identified over 700 associated loci [[9], [10], [11], [12], [13], [14], [15]], whether they all contribute to obesity development via the same pathway remains obscure. The association of these genetic variants with obesity may be elucidated by a polygenic model, wherein the effects of each variant are weak yet contribute to the onset of obesity [16]. Hence, if the genetic architecture of obesity could be elucidated by a polygenic model, we would expect larger sample sizes to correspond to more identified signals, whereas smaller sample sizes would correspond to fewer signals (Fig. S1). Moreover, within a genetically heterogeneous population of obesity, if a few variants exhibit a relatively strong influence leading to obesity in a portion of the subtypes included therein, then dividing the population with obesity into homogeneous groups can detect unique genes in each population, even with a reduced sample size. However, to the best of our knowledge, no GWASs has been conducted to divide people with obesity into more homogeneous populations. In a previous study, Traylor et al. demonstrated that categorizing patients with a complex disease into more homogeneous subgroups offered more insight into hidden heritability in a simulation study [17]. Thus, clustering algorithms for machine learning could reveal novel and more genetically homogeneous clusters.

Despite these circumstances, no studies have examined the genetic architecture of obesity by dividing obese individuals into clusters using machine learning techniques. Herein, we investigated the genetic architecture of obesity by dividing individuals with obesity into clusters using diverse obesity-related factors and machine learning techniques and conducting GWAS on each cluster (cluster-based GWAS) [18,19].

2 Materials and methods

2.1 Resource availability

The data that support the findings of this study are available from the TMM biobank. However, restrictions apply to the availability of these data, which were accessed under the license of the current study. Hence, they are not publicly available. Data are available from the authors upon reasonable request and with permission from the TMM Biobank. All inquiries regarding access to data should be addressed to TMM Biobank at dist@megabank.tohoku.ac.jp.

2.2 Experimental model and study participant details

2.2.1 Population

This study was conducted in accordance with the guidelines of the Declaration of Helsinki [20]. The protocol was reviewed and approved by the Institutional Review Board of the TMM Organization. The main study used data from cohort studies conducted by the TMM Birth and Three-Generation Cohort Study (BirThree Cohort Study) and the TMM Community-Based Cohort Study (CommCohort Study) [[21], [22], [23]]. The BirThree Cohort and CommCohort Study were conducted in the Miyagi and Iwate Prefectures, Japan. Details of the BirThree Cohort and CommCohort Study have been described previously [22,23]. The BirThree Cohort Study was a birth- and three-generation cohort study. Pregnant women were registered between July 2013 and March 2017 [22]. Additionally, the family of the pregnant women and their partners (biological father) were recruited (including maternal and paternal grandparents, siblings of the fetus, and their relatives) [22]. Among the BirThree Cohort Study participants, mothers (n = 22,493), fathers (n = 8,823), and grandparents (n = 8058) of the fetus were included in this study. The TMM CommCohort study is a community-based prospective cohort study that includes men and women aged >20 years living in Miyagi Prefecture, Northeastern Japan [23]. The Type 1 survey (n = 41,097) was conducted at specific municipal health check-up sites. The Type 2 survey (n = 13,855) was conducted at assessment centers [23].

Participant data were excluded based on the following criteria: withdrawal of consent, failure to return the self-reported questionnaire, BMI <18.5 kg/m2, missing information on the food frequency questionnaire (FFQ), extreme energy intake (energy intake > mean ± 3 SD), and duplicate participation in both the BirThree Cohort and the CommCohort Study Type-1 (the data of participation at an earlier date were included). Data from eligible participants of the BirThree Cohort Study (n = 23,479), CommCohort Study Type-1 (n = 34,187), and CommCohort Study Type-2 (n = 12,485) were combined (n = 70,151). In the sub-analysis, a similar analysis was conducted using UKB data [[24], [25], [26]] to compare the results with those of the main study. The methods for analyzing the UKB data are described in the Supplementary Information.

2.2.2 Genotyping, imputation, and quality control

Cohort participants were genotyped using the Affymetrix Axiom Japonica Array (v2) in 19 batches, with 50 plates per batch. Details pertaining to the genotyping conducted in TMM have been described previously [27]. Following batch genotyping, samples with a call rate <0.95 or samples with unusually high IBD values compared to other samples were excluded. Additionally, variants with Hardy–Weinberg equilibrium P-values <1.00 × 10−5, minor allele frequency <0.01, or missing fraction >0.01 were excluded from each batch. A direct genotype dataset in the PLINK BED format was obtained by merging the genotype datasets of the 19 batches. A total of 21,541 participants with missing direct genotypic data were excluded. Principal component analysis (PCA) was conducted using the pca approx tool in PLINK 2.0 [28] on the direct genotype dataset. An additional 245 participants with >4 SD for principal components 1 or 2 were excluded. Finally, 48,365 participants (BirThree Cohort Study: n = 11,674; CommCohort Study Type-1: n = 27,745; and CommCohort Study Type-2: n = 8,946) were included in the analysis (Fig. 1). A plot of the participants (n = 48,365) according to principal components 1 and 2 using PCA is depicted in Fig. S2.Fig. 1 Flow chart of exclusion criteria in this study

Participant data from each cohort are excluded based on these criteria.

Fig. 1

To prepare an imputed genotype dataset, pre-phasing was conducted using SHAPEIT2 [29], along with the Duo tool [30], which incorporates information on the relatedness between individuals to increase phasing accuracy. The phased genotypes were subsequently imputed with a cross-imputed panel of 3.5KJPNv2 [31] and 1KGP3 [32] using IMPUTE4 [26]. To create the cross-imputation panel for 3.5KJPNv2 [31] and 1KGP3 [32], the merge_ref_panels tool in IMPUTE2 was used [33]. Consequently, we obtained an imputed genotype dataset in the Oxford BGEN format (https://www.chg.ox.ac.uk/∼gav/qctool_v1/#overview). For genotype imputation data, those with minor allele frequencies <0.01 and imputation information scores <0.8 were excluded. Finally, 9,868,333 single nucleotide polymorphisms (SNPs) were included in the GWASs.

2.3 Quantification and statistical analysis

2.3.1 Variables

The following variables related to obesity, which were collected from questionnaires completed by the participants at baseline for each cohort, were used for clustering: age, nutrient intake calculated from the FFQ based on frequency of food intake over the past year (energy, protein, fat, carbohydrate, sodium, potassium, calcium, magnesium, phosphorus, iron, zinc, copper, manganese, retinol equivalents, vitamins [C, D, K, B1, B2, B6, and B12], niacin, folate, pantothenic acid, cholesterol, dietary fiber, lycopene, α-carotene, β-carotene, and β-cryptoxanthin), frequency of leisure time physical activity (slow or fast walking, moderate exercise, and strenuous exercise; the choices included: no activity, more than once per month, 1–3 times per month, 1–2 times per week, 3–4 times per week, and almost every day), time typically spent in physical activity per day (strenuous work, walk, standing, and sitting) according to predetermined options (no time, <30 min, >30 min ≤60 min, >1 h ≤ 3 h, >3 h ≤ 5 h, >5 h ≤ 7 h, >7 h ≥ 9 h, >9 h ≤ 11 h, and >11 h) sleep duration (<5 h, >5 h ≤ 6 h, >6 h ≤ 7 h, >7 h ≤ 8 h, >8 h ≤ 9 h, and >9 h), difference between weight at age 20 years and current weight, smoking (smoked >100 cigarettes since birth; yes or no), alcohol consumption (>1 drink per month, quit, rarely, and unable to drink), psychological distress over the past month (total K6 score [Japanese version]) [34,35], and birth weight (unknown; >1,500 g ≤ 2,000 g; >2,000 g ≤ 2,500 g; >2,500 g ≤ 3,000 g; >3,000 g ≤ 3,500 g; >3,500 g ≤ 4,000 g; and >4,000 g). Additionally, cohort type (BirThree Cohort Study, CommCohort Study Type-1, and CommCohort Study Type-2) was added to the variables for clustering.

The missing variables used for clustering were imputed using the k-nearest neighbor (KNN) algorithm [36]. KNN selects k samples close to the missing values in the feature space and inputs the median of the k samples in the case of continuous variables or the most frequent category among the k samples in the case of categorical variables. KNN was implemented using the ‘VIM’ package in the R software (version 4.1.0) [37]. Based on previous reports [36,38], we set k to 219 as an odd number close to the square root of 48,365 participants.

2.3.2 Body mass index

BMI was computed by dividing the weight (kg) by the squared height (m2) using self-reported height and weight on a questionnaire completed by the participants at baseline for each cohort. A BMI >25 kg/m2 was defined as obesity based on the Western Pacific Region of the World Health Organization criteria for Asians [39].

2.3.3 Cluster analysis

The k-prototype is a clustering algorithm that combines k-means and k-modes and enables clustering using continuous and categorical variables [40]. The k-prototype was implemented using the ‘clustMixType’ package in R [41]. The number of clusters was set to five. Continuous variables were standardized by subtracting the mean of each variable and dividing it by the SD before clustering.

2.3.4 Genome-wide association studies

The GWASs with BMI as a continuous variable were conducted in two steps. Step-1: GWAS was conducted on all the 48,365 participants. Step-2:13,067 of the 48,365 participants were separated into five clusters using the k prototype. Thereafter, participants in each of the five obesity clusters and those with BMI <25 kg/m2 were combined. A GWAS was conducted for each cluster (Fig. 2). To identify associations between autosomal SNPs and BMI, fastGWA with the GCTA software was employed [42]. FastGWA is a linear mixed model using a sparse genetic relationship matrix that is reportedly robust for population stratification and familial relationships [43]. The top 20 principal components computed from the PCA of the direct genotyping dataset, sex, age, and cohort type (BirThree Cohort Study, CommCohort Study Type-1, and CommCohort Study Type-2) were included as covariates. We set the Bonferroni genome-wide significance threshold at P < 8.33 × 10−9 (5.0 × 10−8/6), as six GWASs were conducted for Step-1 and -2. The detected SNPs were annotated using ANNOVAR [44]. Manhattan and quantile-quantile plots were generated using R software.Fig. 2 Details of the cluster based GWAS.

Fig. 2

3 Results

3.1 Clustering

Following the assignment of 13,067 participants with obesity into five clusters, clusters 1–5 contained 628, 3,073, 4,111, 2,468, and 2,787 participants, respectively. Table S1 depicts the characteristics of the participants with obesity in each cluster. The variables were characterized by mean and standard deviation (SD) for continuous variables and number and percentage for categorical variables. The participants in Cluster 1 had the highest energy and nutrient intake and a higher frequency of leisure-time exercise. The participants in Cluster 2 were characterized by a higher proportion of older women and the highest percentage of non-smokers. The participants in Cluster 3 had the lowest energy and nutrient intake. A high proportion of participants did not exercise during leisure time or perform their usual physical activities (strenuous work, walking, or standing). The participants in Cluster 4 had the second-highest energy and nutrient intake. The participants in Cluster 5 were characterized by the largest proportion of men, the youngest age, and the longest time spent standing or sitting, and had the highest proportion of smokers, the highest proportion of alcohol drinkers, and the highest score on the psychological distress cale.

3.2 Gene interpretation

We observed several genes that satisfied the P < 8.33 × 10−9 threshold in Step-1 (Fig. 3 and Table S2). Most genes for which associations were detected in Step-1 were reportedly associated with obesity. Specifically, LINC01741 [[45], [46], [47]] (Chromosome [Chr] 1), CRYZL2P-SEC16B [46,48] (Chr 1), SEC16B [47,47] (Chr 1), TMEM18 [49,50] (Chr 2), BDNF [46,48] (Chr 11), LINC00678 [51,52] (Chr 11), BDNF-AS [46,48] (Chr 11), FTO [9,10,46,53] (Chr 16), MC4R [54,55] (Chr 18), GIPR [14] (Chr 19), and FBXO46 [51] (Chr 19) have been previously associated with BMI. Moreover, KIF18A (Chr 11) has been previously associated with visceral fat [56], PMAIP1 (Chr 18) with serum IgE measurement [57] and monocyte count [58,59], RSPH6A (Chr 19) with high- and low-density lipoprotein cholesterol levels [60], SYMPK (Chr 19) with type 2 diabetes mellitus [61] and total cholesterol levels [51], and FOXA3 (Chr 19) with waist-to-hip ratio adjusted for BMI [53].Fig. 3 Manhattan plot of Step-1 A genome-wide association study (GWAS) using body mass index (BMI) as a continuous variable is conducted on 48,365 participants.

Fig. 3

Based on the GWAS results in Step-2, several variants detected in Step-1 were observed in separate clusters (Fig. 4 and Table S3). Genome-wide associations were not detected in Cluster 1. In Cluster 2, the loci that satisfied this threshold were LINC01741, CRYZL2P-SEC16B (Chr 1; intergenic), CRYZL2P-SEC16B, and SEC16B (Chr 1). In Cluster 3, FTO (Chr 16), PMAIP1, and MC4R (Chr 18; intergenic) loci were identified. In Cluster 4, BDNF (Chr 11), BDNF-AS (Chr 11), BDNF-AS, LINC00678 (Chr 11), and BDNF, KIF18A (Chr 11) loci were identified. Additionally, in Cluster 5, BDNF-AS, LINC00678 (Chr 11), LINC00678 (Chr 11), BDNF-AS (Chr 11), BDNF (Chr 11), and KIF18A (Chr 11) loci were identified (Fig. 4). Quantile-quantile plots corresponding to the GWAS results of the main study are depicted in Fig. S3.Fig. 4 Manhattan plots of Step-2 We have clustered 13,067 of 48,365 individuals with a body mass index (BMI) ≥ 25 kg/m2 using the k-prototype. Subsequently, participants with obesity in each of the five clusters and those with a BMI <25 kg/m2 are combined. A cluster-based genome-wide association study (cGWAS) performed according to the five clusters.

Fig. 4

In the sub-analysis, the UK Biobank (UKB) data were used for comparison with the main study results. In Step-1, we verified the association between the representative obesity-related genes and BMI (Table S4 and Fig. S4). In Step-2, the clustering results for the 32,779 obese participants revealed that Clusters 1–5 comprised 5,874, 6,497, 6,919, 6,733, and 6756 participants, respectively. The characteristics of each cluster are presented in Table S5. In the GWAS results for Step-2, several variants detected in Step-1 were found in separate clusters. This was similar to the results of the Tohoku Medical Megabank Project (TMM) cohort analysis (Table S6 and Fig. S5).

4 Discussion

Herein, we conducted a GWAS on the participants according to their BMI in Step-1. In Step-2, participants with obesity were divided into five clusters based on obesity-related factors. A GWAS was conducted for each cluster. Consequently, several genes identified in previous studies were verified in Step-1. Of the 18 genes detected in Step-1, LINC01741, CRYZL2P-SEC16B, and SEC16B were significantly associated with Cluster 2. FTO, PMAIP1, and MC4R were linked to Cluster 3. BDNF, BDNF-AS, LINC00678, and KIF18A were associated with Clusters 4 and 5. A similar phenomenon was observed in the sub-analysis using UKB data, wherein unique obesity-related genes were detected in each cluster.

It is important to consider how the cluster characteristics relate to the variants identified in each cluster. The GWAS results in Step-2 may be partially elucidated by cluster characteristics. In Cluster 1, significant associations were not detected. This might be due to the low number of participants with BMI >25.0 kg/m2 as this cluster contained the fewest participants with obesity. Hence, the detection power was insufficient.

FTO, PMAIP1, and MC4R (intergenic) variants were associated with BMI in Cluster 3. Variants in the FTO region regulate IRX3 and IRX5 expression [62]. These promote fat accumulation and cause obesity. Moreover, melanocortin-4-receptors (MC4R), transcribed by the MR4C gene, regulate food intake and energy expenditure [63,64]. MC4R in the paraventricular hypothalamus or amygdala controls food intake, whereas its expression elsewhere is responsible for energy expenditure [63]. Therefore, the genetic variants in FTO and MC4R, which have been reported to increase body fat accumulation and reduce energy expenditure, may partially account for the obesity of individuals in Cluster 3 despite low energy intake.

In Clusters 4 and 5, BDNF and BDNF-AS variants were identified. The participants with obesity in Cluster 4 had the second highest energy and nutrient intake, whereas those in Cluster 5 had the highest mean psychological distress score (K6 total score). Brain-derived neurotrophic factor (BDNF), which is transcribed by the BDNF gene, promotes the development and growth of nerve cells and has anti-obesity effects [65,66]. Furthermore, transcription of the BDNF-AS (antisense RNA) gene regulates BDNF expression [67]. Thus, altering BDNF regulation may affect the central nervous system and alter eating behaviors and psychiatric conditions, as observed in this cluster.

In Cluster 2, SEC16B variants were detected. Participants with obesity in this cluster had the highest proportion of women who were older and were nonsmokers. Variants of SEC16B may be associated with obesity via the regulation of dietary lipid absorption and appetite [68,69]. To the best of our knowledge, no previous study has reported direct connections between the characteristics of Cluster 2 participants and SEC16B variants. However, it should be noted that the characteristics of clusters are not always recognizable by humans. Given that clustering algorithms extract latent features by combining numerous variables, the resulting clusters are not necessarily comprehensible, although they are more homogeneous. Thus, it is essential to define the obscure clusters identified using clustering algorithms.

This study has several strengths. The GWAS results had high validity. Most genes detected in this study were previously reported to be associated with BMI. Therefore, the GWAS data were considered appropriate. The TMM and UKB cohorts included diverse obesity-related factors. Using these two cohorts, it was possible to cluster the obese population into more homogeneous groups using a rich set of obesity-related factors. Sub-analysis replicated the phenomenon, wherein unique obesity-related genes were detected in each cluster. This finding supports the hypothesis that obesity comprises an aggregation of heterogeneous subgroups. Our findings suggest that dividing obese populations into homogeneous subpopulations would yield fewer genetic variants that could elucidate obesity in each subgroup. Although several issues remain to be addressed to elucidate the complete genetic architecture of obesity, this study provides important insights into the potential to inform the development of personalized treatments or nutritional support for obesity. More specifically, once clusters are identified, a classifier can be created using the cluster numbers as training data, which can then be applied to classify obesity into subgroups and verify the effectiveness of obesity treatment according to these subgroups.

This study had certain limitations. It is unclear whether the selection of variables, algorithms, or number of clusters is optimal. Herein, several obesity-related factors were selected. However, the existence of unknown obesity-related factors cannot be ruled out. Additionally, the number of clusters in this study was arbitrarily set to five, which should be explored in the future. Obesity was assessed at a temporal point; therefore, misclassification may have occurred. Even those who were not obese at the time of measurement have the potential to develop obesity with age. The BMI was computed using height and weight from self-reported questionnaires in the main study. Previous studies have demonstrated no substantial differences between BMI computed from self-reported or measured height and weight, indicating the usefulness of self-reported data [70]. Therefore, it is unlikely that the use of self-reported height and weight data significantly distorted our results. However, the impact of self-report bias in this study cannot be ruled out, as all data on obesity-related factors, not just weight and height, were collected through self-report questionnaires. Despite our best efforts to select similar variables for obesity-related factors in the TMM and UKB, we were unable to perfectly match obesity-related factors. This may have made it difficult to replicate the results of this study. Clustering the obese population reduces the sample size, leading to lower statistical power. Additionally, we could not assess the heritability of each cluster because of the small sample size. In the future, computing heritability in a population with a larger sample size and demonstrating that heritability is greater post clustering could provide strong evidence to indicate that obesity is an aggregation of heterogeneous genetic populations.

5 Conclusion

Our data suggest that a decreased sample size with increased homogeneity may provide insights into the genetic architecture of obesity.

Ethics declarations

This study was conducted in accordance with the guidelines of the Declaration of Helsinki. The protocol was reviewed and approved by the Institutional Review Board of the Tohoku Medical Megabank Organization (Approval number: 2022-4-089).

Funding statement

The Tohoku Medical Megabank Project Birth and Three-Generation Cohort Study and Community-Based Cohort Study were supported by the 10.13039/100009619 Japan Agency for Medical Research and Development (10.13039/100009619 AMED ) (grant numbers JP20km0105001 and JP21km0105002). This study was also supported by the 10.13039/501100001700 Ministry of Education, Culture, Sports, Science and Technology (10.13039/501100001700 MEXT ), 10.13039/501100001691 KAKENHI , Japan (grant numbers 19H03894 and 22H03346). AMED and MEXT played no role in the design or execution of the study. This research was supported in part by the 10.13039/100009619 Japan Agency for Medical Research and Development (10.13039/100009619 AMED ) under grant number JP21tm0424601.

Data statement

For the TMM Biobank, data are available from the authors upon reasonable request and with the permission of the TMM Biobank. All inquiries regarding access to data should be addressed to TMM Biobank at dist@megabank.tohoku.ac.jp. The data for the UKB are available to the public upon request from the UKB.

CRediT authorship contribution statement

Ippei Takahashi: Writing – review & editing, Writing – original draft, Visualization, Software, Methodology, Formal analysis, Conceptualization. Hisashi Ohseto: Writing – review & editing, Methodology, Conceptualization. Fumihiko Ueno: Writing – review & editing, Conceptualization. Tomomi Oonuma: Writing – review & editing, Conceptualization. Akira Narita: Writing – review & editing, Methodology, Conceptualization. Taku Obara: Writing – review & editing, Project administration, Investigation, Funding acquisition, Data curation. Mami Ishikuro: Writing – review & editing. Keiko Murakami: Writing – review & editing. Aoi Noda: Writing – review & editing. Atsushi Hozawa: Writing – review & editing, Project administration, Investigation, Data curation. Junichi Sugawara: Writing – review & editing. Gen Tamiya: Writing – review & editing, Methodology, Conceptualization. Shinichi Kuriyama: Writing – review & editing, Writing – original draft, Supervision, Project administration, Methodology, Investigation, Funding acquisition, Data curation, Conceptualization.

Declaration of competing interest

The authors declare the following financial interests/personal relationships which may be considered as potential competing interests:Shinichi Kuriyama reports financial support was provided by 10.13039/501100009619 Japan Agency for Medical Research and Development (AMED) and 10.13039/501100001700 Ministry of Education, Culture, Sports, Science and Technology (MEXT) . If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Abbreviations

AS Antisense RNA

BDNF Brain-Derived Neurotrophic Factor

BirThree Cohort Study TMM Birth and Three-Generation Cohort Study

BMI Body Mass Index

cGWAS Cluster-Based GWAS

Chr Chromosome

CommCohort Study TMM Community-Based Cohort Study

FFQ Food Frequency Questionnaire

GRM Genetic relationship matrix

GWAS Genome-Wide Association Study

AMED Japan Agency for Medical Research and Development

KNN k-Nearest Neighbor

MAF Minor allele frequency

MC4R : MelanoCortin-4-Receptors

PCA Principal Component Analysis

SD Standard Deviation

SNPs Single Nucleotide Polymorphisms

TMM Tohoku Medical Megabank Project

UKB UK Biobank

Appendix A Supplementary data

The following are the Supplementary data to this article.Multimedia component 1

Multimedia component 1

Multimedia component 2

Multimedia component 2

figs1 figs1

figs2 figs2

figs3 figs3

figs4 figs4

figs5 figs5

Acknowledgment

The authors would like to thank the participants of the Tohoku Medical Megabank Project Birth and Three-Generation Cohort Study, Community-Based Cohort Study, and UK Biobank. The authors would also like to thank the staff members of the Tohoku Medical Megabank Organization (https://www.megabank.tohoku.ac.jp/english/a210901/) and UK Biobank.

Appendix A Supplementary data to this article can be found online at https://doi.org/10.1016/j.heliyon.2024.e36023.
==== Refs
References

1 Haslam D.W. James W.P. Obesity Lancet 366 2005 1197 1209 10.1016/S0140-6736(05)67483-1 16198769
2 Eckel R.H. Grundy S.M. Zimmet P.Z. The metabolic syndrome Lancet 365 2005 1415 1428 10.1016/S0140-6736(05)66378-7 15836891
3 Ng M. Fleming T. Robinson M. Thomson B. Graetz N. Margono C. Mullany E.C. Biryukov S. Abbafati C. Abera S.F. Global, regional, and national prevalence of overweight and obesity in children and adults during 1980–2013: a systematic analysis for the Global Burden of Disease Study 2013 Lancet 384 2014 766 781 10.1016/S0140-6736(14)60460-8 24880830
4 Lin X. Li H. Obesity: epidemiology, pathophysiology, and therapeutics Front. Endocrinol. 12 2021 706978 10.3389/fendo.2021.706978
5 Lyon H.N. Hirschhorn J.N. Genetics of common forms of obesity: a brief overview Am. J. Clin. Nutr. 82 Suppl 2005 215S 217S 10.1093/ajcn/82.1.215S 16002823
6 Feng R. How much do we know about the heritability of BMI? Am. J. Clin. Nutr. 104 2016 243 244 10.3945/ajcn.116.139451 27413132
7 Elks C.E. den Hoed M. Zhao J.H. Sharp S.J. Wareham N.J. Loos R.J. Ong K.K. Variability in the heritability of body mass index: a systematic review and meta-regression Front. Endocrinol. 3 2012 29 10.3389/fendo.2012.00029
8 Min J. Chiu D.T. Wang Y. Variation in the heritability of body mass index based on diverse twin studies: a systematic review Obes. Rev. 14 2013 871 882 10.1111/obr.12065 23980914
9 Akiyama M. Okada Y. Kanai M. Takahashi A. Momozawa Y. Ikeda M. Iwata N. Ikegawa S. Hirata M. Matsuda K. Genome-wide association study identifies 112 new loci for body mass index in the Japanese population Nat. Genet. 49 2017 1458 1467 10.1038/ng.3951 28892062
10 Locke A.E. Kahali B. Berndt S.I. Justice A.E. Pers T.H. Day F.R. Powell C. Vedantam S. Buchkovich M.L. Yang J. Genetic studies of body mass index yield new insights for obesity biology Nature 518 2015 197 206 10.1038/nature14177 25673413
11 Yengo L. Sidorenko J. Kemper K.E. Zheng Z. Wood A.R. Weedon M.N. Frayling T.M. Hirschhorn J. Yang J. Visscher P.M. GIANT Consortium Meta-analysis of genome-wide association studies for height and body mass index in ∼700000 individuals of European ancestry Hum. Mol. Genet. 27 2018 3641 3649 10.1093/hmg/ddy271 30124842
12 Yang J. Bakshi A. Zhu Z. Hemani G. Vinkhuyzen A.A. Lee S.H. Robinson M.R. Perry J.R. Nolte I.M. van Vliet-Ostaptchouk J.V. Genetic variance estimation with imputed variants finds negligible missing heritability for human height and body mass index Nat. Genet. 47 2015 1114 1120 10.1038/ng.3390 26323059
13 Speliotes E.K. Willer C.J. Berndt S.I. Monda K.L. Thorleifsson G. Jackson A.U. Lango Allen H. Lindgren C.M. Luan J. Mägi R. Association analyses of 249,796 individuals reveal 18 new loci associated with body mass index Nat. Genet. 42 2010 937 948 10.1038/ng.686 20935630
14 Wen W. Zheng W. Okada Y. Takeuchi F. Tabara Y. Hwang J.Y. Dorajoo R. Li H. Tsai F.J. Yang X. Meta-analysis of genome-wide association studies in East Asian-ancestry populations identifies four new loci for body mass index Hum. Mol. Genet. 23 2014 5492 5504 10.1093/hmg/ddu248 24861553
15 Scuteri A. Sanna S. Chen W.M. Uda M. Albai G. Strait J. Najjar S. Nagaraja R. Orrú M. Usala G. Genome-wide association scan shows genetic variants in the FTO gene are associated with obesity-related traits PLoS Genet. 3 2007 e115 10.1371/journal.pgen.0030115
16 Khera A.V. Chaffin M. Wade K.H. Zahid S. Brancale J. Xia R. Distefano M. Senol-Cosar O. Haas M.E. Polygenic prediction of weight and obesity trajectories from birth to adulthood Cell 177 2019 587 596.e9 10.1016/j.cell.2019.03.028 31002795
17 Traylor M. Markus H. Lewis C.M. Homogeneous case subgroups increase power in genetic association studies Eur. J. Hum. Genet. 23 2015 863 869 10.1038/ejhg.2014.194 25271086
18 Ueno F. Onuma T. Takahashi I. Ohseto H. Narita A. Obara T. Ishikuro M. Murakami K. Noda A. Matsuzaki F. Deep embedded clustering by relevant scales and genome-wide association study in autism bioRxiv 2022 10.1101/2022.07.25.500917 Published online July 25.
19 Narita A. Nagai M. Mizuno S. Ogishima S. Tamiya G. Ueki M. Sakurai R. Makino S. Obara T. Ishikuro M. Clustering by phenotype and genome-wide association study in autism Transl. Psychiatry 10 2020 290 10.1038/s41398-020-00951-x 32807774
20 World Medical Association World Medical Association Declaration of Helsinki: ethical principles for medical research involving human subjects JAMA 310 2013 2191 2194 10.1001/jama.2013.281053 24141714
21 Kuriyama S. Yaegashi N. Nagami F. Arai T. Kawaguchi Y. Osumi N. Sakaida M. Suzuki Y. Nakayama K. Hashizume H. The tohoku medical megabank Project: design and mission J. Epidemiol. 26 2016 493 511 10.2188/jea.JE20150268 27374138
22 Kuriyama S. Metoki H. Kikuya M. Obara T. Ishikuro M. Yamanaka C. Nagai M. Matsubara H. Kobayashi T. Cohort profile: tohoku medical megabank Project birth and three-generation cohort study (TMM BirThree cohort study): rationale, progress and perspective Int. J. Epidemiol. 49 2020 18 19m 10.1093/ije/dyz169 31504573
23 Hozawa A. Tanno K. Nakaya N. Nakamura T. Tsuchiya N. Hirata T. Narita A. Kogure M. Nochioka K. Sasaki R. Study profile of the Tohoku medical Megabank community-based cohort study J. Epidemiol. 31 2021 65 76 10.2188/jea.JE20190271 31932529
24 Sudlow C. Gallacher J. Allen N. Beral V. Burton P. Danesh J. Downey P. Elliott P. Green J. Landray M. UK Biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age PLoS Med. 12 2015 e1001779 10.1371/journal.pmed.1001779
25 Bycroft C. Freeman C. Petkova D. Band G. Elliott L.T. Sharp K. Motyer A. Vukcevic D. Delaneau O. Connell J.O. Genome-wide genetic data on ∼500,000 UK Biobank participants bioRxiv 2017 10.1101/166298 Published online July 20
26 Bycroft C. Freeman C. Petkova D. Band G. Elliott L.T. Sharp K. Motyer A. Vukcevic D. Delaneau O. O'Connell J. The UK biobank resource with deep phenotyping and genomic data Nature 562 2018 203 209 10.1038/s41586-018-0579-z 30305743
27 Yamada M. Motoike I.N. Kojima K. Fuse N. Hozawa A. Kuriyama S. Katsuoka F. Tadaka S. Shirota M. Sakurai M. Genetic loci for lung function in Japanese adults with adjustment for exhaled nitric oxide levels as airway inflammation indicator Commun. Biol. 4 2021 1288 10.1038/s42003-021-02813-8 34782693
28 Chang C.C. Chow C.C. Tellier L.C. Vattikuti S. Purcell S.M. Lee J.J. Second-generation PLINK: rising to the challenge of larger and richer datasets GigaScience 4 2015 7 10.1186/s13742-015-0047-8 25722852
29 Delaneau O. Zagury J.F. Marchini J. Improved whole-chromosome phasing for disease and population genetic studies Nat. Methods 10 2013 5 6 10.1038/nmeth.2307 23269371
30 O'Connell J. Gurdasani D. Delaneau O. Pirastu N. Ulivi S. Cocca M. Traglia M. Huang J. Huffman J.E. Rudan I. A general approach for haplotype phasing across the full spectrum of relatedness PLoS Genet. 10 2014 e1004234 10.1371/journal.pgen.1004234
31 Tadaka S. Katsuoka F. Ueki M. Kojima K. Makino S. Saito S. Otsuki A. Gocho C. Sakurai-Yageta M. 3.5KJPNv2: an allele frequency panel of 3552 Japanese individuals including the X chromosome Hum. Genome Var. 6 2019 28 10.1038/s41439-019-0059-5 31240104
32 1000 Genomes Project ConsortiumAuton A. Brooks L.D. Durbin R.M. Garrison E.P. Kang H.M. Korbel J.O. Marchini J.L. McCarthy S. McVean G.A. A global reference for human genetic variation Nature 526 2015 68 74 10.1038/nature15393 26432245
33 Howie B.N. Donnelly P. Marchini J. A flexible and accurate genotype imputation method for the next generation of genome-wide association studies PLoS Genet. 5 2009 e1000529
34 Kessler R.C. Andrews G. Colpe L.J. Hiripi E. Mroczek D.K. Normand S.L. Walters E.E. Zaslavsky A.M. Short screening scales to monitor population prevalences and trends in non-specific psychological distress Psychol. Med. 32 2002 959 976 10.1017/s0033291702006074 12214795
35 Furukawa T.A. Kawakami N. Saitoh M. Ono Y. Nakane Y. Nakamura Y. Tachimori H. Iwata N. Uda H. Nakane H. The performance of the Japanese version of the K6 and K10 in the World mental health survey Japan. Int. J. Methods psychiatr Res. 17 2008 152 158 10.1002/mpr.257
36 Zhang Z. Introduction to machine learning: k-nearest neighbors Ann. Transl. Med. 4 2016 218 10.21037/atm.2016.03.37 27386492
37 Templ M. Kowarik A. Alfons A. de Cilia G. Prantner B. Rannetbauer W. Visualization and imputation of missing values https://cran.r-project.org/web/packages/VIM 2022
38 Lantz B. Machine Learning with R: Expert Techniques for Predictive Modeling to Solve All Your Data Analysis Problems second ed. 2015 Packt Publishing https://www.packtpub.com/en-us/product/machine-learning-with-r-9781784393908
39 World Health Organization Regional office for the western pacific. The asia-pacific perspective: redefining obesity and its treatment 2000 Sydney: Health Communications https://iris.who.int/handle/10665/206936 (Accessed 8 July 2024)
40 Huang Z. Extensions to the k-means algorithm for clustering large data sets with categorical values Data Min. Knowl. Discov. 2 1998 283 304 10.1023/A:1009769707641
41 Package ‘clustMixType’ k-prototypes clustering for mixed variable-type data https://cran.r-project.org/web/packages/clustMixType/clustMixType.pdf 2021
42 Wang Y. Ding X. Tan Z. Ning C. Xing K. Yang T. Pan Y. Sun D. Wang C. Genome-wide association study of piglet uniformity and farrowing interval Front. Genet. 8 2017 194 10.3389/fgene.2017.00194 29234349
43 Jiang L. Zheng Z. Fang H. Yang J. A generalized linear mixed model association tool for biobank-scale data Nat. Genet. 53 2021 1616 1621 10.1038/s41588-021-00954-4 34737426
44 Wang K. Li M. Hakonarson H. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data Nucleic Acids Res. 38 2010 e164 10.1093/nar/gkq603
45 Ng M.C.Y. Graff M. Lu Y. Justice A.E. Mudgal P. Liu C.T. Young K. Yanek L.R. Feitosa M.F. Wojczynski M.K. Discovery and fine-mapping of adiposity loci using high density imputation of genome-wide association studies in individuals of African ancestry: african Ancestry Anthropometry Genetics Consortium PLoS Genet. 13 2017 e1006719 10.1371/journal.pgen.1006719
46 Wojcik G.L. Graff M. Nishimura K.K. Tao R. Haessler J. Gignoux C.R. Highland H.M. Patel Y.M. Sorokin E.P. Avery C.L. Genetic analyses of diverse populations improves discovery for complex traits Nature 570 2019 514 518 10.1038/s41586-019-1310-4 31217584
47 Monda K.L. Chen G.K. Taylor K.C. Palmer C. Edwards T.L. Lange L.A. Ng M.C. Adeyemo A.A. Allison M.A. Bielak L.F. A meta-analysis identifies new loci associated with body mass index in individuals of African ancestry Nat. Genet. 45 2013 690 696 10.1038/ng.2608 23583978
48 Thorleifsson G. Walters G.B. Gudbjartsson D.F. Steinthorsdottir V. Sulem P. Helgadottir A. Styrkarsdottir U. Gretarsdottir S. Thorlacius S. Jonsdottir I. Genome-wide association yields new sequence variants at seven loci that associate with measures of obesity Nat. Genet. 41 2009 18 24 10.1038/ng.274 19079260
49 Pei Y.F. Zhang L. Liu Y. Li J. Shen H. Liu Y.Z. Tian Q. He H. Wu S. Ran S. Meta-analysis of genome-wide association data identifies novel susceptibility loci for obesity Hum. Mol. Genet. 23 2014 820 830 10.1093/hmg/ddt464 24064335
50 Willer C.J. Speliotes E.K. Loos R.J. Li S. Lindgren C.M. Heid I.M. Berndt S.I. Elliott A.L. Jackson A.U. Lamina C. Six new loci associated with body mass index highlight a neuronal influence on body weight regulation Nat. Genet. 41 2009 25 34 10.1038/ng.287 19079261
51 Zhu Z. Guo Y. Shi H. Liu C.L. Panganiban R.A. Chung W. O'Connor L.J. Himes B.E. Gazal S. Hasegawa K. Shared genetic and experimental links between obesity-related traits and asthma subtypes in UK Biobank J. Allergy Clin. Immunol. 145 2020 537 549 10.1016/j.jaci.2019.09.035 31669095
52 Tachmazidou I. Süveges D. Min J.L. Ritchie G.R.S. Steinberg J. Walter K. Iotchkova V. Schwartzentruber J. Huang J. Memari Y. Whole-genome sequencing coupled to imputation discovers genetic signals for anthropometric traits Am. J. Hum. Genet. 100 2017 865 884 10.1016/j.ajhg.2017.04.014 28552196
53 Sakaue S. Kanai M. Tanigawa Y. Karjalainen J. Kurki M. Koshiba S. Narita A. Konuma T. Yamamoto K. Akiyama M. A cross-population atlas of genetic associations for 220 human phenotypes Nat. Genet. 53 2021 1415 1424 10.1038/s41588-021-00931-x 34594039
54 Barton A.R. Sherman M.A. Mukamel R.E. Loh P.R. Whole-exome imputation within UK Biobank powers rare coding variant association and fine-mapping analyses Nat. Genet. 53 2021 1260 1269 10.1038/s41588-021-00892-1 34226706
55 Akbari P. Gilani A. Sosina O. Kosmicki J.A. Khrimian L. Fang Y.Y. Persaud T. Garcia V. Sun D. Li A. Sequencing of 640,000 exomes identifies GPR75 variants associated with protection from obesity Science 373 2021 eabf8683
56 Shin J. Syme C. Wang D. Richer L. Pike G.B. Gaudet D. Paus T. Pausova Z. Novel genetic locus of visceral fat and systemic inflammation J. Clin. Endocrinol. Metab. 104 2019 3735 3742 10.1210/jc.2018-02656 30942860
57 Akenroye A.T. Brunetti T. Romero K. Daya M. Kanchan K. Shankar G. Chavan S. Preethi Boorgula M. Ampleford E.A. Fonseca H.F. Genome-wide association study of asthma, total IgE, and lung function in a cohort of Peruvian children J. Allergy Clin. Immunol. 148 2021 1493 1504 10.1016/j.jaci.2021.02.035 33713768
58 Chen M.H. Raffield L.M. Mousas A. Sakaue S. Huffman J.E. Moscati A. Trivedi B. Jiang T. Akbari P. Vuckovic D. Trans-ethnic and ancestry-specific blood-cell genetics in 746,667 individuals from 5 global populations Cell 182 2020 1198 1213.e14 10.1016/j.cell.2020.06.045 32888493
59 Vuckovic D. Bao E.L. Akbari P. Lareau C.A. Mousas A. Jiang T. Chen M.H. Raffield L.M. Tardaguila M. Huffman J.E. The polygenic and monogenic basis of blood traits and diseases Cell 182 2020 1214 1231.e11 10.1016/j.cell.2020.08.008 32888494
60 Sinnott-Armstrong N. Tanigawa Y. Amar D. Mars N. Benner C. Aguirre M. Venkataraman G.R. Wainberg M. Ollila H.M. Kiiskinen T. Genetics of 35 blood and urine biomarkers in the UK Biobank Nat. Genet. 53 2021 185 194 33462484
61 Mahajan A. Taliun D. Thurner M. Robertson N.R. Torres J.M. Rayner N.W. Payne A.J. Steinthorsdottir V. Scott R.A. Grarup N. Fine-mapping type 2 diabetes loci to single-variant resolution using high-density imputation and islet-specific epigenome maps Nat. Genet. 50 2018 1505 1513 30297969
62 Claussnitzer M. Dankel S.N. Kim K.H. Quon G. Meuleman W. Haugen C. Glunk V. Sousa I.S. Beaudry J.L. Puviindran V. FTO obesity variant circuitry and adipocyte browning in humans N. Engl. J. Med. 373 2015 895 907 10.1056/NEJMoa1502214 26287746
63 Balthasar N. Dalgaard L.T. Lee C.E. Yu J. Funahashi H. Williams T. Ferreira M. Tang V. McGovern R.A. Kenny C.D. Divergence of melanocortin pathways in the control of food intake and energy expenditure Cell 123 2005 493 505 10.1016/j.cell.2005.08.035 16269339
64 Krashes M.J. Lowell B.B. Garfield A.S. Melanocortin-4 receptor–regulated energy homeostasis Nat. Neurosci. 19 2016 206 219 10.1038/nn.4202 26814590
65 Noble E.E. Billington C.J. Kotz C.M. Wang C. The lighter side of BDNF Am. J. Physiol. Regul. Integr. Comp. Physiol. 300 2011 R1053 R1069 10.1152/ajpregu.00776.2010 21346243
66 Pandit M. Behl T. Sachdeva M. Arora S. Role of brain derived neurotropic factor in obesity Obes. Med. 17 2020 100189 10.1016/j.obmed.2020.100189
67 Ghafouri-Fard S. Khoshbakht T. Taheri M. Ghanbari M. A concise review on the role of BDNF-AS in human disorders Biomed. Pharmacother. 142 2021 112051 10.1016/j.biopha.2021.112051
68 Shi R. Lu W. Tian Y. Wang B. Ave L. Intestinal SEC16B modulates obesity by controlling dietary lipid absorption bioRxiv 2021 10.1101/2021.12.07.471468 Published online Dec 7
69 Hotta K. Nakamura M. Nakamura T. Matsuo T. Nakata Y. Kamohara S. Miyatake N. Kotani K. Komatsu R. Itoh N. Association between obesity and polymorphisms in SEC16B, TMEM18, GNPDA2, BDNF, FAIM2 and MC4R in a Japanese population. J Hum. Genet. 54 2009 727 731 10.1038/jhg.2009.106
70 Haakstad L.A.H. Stensrud T. Gjestvang C. Does self-perception equal the truth when judging own body weight and height? Int. J. Environ. Res. Public Health 18 2021 8502 10.3390/ijerph18168502 34444251
