==== Front Transl Psychiatry Transl Psychiatry Translational Psychiatry 2158-3188 Nature Publishing Group UK London 37386009 2531 10.1038/s41398-023-02531-1 Article Classification and deep-learning–based prediction of Alzheimer disease subtypes by using genomic data http://orcid.org/0000-0002-4412-0552 Shigemizu Daichi daichi@ncgg.go.jp 12 Akiyama Shintaro 1 Suganuma Mutsumi 1 Furutani Motoki 13 Yamakawa Akiko 1 Nakano Yukiko 3 http://orcid.org/0000-0003-0028-1570 Ozaki Kouichi 123 Niida Shumpei 4 1 grid.419257.c 0000 0004 1791 9005 Medical Genome Center, Research Institute, National Center for Geriatrics and Gerontology, Obu, Aichi 474-8511 Japan 2 grid.509459.4 0000 0004 0472 0267 RIKEN Center for Integrative Medical Sciences, Yokohama, Kanagawa 230-0045 Japan 3 grid.257022.0 0000 0000 8711 3200 Department of Cardiovascular Medicine, Hiroshima University Graduate School of Biomedical and Health Sciences, Hiroshima, 734-8553 Japan 4 grid.419257.c 0000 0004 1791 9005 Core Facility Administration, Research Institute, National Center for Geriatrics and Gerontology, Obu, Aichi 474-8511 Japan 29 6 2023 29 6 2023 2023 13 2322 3 2023 16 6 2023 19 6 2023 © The Author(s) 2023 https://creativecommons.org/licenses/by/4.0/ Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons license, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons license and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/. Late-onset Alzheimer’s disease (LOAD) is the most common multifactorial neurodegenerative disease among elderly people. LOAD is heterogeneous, and the symptoms vary among patients. Genome-wide association studies (GWAS) have identified genetic risk factors for LOAD but not for LOAD subtypes. Here, we examined the genetic architecture of LOAD based on Japanese GWAS data from 1947 patients and 2192 cognitively normal controls in a discovery cohort and 847 patients and 2298 controls in an independent validation cohort. Two distinct groups of LOAD patients were identified. One was characterized by major risk genes for developing LOAD (APOC1 and APOC1P1) and immune-related genes (RELB and CBLC). The other was characterized by genes associated with kidney disorders (AXDND1, FBP1, and MIR2278). Subsequent analysis of albumin and hemoglobin values from routine blood test results suggested that impaired kidney function could lead to LOAD pathogenesis. We developed a prediction model for LOAD subtypes using a deep neural network, which achieved an accuracy of 0.694 (2870/4137) in the discovery cohort and 0.687 (2162/3145) in the validation cohort. These findings provide new insights into the pathogenic mechanisms of LOAD. Subject terms Medical genetics Genomics https://doi.org/10.13039/100009619 Japan Agency for Medical Research and Development (AMED) JP21de0107002 JP21dk0207052 JP21dk0207045 JP22dk0207060 JP21km0405501 JP18kk0205009 JP18kk0205012 JP21dk0207045 JP22dk0207060 JP21de0107002 JP21dk0207052 Shigemizu Daichi Ozaki Kouichi Niida Shumpei https://doi.org/10.13039/501100001691 MEXT | Japan Society for the Promotion of Science (JSPS) JP21H02470 Shigemizu Daichi Research Funding for Longevity Sciences from the NCGG (21-24), The Hori Science and Arts Foundation, The Chukyo Longevity Medical Research and Promotion Foundation.Research Funding for Longevity Sciences from the NCGG (21-23), Japanese Ministry of Health, Labour, and Welfare for Research on Dementia.Japan Foundation For Aging and Healthissue-copyright-statement© Springer Nature Limited 2023 ==== Body pmcIntroduction Alzheimer’s disease (AD) is the most common cause of dementia among elderly people [1, 2]. The majority of AD cases are sporadic late-onset AD (LOAD), diagnosed in people ≥65 years [3]. The diagnosis of LOAD is characterized by the accumulation of amyloid-beta (Aβ) plaques and tau neurofibrillary tangles in the neurodegenerative brain [4], but it is a heterogeneous disease and the symptoms vary among patients. Previous studies have reported some subtypes of LOAD [5, 6]. Bredesen reported three LOAD subtypes (inflammatory, non-inflammatory, and cortical) after a metabolic profiling analysis of patients with cognitive decline [5]. Byun et al. identified four LOAD subtypes with heterogeneous patterns of regional atrophy from MRI-measured volumes of the hippocampus and cortical regions (both impaired, hippocampal atrophy only, cortical atrophy only, and both spared) [7]. However, these findings were derived from small samples, limiting their interpretation. Genome-wide association studies (GWAS) have been highly successful in identifying genetic risk factors associated with LOAD [8]. The strongest genetic risk factor for LOAD is the apolipoprotein E ε4 (APOE4) polymorphism [9]. Others include bridging integrator 1 (BIN1) [10], clusterin (CLU) [10], and triggering receptors expressed on myeloid cells 2 (TREM2) [11], also identified from LOAD GWAS analyses. Risk prediction models using these genetic risk factors have also been developed for LOAD using various computational approaches [12, 13], including machine learning techniques [14]. With a large sample size, determining genetic risk will likely not only contribute to our understanding of the underlying pathologic mechanisms of LOAD but will also provide new insights into the classification of distinct LOAD subtypes. However, no previous studies have been conducted on genetic risks for the determination of LOAD subtypes. Here, we comprehensively investigated the genetic architecture of LOAD based on Japanese GWAS data from a large number of patients and controls and examined the presence of LOAD subtypes using the energy landscape [15]. Each subject was represented as a binary vector based on the association signals obtained from GWAS. The energy landscape was visualized with disconnectivity graphs. We revealed the presence of two distinct groups of LOAD patients. Subsequent analysis of representative association signals from each group provided evidence that one group was characterized by both major risk genes for developing LOAD and immune-related genes, and the other group was characterized by genes associated with kidney disorders. We further found that LOAD patients had significantly decreased levels of serum albumin and hemoglobin, which are used to assess kidney function. These results suggest that impaired kidney function could be involved in LOAD pathogenesis. Our findings will contribute to a better understanding of the mechanisms driving heterogeneity in LOAD and will provide novel insight into potential subtype-specific therapies as a step toward precision medicine. Materials and methods Ethics statements This study was approved by the ethics committee of the National Center for Geriatrics and Gerontology (NCGG). The design and performance of the current study involving human subjects were clearly described in a research protocol. All participation was voluntary, and all participants completed informed consent in writing before registering with the NCGG Biobank. Clinical samples Genome-wide genotyping data from the blood samples of a total of 7284 subjects used in this study and their associated clinical data were downloaded from the NCGG Biobank database. Of the total, 4139 subjects, composed of 1947 AD subjects and 2192 cognitively normal (CN) subjects, were genotyped using the Affymetrix Japonica Array (called the discovery cohort). The remaining 3145 subjects, composed of 847 AD subjects and 2298 CN subjects, were genotyped using the Infinium Asian Screening Array (called the validation cohort). The AD subjects were diagnosed with probable or possible AD using the criteria of the National Institute on Aging and Alzheimer’s Association workgroups [16, 17]. The CN subjects had subjective cognitive complaints, but normal cognition (score >23) on a neuropsychological assessment with a comprehensive neuropsychological test, the Mini-Mental State Examination. The diagnosis of all subjects was conducted based on medical history, physical examination and diagnostic tests, neurological examination, neuropsychological tests, and brain imaging with magnetic resonance imaging or computerized tomography by experts including neurologists, psychiatrists, geriatricians, or a neurosurgeon, all experts in dementia who are familiar with its diagnostic criteria. All subjects were ≥60 years of age. Quality control in the GWAS Genotype imputation was conducted by using IMPUTE2 [18] with the 3.5k Japanese reference panel developed by the Tohoku Medical Megabank Organization (https://www.megabank.tohoku.ac.jp/english/) for the discovery cohort and the 5.6k in-house reference panel for the validation set. We used imputed variants with an INFO score ≥0.4. Quality control was performed in each dataset separately after imputation using PLINK software [19]. We first applied quality control (QC) filters: (1) sex inconsistencies (--check-sex); (2) inbreeding coefficient (--het 0.1); (3) genotype missingness (--missing 0.05); (4) PI_HAT >0.25, where PI_HAT is a statistic for the proportion of identity by descent (--genome); and (5) exclusion of outliers from the clusters of East Asian populations in a principal component analysis (PCA) that was conducted together with 1000 Genomes Phase 3 data. We next applied QC filters to the genetic markers (single-nucleotide polymorphisms [SNPs] and short insertions and deletions [Indels]): (1) genotyping efficiency or call rate (--geno 0.95), (2) minor allele frequency (--freq 0.001), and (3) Hardy–Weinberg equilibrium (--hwe 0.001). To generate a set of independent SNPs, we further performed linkage disequilibrium–based SNP pruning using the statistical analysis program PLINK, version 1.90b [20] with a window size of 50 SNPs, a step of 5 SNPs, and a pairwise r2 threshold of 0.1 (--indep-pairwise 50 5 0.1). Disconnectivity graph The autosomal variants that passed QC criteria as described above were assessed with a logistic regression model, adjusting for sex and age with PLINK software (--logistic) [19] in the following way:logitPi=β0+β1×agei+β2×sexi+β3×Varianti. The coefficient and p-value of each variant were obtained from the discovery set. Using the variants with p < 0.01 weighted by their coefficients, a PCA analysis was performed. The principal component (PC) scores were binarized (i.e., −1 or +1) by the mean of the eigenvector values. Each subject was represented by a binary vector Vk = (σ1,σ2,⋯,σN) of 2N, the appearance probability P(Vk) of which was applied to the pairwise maximum entropy model (i.e., Boltzmann distribution). Energy values of 2N binary vectors were compared to identify local energy minimums. By using a disconnectivity graph labeling the energy local minimum states, energy landscapes [15] can be visualized using the R packages akima and rgl (Supplementary Fig. 1). RNA-sequencing data analysis All RNA-sequencing (RNA-seq) data were downloaded from the NCGG Biobank database [21]. The quality of the read sequences (fastq files) was assessed by using FastQC (version 0.11.7). The low-quality reads (