
==== Front
PLoS Med
PLoS Med
plos
PLOS Medicine
1549-1277
1549-1676
Public Library of Science San Francisco, CA USA

39213443
10.1371/journal.pmed.1004451
PMEDICINE-D-24-00964
Research Article
Research and Analysis Methods
Mathematical and Statistical Techniques
Statistical Methods
Forecasting
Physical Sciences
Mathematics
Statistics
Statistical Methods
Forecasting
Medicine and Health Sciences
Epidemiology
Medical Risk Factors
Biology and Life Sciences
Anatomy
Musculoskeletal System
Skeleton
Bone
Bone Density
Medicine and Health Sciences
Anatomy
Musculoskeletal System
Skeleton
Bone
Bone Density
Biology and Life Sciences
Anatomy
Biological Tissue
Connective Tissue
Bone
Bone Density
Medicine and Health Sciences
Anatomy
Biological Tissue
Connective Tissue
Bone
Bone Density
Medicine and Health Sciences
Rheumatology
Connective Tissue Diseases
Osteoporosis
Medicine and Health Sciences
Critical Care and Emergency Medicine
Trauma Medicine
Traumatic Injury
Bone Fracture
Biology and Life Sciences
Genetics
Biology and Life Sciences
Genetics
Single Nucleotide Polymorphisms
Biology and Life Sciences
Physiology
Physiological Parameters
Body Weight
Variability in performance of genetic-enhanced DXA-BMD prediction models across diverse ethnic and geographic populations: A risk prediction study
Ethnic and geographical variability in genetic-enhanced DXA-BMD prediction models
https://orcid.org/0000-0001-9360-1304
Liu Yong Data curation Formal analysis Methodology Software Validation Visualization Writing – original draft 1
https://orcid.org/0000-0001-8731-2899
Meng Xiang-He Methodology Writing – review & editing 2
Wu Chong Methodology Supervision Writing – review & editing 3
https://orcid.org/0000-0002-5163-9774
Su Kuan-Jui Data curation Writing – review & editing 4
Liu Anqi Data curation Writing – review & editing 4
Tian Qing Data curation Writing – review & editing 4
Zhao Lan-Juan Data curation Writing – review & editing 4
Qiu Chuan Data curation Writing – review & editing 4
https://orcid.org/0000-0001-6495-408X
Luo Zhe Data curation Writing – review & editing 4
Gonzalez-Ramirez Martha I Methodology Writing – review & editing 4
Shen Hui Data curation Writing – review & editing 4
https://orcid.org/0000-0002-8121-9498
Xiao Hong-Mei Conceptualization Data curation Funding acquisition Supervision Writing – review & editing 1 5 *
https://orcid.org/0000-0002-0387-8818
Deng Hong-Wen Conceptualization Data curation Funding acquisition Project administration Supervision Writing – review & editing 4 *
1 Center for System Biology, Data Sciences, and Reproductive Health, School of Basic Medical Science, Central South University, Changsha, Hunan Province, China
2 Hunan Provincial Key Laboratory of Regional Hereditary Birth Defects Prevention and Control, Changsha Hospital for Maternal & Child Health Care Affiliated to Hunan Normal University, Changsha, Hunan Province, China
3 Department of Biostatistics, The University of Texas MD Anderson Cancer Center, Houston, Texas, United States of America
4 Tulane Center of Biomedical Informatics and Genomics, Deming Department of Medicine, School of Medicine, Tulane University, New Orleans, Louisiana, United States of America
5 Key Laboratory of Biological, Nanotechnology of National Health Commission, Xiangya Hospital, Central South University, Changsha, Hunan Province, China
The authors have declared that no competing interests exist.

* E-mail: hmxiao@csu.edu.cn (H-MX); hdeng2@tulane.edu (H-WD)
30 8 2024
8 2024
21 8 e100445126 3 2024
23 7 2024
© 2024 Liu et al
2024
Liu et al
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Background

Osteoporosis is a major global health issue, weakening bones and increasing fracture risk. Dual-energy X-ray absorptiometry (DXA) is the standard for measuring bone mineral density (BMD) and diagnosing osteoporosis, but its costliness and complexity impede widespread screening adoption. Predictive modeling using genetic and clinical data offers a cost-effective alternative for assessing osteoporosis and fracture risk. This study aims to develop BMD prediction models using data from the UK Biobank (UKBB) and test their performance across different ethnic and geographical populations.

Methods and findings

We developed BMD prediction models for the femoral neck (FNK) and lumbar spine (SPN) using both genetic variants and clinical factors (such as sex, age, height, and weight), within 17,964 British white individuals from UKBB. Models based on regression with least absolute shrinkage and selection operator (LASSO), selected based on the coefficient of determination (R2) from a model selection subset of 5,973 individuals from British white population. These models were tested on 5 UKBB test sets and 12 independent cohorts of diverse ancestries, totaling over 15,000 individuals. Furthermore, we assessed the correlation of predicted BMDs with fragility fractures risk in 10 years in a case-control set of 287,183 European white participants without DXA-BMDs in the UKBB.

With single-nucleotide polymorphism (SNP) inclusion thresholds at 5×10−6 and 5×10−7, the prediction models for FNK-BMD and SPN-BMD achieved the highest R2 of 27.70% with a 95% confidence interval (CI) of [27.56%, 27.84%] and 48.28% (95% CI [48.23%, 48.34%]), respectively. Adding genetic factors improved predictions slightly, explaining an additional 2.3% variation for FNK-BMD and 3% for SPN-BMD over clinical factors alone. Survival analysis revealed that the predicted FNK-BMD and SPN-BMD were significantly associated with fragility fracture risk in the European white population (P < 0.001). The hazard ratios (HRs) of the predicted FNK-BMD and SPN-BMD were 0.83 (95% CI [0.79, 0.88], corresponding to a 1.44% difference in 10-year absolute risk) and 0.72 (95% CI [0.68, 0.76], corresponding to a 1.64% difference in 10-year absolute risk), respectively, indicating that for every increase of one standard deviation in BMD, the fracture risk will decrease by 17% and 28%, respectively. However, the model’s performance declined in other ethnic groups and independent cohorts. The limitations of this study include differences in clinical factors distribution and the use of only SNPs as genetic factors.

Conclusions

In this study, we observed that combining genetic and clinical factors improves BMD prediction compared to clinical factors alone. Adjusting inclusion thresholds for genetic variants (e.g., 5×10−6 or 5×10−7) rather than solely considering genome-wide association study (GWAS)-significant variants can enhance the model’s explanatory power. The study highlights the need for training models on diverse populations to improve predictive performance across various ethnic and geographical groups.

Yong Liu and co-workers compare

Author summary

Why was this study done?

Osteoporosis diagnosis via bone mineral density (BMD) measurements by dual-energy X-ray absorptiometry (DXA) is impractical for large-scale screening, especially in resource-limited areas.

Genomic data offers a cost-effective alternative for predicting disease risk, but current methods often overlook sub-significant variants and clinical factors.

Most existing genomic prediction methods are based on European ancestry, with limited evaluation in other ethnic populations.

What did the researchers do and find?

We developed BMD prediction models for femoral neck (FNK) and lumbar spine (SPN) using a training set of 17,964 individuals from British white ancestry in UK Biobank (UKBB), integrating clinical and genetic factors.

We observed that strong correlations between predicted and true BMDs (R2≈25% for FNK-BMD and R2≈45% for SPN-BMD) and significant associations with fracture risk in European ancestry populations. And we identified the optimal P-value thresholds for FNK-BMD (5×10−6) and SPN-BMD (5×10−7), noting that these thresholds vary by trait and sample size.

By applying the prediction models on 5 UKBB test sets and 12 independent cohorts of diverse ancestries, totaling over 15,000 individuals, we observed that the BMD prediction models performed well in UKBB European populations but less effectively in other ancestry groups and independent cohorts.

What do these findings mean?

We show that genetic factors could improve the performance of DXA-BMD prediction. Within the same population, the predicted BMDs can help prioritize individuals at high risk of fragility fracture for tailored treatments.

Genetic prediction methods need rigorous evaluation before application to different populations, emphasizing the importance of diverse population training.

Study limitations include differences in the distribution of clinical factors such as sex and age between the UKBB data sets and the independent cohorts, as well as the inclusion of only single-nucleotide polymorphisms (SNPs) as genetic factors.

National Key Research and Development Plan of China 2017YFC1001103, 2016YFC1201805 https://orcid.org/0000-0002-8121-9498
Xiao Hong-Mei http://dx.doi.org/10.13039/501100001809 National Natural Science Foundation of China #81471453 https://orcid.org/0000-0002-8121-9498
Xiao Hong-Mei Jiangwang Educational Endowment https://orcid.org/0000-0002-8121-9498
Xiao Hong-Mei H-MX was supported by the National Key Research and Development Plan of China (2017YFC1001103, 2016YFC1201805), National Natural Science Foundation of China (#81471453) (URL: https://www.nsfc.gov.cn/), and Jiangwang Educational Endowment. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. PLOS Publication Stagevor-update-to-uncorrected-proof
Publication Update2024-09-16
Data AvailabilityThe UK Biobank data utilized in this study can be accessed through an application to the UK Biobank Resource at https://www.ukbiobank.ac.uk; Data from the Osteoporotic Fractures in Men Study (MrOS, accession id: phs000373.v1.p1), the Women's Health Initiative Clinical Trial and Observational Study (WHI, accession id: phs000200.v12.p3) and the Cardiovascular Health Study study (CHS, accession id: phs000287.v7.p1) can be requested via the database of Genotypes and Phenotypes (dbGaP) at https://www.ncbi.nlm.nih.gov/gap. The anonymized, lightweight dataset from the Louisiana Osteoporosis Study (LOS), the Kansas City Osteoporosis Study (KCOS), the China Osteoporosis Study (COS)) could obtained at Mendeley Data (https://data.mendeley.com/datasets/p78t84md5h/1). The raw whole genome array or sequencing (WGS) data will also be available on the Research Aging Biobank (https://agingresearchbiobank.nia.nih.gov).
Data Availability

The UK Biobank data utilized in this study can be accessed through an application to the UK Biobank Resource at https://www.ukbiobank.ac.uk; Data from the Osteoporotic Fractures in Men Study (MrOS, accession id: phs000373.v1.p1), the Women's Health Initiative Clinical Trial and Observational Study (WHI, accession id: phs000200.v12.p3) and the Cardiovascular Health Study study (CHS, accession id: phs000287.v7.p1) can be requested via the database of Genotypes and Phenotypes (dbGaP) at https://www.ncbi.nlm.nih.gov/gap. The anonymized, lightweight dataset from the Louisiana Osteoporosis Study (LOS), the Kansas City Osteoporosis Study (KCOS), the China Osteoporosis Study (COS)) could obtained at Mendeley Data (https://data.mendeley.com/datasets/p78t84md5h/1). The raw whole genome array or sequencing (WGS) data will also be available on the Research Aging Biobank (https://agingresearchbiobank.nia.nih.gov).
==== Body
pmcIntroduction

Osteoporosis is a major global health problem that affects bone quality and increases the risk of fragility fractures [1]. These fractures can impair the quality of life and increase the mortality of the affected individuals, especially in the elderly population over 50 years old [2]. As the world population ages rapidly, osteoporosis and its related fractures pose a huge challenge to the health care system [3]. For example, it is estimated that the annual incidence of hip fractures will exceed 500,000 in the United States of America by 2040, and the direct medical costs for each hip fracture surgery will range from $65,000 to $68,000 [4]. Therefore, early risk prediction for osteoporosis is essential for implementing more effective intervention strategies.

The most common method for diagnosing osteoporosis and assessing the fracture risk is bone mineral density (BMD) measurement by dual-energy X-ray absorptiometry (DXA) [1]. BMD has been shown to be a reliable predictor of osteoporosis risk [1,5,6]. However, DXA has some limitations, such as high cost, low availability, and radiation exposure, which make it unsuitable for mass screening, especially in low-resource or underdeveloped regions [7]. Moreover, previous osteoporosis screening programs have revealed that only a small fraction of the screened subjects were at high risk and required early intervention, indicating that a large amount of the screening resources were spent on individuals who did not qualify for intervention [8,9]. Therefore, there is a great need for developing alternative approaches based on clinical or genetic factors to predict BMD or fracture risk, which would be beneficial for the prevention and management of osteoporosis.

BMD is determined by both genetic and environmental factors, such as sex, age, and lifestyle [10,11]. The heritability of BMD variation is estimated to range from 50% to 85%, depending on the measurement site, the age and sex of the individuals, and the population studied [12]. With the advancement of genome detection technology, large-scale genetic studies have made remarkable progress in identifying genetic variants associated with BMD or fracture risk. In the past decade, genome-wide association studies (GWASs) have identified more than 100 loci associated with BMD, but these loci collectively explain only a small proportion (less than 20%) of the total trait heritability [11]. The variants associated with BMD are common in the population but have small effect sizes [13]. The small effect size of each variant limits its predictive value for osteoporosis. Therefore, prediction methods that combine multiple genetic variants, even with small effect sizes, could be useful for osteoporosis prediction, such as polygenic risk score (PRS) [14]. Several studies have developed and validated PRS models for BMD or fracture risk prediction using different sets of single-nucleotide polymorphisms (SNPs) and cohorts. For example, in a previous work, the researchers conducted a PRS model including 21 SNPs in 19 genes and showed that their PRS was moderately associated with nonvertebral fracture risk in Korean postmenopausal women [15]. Another study generated a PRS model based on 62 SNPs associated with femoral neck (FNK) or lumbar spine (SPN) BMD, which could only explain about 2% of the variation in BMD [16]. More recently, a study developed a PRS consisting of 21,717 genetic variants that was strongly correlated with estimated BMD of the heel (eBMD) (could explain about 23.2% of the variation in eBMD) using data from the UK Biobank (UKBB) project [17]. Despite these advancements, the existing PRS models have some limitations. First, they can only reflect a small part of the DXA-BMD variation because the current cohorts lack sufficient sample size. Second, some of them are based on eBMD, which is a surrogate measure of BMD derived from quantitative ultrasound (QUS) of the heel. Although eBMD is moderately correlated with DXA-BMD [18], it cannot reflect the BMD status of specific sites, which is crucial for assessing site-specific fracture risks [19]. Third, nongenetic risk factors are often ignored in the current PRS studies, but, for most common diseases such as osteoporosis, unglamorous but well-established risk factors like obesity, nutrition, and lifestyle may matter more than a person’s genetic background [20]. Last, most of the BMD prediction models were established in European populations, and their predictive performance in other populations is still lacking evaluation [20,21]. Our study aims to address these gaps by developing prediction models that integrate genetic and clinical factors, for DXA-BMD at 2 clinically relevant sites: the femoral neck (FNK) and lumbar spine (SPN). These sites are not only pivotal for osteoporosis diagnosis but also serve as the most significant predictors of fracture risk [22]. By leveraging a larger cohort and incorporating a broader array of genetic variants, our model seeks to provide a more accurate and comprehensive reflection of BMD variation, thereby enhancing the predictive precision for osteoporosis and associated fracture risks. Further, to investigate whether the prediction models developed in the UKBB European population could be applied to other groups, we tested our models in multiple cohorts of diverse ancestry and geographic backgrounds across populations.

Methods

The aim and design of the study

In this study, we aimed to develop a prediction model for DXA-BMD of FNK and SPN using a large cohort from UKBB and test its performance in diverse ethnic and geographical populations. First, we used the DXA assessment data and genome-wide genetic data to build a series of prediction models to predict the DXA-BMD at FNK and SPN separately (FNK-BMD and SPN-BMD) in UKBB, which are important indicators for evaluating osteoporosis and fragility fractures risk. Then, we selected the model with best performance and assessed the predictive performance in 5 UKBB test sets and 12 independent cohorts. We also examined whether the predicted BMDs were associated with fragility fracture risk. The workflow of our study was outlined in Fig 1.

10.1371/journal.pmed.1004451.g001 Fig 1 The workflow of our study.

Study overview. At Phase 1, we built a series of prediction models to predict the DXA-BMD at FNK and SPN separately (FNK-BMD and SPN-BMD) in UKBB Training set. At Phase 2, we selected the model with highest R2 in UKBB Model selection set as the prediction for further analysis. At Phase 3, we evaluated the predictive performance in 3 UKBB testing data sets and 12 independent data sets. At Phase 4, we examined the association between predicted BMDs and fragility fracture risk in UKBB Fragility facture case-control set and the independent data sets. BMD, bone mineral density; DXA, dual-energy X-ray absorptiometry; FNK, femoral neck; SPN, lumbar spine; UKBB, UK Biobank.

Study cohorts

We obtained the data sets for our study from 7 studies: the UKBB Study [23], the Louisiana Osteoporosis Study (LOS) [24], the Kansas City Osteoporosis Study (KCOS), the China Osteoporosis Study (COS), the Osteoporotic Fractures in Men Study (MrOS) [25], the Women’s Health Initiative Clinical Trial and Observational Study (WHI) [26], and the Cardiovascular Health Study (CHS) [27].

The UKBB is a large-scale biomedical database and research resource containing genetic, lifestyle, and health information from more than 500,000 UK participants, enrolled at ages from 40 to 69 [28]. In total, 32,999 independent participants with DXA-BMD measurements in UKBB were used for model training and evaluation in our study. To validate the association between predicted BMDs and fragility fracture risk, we applied our prediction approaches and evaluated the association between the predicted BMDs and incidence of fragility fracture in 287,183 participants (17,490 fragility fracture cases and 269,693 controls) without DXA-BMD in UKBB. The current research project (application number 63047) was approved by UKBB.

The LOS study is an ongoing cross-sectional study with more than 17,000 subjects so far since 2011 for investigating genetic and nongenetic determinants of osteoporosis and other complex diseases/traits. In total, we included 2,863 Caucasians and 2,097 African Americans randomly selected (stratified by sex and race groups) from the whole LOS cohort [24]. We separated the cohort into LOS African Americans (LOS_AFR) set and LOS Caucasians (LOS_CAU) set when evaluating the performance of prediction models.

The KCOS study included 2,271 unrelated subjects of European ancestry from the Kansas City osteoporosis study (KCOS_CAU). The COS study included 1,569 unrelated subjects of East Asian (Chinese Han) ancestry from the China osteoporosis study (COS_EAS). More details about the sample processing, DNA extraction as well as variant calling can be found in previous published studies [29].

The MrOS is an international multicenter longitudinal study of elderly men. The MrOS cohort recruited 5,995 men aged ≥65 years at multiple assessment centers in the USA between 2000 and 2002 [25]. Among them, we included participants with both genotype data and DXA-BMD measurements. Subsequently, we segregated them based on ancestry, creating 4 distinct data sets: 4,587 individuals of Caucasian ancestry (MrOS_CAU), 181 individuals of African American ancestry (MrOS_AFR), 165 individuals of East Asian ancestry (MrOS_EAS), and 111 individuals of Hispanic ancestry (MrOS_HIS).

The WHI study is a long-term national health investigation that has focused on strategies for preventing heart disease, breast and colorectal cancer, as well as osteoporotic fractures in postmenopausal women [26]. In this study, we included 1,064 individuals with both genotype data and DXA-BMD measurements. Subsequently, we segregated them by ancestry: 671 individuals of black or African-American ancestry (WHI_AFR) and 393 individuals of Hispanic/Latino ancestry (WHI_HIS).

The CHS is a prospective investigation aimed at identifying risk factors for the development and progression of coronary heart disease and stroke in individuals aged 65 years and older [27]. We categorized the CHS cohort by ancestry into 2 sets: 432 individuals of white/Caucasian ancestry (CHS_CAU) and 190 individuals of black/African-American ancestry (CHS_AFR).

All samples were approved by the respective institutional ethics review boards, and all participants provided written informed consent.

Data preprocessing

For UKBB cohort, we extracted the participants who underwent the DXA measurement to develop the BMD prediction models. Considering the requirement for a substantial number of samples for the model training and promotion, we have adopted a relatively stringent criterion for the exclusion of outlier samples. Specifically, we exclude samples with extreme BMD values below or above the 0.1th percentile to mitigate potential errors that may arise from handing factors. To ensure the independence of our samples, the kinship between each pair of individuals was inferred using KING software [30]. We randomly retained only 1 individual from the inferred kinship relationships within third-degree or closer, while the others were excluded. For more detailed data preprocessing procedures, please refer to S1 Appendix.

Splitting of UKBB data

To reduce the impact of population heterogeneity, we limited the model training and selection to British white individuals, but we also tested the best-performing models in other ethnic populations. Participants of British white ancestry in UKBB, with DXA-BMD measurements, and genotyping information (N = 29,914) were randomly assigned to the UKBB British Training set (N = 17,964, 60% of British white individuals), the UKBB British Model Selection set (N = 5,973, 20% of British white individuals), or the UKBB British Testing set (N = 5,977, 20% of British white individuals). The rest of the participants with DXA-BMD measurements in UKBB were divided to UKBB other white set (N = 1,960), UKBB African ancestry set (N = 218), UKBB East Asian ancestry set (N = 112), and UKBB other ancestry set (N = 795). In addition, we included 287,183 white individuals (17,490 fragility fracture cases and 269,693 controls) without DXA-BMDs to assess the association between predicted BMDs and fragility fractures (the UKBB Fracture case-control set). The UKBB Fracture case-control set had no overlap with the UKBB British Training, Model Selection, Testing, African, East Asian, or other ancestry sets.

Phenotype measurements and quality control

The BMD measurements were performed at an imaging assessment center for UKBB with a GE-Lunar iDXA instrument. The clinical risk factors were collected at the imaging visit (prior to DXA scan), including age, sex, height, and weight. The lifestyle variables were collected from the touchscreen questionnaire completed at the UKBB assessment center, including smoking, drinking, and exercise. The individuals in the UKBB Fracture case-control set did not attend the DXA measurement, so their variables were obtained at the initial assessment visit. Fragility fracture cases were identified based on the 10th revision of the International Statistical Classification of Diseases and Related Health Problems (ICD10) codes of primary or secondary diagnoses and self-reported codes following the initial assessment. Individuals without any hospital inpatient data were excluded from this study. The rheumatoid arthritis (RA) cases were identified based on ICD10 codes of primary or secondary diagnoses, and the glucocorticoid using information was obtained from the record-level primary care linked data. Individuals lacking confirmed sex or age information were excluded. For other missing clinical factors, we imputed the values using the median of the respective cohort. A full list of ICD10 and self-reported codes used can be found in S1 Table. More details of phenotype measurements are available in S1 Appendix.

Processing of genetic data

Except for LOS, the genetic data from UKBB, KCOS, COS, MrOS, WHI, and CHS that were genotyped using chip technology was first imputed with the Haplotype Reference Consortium (HRC) [31], UK10K [32], or 1000 Genomes (1000G) haplotype reference panels [33]. For UKBB, the imputed genetic data was obtained from the official release, which was with the HRC and UK10K haplotype resource [31,32]. For KCOS, COS, and MrOS, we performed the genetic imputation using the Michigan Imputation Server with the HRC and 1000G panels [32–34]. Since the genetic data in LOS was genotyped by the next-generation whole genome sequencing with an enough read depth of 22×, we did not perform imputation for this cohort. All genetic data was converted to GRCh37 version. After imputation, we applied quality control to the million imputed SNPs with the following criterions: (1) Imputation scores >0.3; (2) Minimal allele frequency >0.001; (3) Missing rate <0.1; (4) P-value of Hardy–Weinberg equilibrium test > 1×10−6.

Individual-level genetic distance (GD) from training data set

In the Model Selection set and Testing sets (both UKBB and independent data sets), we calculated the GD for each individual. As described in the previous literature [35], the GD is defined as the Euclidean distance between a target individual and the center of training data on the principal component (PC) space of training data. Specifically, we first performed linkage disequilibrium (LD) pruning with plink1.9 (—indep-pairwise 1000 50 0.05) (https://www.cog-genomics.org/plink) and excluded the long-range LD regions. Next, we conducted principal component analysis (PCA) with flashpca2 [36] on the UKBB British Training set to obtain the top 20 PCs. Then, we projected the individuals in the Model Selection set and Testing sets onto the PC space of training data by using SNP loadings (with their means and standard deviations) output from flashpca2. In the end, we computed the GD for each individual as the Euclidean distance of their PCs from the center of training data with the equation di=∑j=120(pcij)2, in which pcij represents the jth PC of individual i.

Genome-wide association study (GWAS)

We performed a GWAS for FNK-BMD and SPN-BMD in the UKBB Training set (N = 17,964). After imputation and QC, ~19 million SNPs were included in the GWAS. We used the BOLT-LMM software [37] to perform a linear mixed model to perform the association test for FNK-BMD and SPN-BMD separately, with adjustment for age, square of age, sex, height, weight, the first 20th genetic PCs, assessment center and genotyping array. Although we have utilized the largest available single ancestry DXA-BMD population to train our model, it still only encompasses a magnitude of tens of thousands, which is relatively small compared to the vast number of genomic variations. Additionally, there is a high correlation between linked SNPs. Including all of them in the model could lead to multicollinearity issues, potentially making the model difficult to train and generalize. To reduce the input features of prediction models, we obtained LD-independent associations using PLINK1.9 by clumping SNPs in LD at a squared correlation (r2) > 0.05 and selected the most significant SNP from within each LD window.

Prediction models of BMDs

In the UKBB Training set (N = 17,964), we built 4 prediction models that integrated with clinical and genetic factors to predict FNK-BMD and SPN-BMD, separately. These models were:

Clumping and Thresholding (C+T) based PRS: Initially, we derived the C+T based PRSs from the GWAS for FNK-BMD and SPN-BMD in the UKBB British Training set using PRSice-2 [38,39]. Subsequently, we integrated these calculated PRS with clinical factors to train a regression model for predicting BMDs.

Linear regression (LR): Differing from the C+T based PRS approach, this method does not require predetermined SNP weights from GWAS but learns them directly from the training data [40].

Regression with least absolute shrinkage and selection operator (LASSO) [41]: In contrast to a simple linear regression model, LASSO models incorporate a regularization parameter (λ) for both variable selection and regularization [41].

Convolutional neural network (CNN) [42]: The CNN model, a deep learning architecture widely used dimensionality reduction and feature recognition in image processing tasks, was employed in our study. The network architecture diagram is shown in S1 Fig. Compared to traditional regression models, neural network models with activation functions [the rectified linear unit (ReLU) in this study] are advantageous for capturing nonlinear relationships between SNPs and BMDs [43].

For each type of model, we fitted 5 models that integrated clinical factors and SNPs with p-values smaller than a chosen set of thresholds (5×10−8, 5×10−7, 5×10−6, 5×10−5, and 5×10−4). More details about the models and the parameter selection process are available in S1 Appendix, S2 and S3 Figs. For each type of model and p-value threshold, the model with the highest R2 in UKBB British Model Selection set was then taken forward for testing in the testing sets. In total, we included 18 clinical factors in the prediction models: age, square of age, sex, height, weight, smoking, drinking, exercise, and the first 10th genetic PCs. The clinical risk factors chosen for BMD and fracture risk were determined associated with osteoporosis and osteoporotic fracture [10]. They are also the most commonly used variables in genetic studies on osteoporosis, such as sex, age, body weight, etc. Furthermore, we constructed an LR model that includes only clinical factors to evaluate whether incorporating genetic factors could improve the prediction for BMDs. The diagram that shows the structure of models is presented in S4 Fig. All continuous variables (age, square of age, height, weight, and genetic PCs) were normalized with zero-mean normalization, and categorical variables (sex, smoking, drinking, exercise, and SNPs) were encoded with one-hot encoding. The S2 Table provides detailed information on the input features of the models.

Model training and evaluation

The model training and evaluation was performed separately for FNK-BMD and SPN-BMD. For each type of model, we trained the model in the UKBB British Training set using a series of parameters (see the Prediction models of BMDs section). All weights in the models were initialized using random initialization, and the model parameters were adjusted based on the mean squared error (MSE) [44]. Due to the large total number of variables and sample size, it is not feasible to process all data at once. Therefore, we have adopted the mini-batch gradient descent (MBGD) approach [45], with a batch size set to 32 and an initial learning rate of 10−4, utilizing the Adam optimizer [46] to adjust the learning rates of the parameters. We optimized the models by minimizing the MSE and evaluated the models with the coefficient of determination (R2) and Pearson correlation coefficient (PCC) between the predicted and actual BMD values. The model conduction and evaluation were performed using Python3.9 and TensorFlow2.

Uncertainty and sensitivity analyses

To estimate the bias and confidence intervals (CIs) of model performances, we implemented a bootstrap strategy, randomly resampling the training data set 50 times for independent training and testing. The bootstrap method is particularly robust, providing estimates of the model’s mean performance and variability. Its key advantage lies in its non-reliance on rigid parametric assumptions, offering flexibility in addressing a range of statistical issues, such as nonlinear regression, determination of CIs, and evaluation of bias [47].

Evaluating the association between predicted BMDs and fragility fracture risk

To examine the association between the predicted BMDs and fragility fracture risk, we first compared the mean value of the predicted BMDs the fragility fracture cases and controls using a T test in the UKBB Fracture case-control set, in which the DXA-BMD measurements were not available. Next, we grouped all individuals by different quantile ranges of predicted BMDs: ≤5%, 5%–20%, 20%–40%, 40%–60%, 60%–80%, 80%–95%, and >95%, and quantified the incidence of fragility fractures in each group.

To investigate whether the predicted BMDs were useful in predicting future fracture risk, we performed survival analyses to evaluate the cumulative incidence of fragility fracture in the UKBB Fracture case-control set, censored by 10 years. First, we performed a univariate survival analysis with the Kaplan–Meier model to examine whether there was a difference in cumulative fracture risk among different groups based on different quantile ranges of predicted BMDs. Then, we performed a multivariate survival analysis with the Cox proportional hazards regression (Cox) model, integrating the predicted BMDs and clinical factors: age, sex, height, weight, body mass index (BMI), previous fracture, smoking, drinking, exercise, glucocorticoid using, and RA. We incorporated additional variables such as glucocorticoid use and RA status into the fracture prediction model, as these factors have been proven to be closely associated with osteoporotic fractures [48,49] and are utilized in the commonly used fracture prediction tool, Fracture Risk Assessment Tool (FRAX, https://frax.shef.ac.uk/FRAX/) [50]. The hazard ratios (HRs) and 95% CIs for fracture incidence based on predicted BMDs were estimated using a Cox model, and the absolute risk was also calculated [51]. We assessed the predictive performance of the Cox model using the Concordance Index (C-index). Survival analyses were performed using the R software (version 4.1.1) with the “survival” package (version 3.2) and “survminer” package (version 4.1.2).

Results

Characteristics of the study population

In total, we included 320,182 individuals from UKBB for model construction and evaluation in our study, which were split to 8 data sets: the UKBB British Training set (N = 17,964), the UKBB British Model Selection set (N = 5,973), the UKBB British Test set (N = 5,977), the UKBB other White Test set (N = 1,960), the UKBB African ancestry set (N = 218), the UKBB East Asian ancestry set (N = 112), the UKBB other ancestry set (N = 795), and the UKBB Fracture case-control set (N = 287,183). We also included 12 data sets from 6 independent cohorts to evaluate the selected prediction models: LOS_CAU (N = 2,863), LOS_AFR (N = 2,097), KCOS_CAU (N = 2,271), COS_EAS (N = 1,569), MrOS_CAU (N = 4,587), MrOS_AFR (N = 181), MrOS_EAS (N = 165), MrOS_HIS (N = 111), WHI_AFR (N = 671), WHI_HIS (N = 393), CHS_CAU (N = 432), and CHS_AFR (N = 190). The demographic characteristics of each data set are summarized in Table 1. The first 2 genetic PCs of all individuals and the GD to the training data are shown in S5 Fig. As expected, the data set originating from the European ancestry population demonstrated a notably lower GD compared to data sets representing other racial groups. Although the first 2 genetic PCs of these racial groups exhibited distinct clustering patterns based on ancestry, no clear demarcation between these categories was evident. In line with previous research findings [52,53], human diversity is observed along a genetic ancestry continuum, devoid of well-defined clusters.

10.1371/journal.pmed.1004451.t001 Table 1 Participant characteristics by data set.

Datasets	Sample
size	Age
Mean
(SD)	Women
N
(%)	Height
Mean
(SD)	Weight
Mean
(SD)	Smoking
N
(%)	Drinking
N
(%)	Exercise
N
(%)	FNK-BMD
mean
(SD)	SPN-BMD
mean
(SD)	Fracture
N
(%)	
UKBB
Training	17,964	63.95
(7.56)	9,113
(50.73)	170.56
(9.42)	75.62
(14.99)	6,631
(36.91)	17,456
(97.17)	16,289
(90.68)	0.94
(0.14)	1.1
(0.18)	616
(3.43)	
UKBB
Model Selection	5,973	64.15
(7.55)	2,973
(49.77)	170.56
(9.39)	75.82
(15.47)	2,273
(38.05)	5,797
(97.05)	5,394
(90.31)	0.94
(0.14)	1.1
(0.18)	207
(3.47)	
UKBB
British Test	5,977	63.82
(7.6)	3,082
(51.56)	170.47
(9.3)	75.68
(15.18)	2,286
(38.25)	5,808
(97.17)	5,405
(90.43)	0.94
(0.14)	1.1
(0.18)	212
(3.55)	
UKBB
Other White	1,960	62.86
(7.75)	1,086
(55.41)	169.88
(9.48)	75.13
(15.81)	890
(45.41)	1,886
(96.22)	1,785
(91.07)	0.92
(0.14)	1.08
(0.18)	75
(3.83)	
UKBB
AFR	218	58.41 (7.09)	117 (53.67)	169.5 (9.36)	80.33 (17.33)	68
(31.19)	191 (87.61)	186 (85.32)	1.05 (0.16)	1.21 (0.19)	3
(1.38)	
UKBB
EAS	112	59.96 (6.99)	63
(56.25)	163.82 (7.91)	62.5
(11.2)	23
(20.54)	96
(85.71)	91
(81.25)	0.9
(0.15)	1.06 (0.19)	2
(1.79)	
UKBB
Other Ancestry	795	61.05 (8.18)	390 (49.06)	166.93 (9.47)	72.43 (14.33)	246 (30.94)	650 (81.76)	689 (86.67)	0.95 (0.14)	1.1
(0.18)	13 (1.64)	
UKBB
Case-control	287,183	57.1
(8.02)	154,601
(53.83)	168.75
(9.2)	78.68
(16.02)	136,125
(47.4)	277,919
(96.77)	175,539
(61.12)	NA	NA	17,490
(6.09)	
LOS_CAU	2,863	38.83
(12.3)	1,510
(52.74)	169.59
(9.23)	76.25
(18.24)	1,477
(51.59)	2,391
(83.51)	2,172
(75.86)	0.83
(0.14)	1.02
(0.14)	925
(32.31)	
LOS_AFR	2,097	39.62
(9.53)	981
(46.78)	170.16
(9.03)	84.6
(21.47)	1,145
(54.6)	1,313
(62.61)	1,381
(65.86)	0.93
(0.15)	1.1
(0.15)	335
(15.98)	
KCOS_CAU	2,271	51.26
(13.71)	1,715
(75.52)	166.37
(8.46)	75.28
(17.53)	761
(33.51)	1,494
(65.79)	1,711
(75.34)	0.79
(0.13)	1.02
(0.15)	967
(42.58)	
COS_EAS	1,569	34.52
(13.23)	798
(50.86)	162.8
(42.32)	58.8
(39.23)	122
(7.78)	141
(8.99)	750
(47.8)	0.81
(0.13)	0.95
(0.13)	124
(7.9)	
MrOS_CAU	4,587	73.97
(5.96)	0(0)	174.49
(6.62)	83.5
(13.07)	2,862
(62.39)	2,977
(64.9)	NA	0.78
(0.12)	1.07
(0.18)	684
(14.91)	
MrOS_AFR	181	72.04
(5.42)	0 (0)	174.37
(7.19)	87.08
(15.76)	114
(62.98)	92
(50.83)	NA	0.87
(0.15)	1.13
(0.21)	11
(6.08)	
MrOS_EAS	165	72.88
(5.04)	0 (0)	167.07
(5.98)	70.29
(8.61)	88
(53.33)	75
(45.45)	NA	0.75
(0.11)	1.04
(0.18)	20
(12.12)	
MrOS_HIS	111	71.84
(4.44)	0 (0)	170.39
(5.99)	81.83
(12.51)	67
(60.36)	84
(75.68)	NA	0.8
(0.12)	1.04
(0.18)	12
(10.81)	
WHI_AFR	671	60.52
(6.76)	671
(100)	162.33
(5.52)	82.49
(17.06)	NA	NA	NA	0.83
(0.13)	1.05
(0.17)	164
(24.44)	
WHI_HIS	393	60.04
(7.07)	393
(100)	157.79
(5.46)	73.47
(14.86)	NA	NA	NA	0.73
(0.11)	0.97
(0.15)	108
(27.48)	
CHS_CAU	432	71.06
(4.58)	131
(30.32)	169.32
(8.6)	170.41
(27.6)	257
(59.49)	313
(72.45)	408
(94.44)	0.77
(0.12)	1.15
(0.23)	75
(17.36)	
CHS_AFR	190	71.63
(4.54)	92
(48.42)	167.51
(8.87)	177.9
(29.53)	110
(57.89)	84
(44.21)	176
(92.63)	0.84
(0.13)	1.19
(0.23)	38
(20)	
FNK-BMD: Bone mineral density at femoral neck.

SPN-BMD: Bone mineral density at lumbar spine.

“NA” indicates that the covariate was not available at that data set.

“SD” indicates standard deviation.

“UKBB” indicates the UK biobank.

“LOS” indicates the Louisiana Osteoporosis Study.

“KCOS” indicates the Kansas City Osteoporosis Study.

“COS” indicates the China Osteoporosis Study.

“MrOS” indicates the Osteoporotic Fractures in Men Study.

“WHI” indicates the Women’s Health Initiative Clinical Trial and Observational Study.

“CHS” indicates the Cardiovascular Health Study.

“AFR” indicates African American.

“CAU” indicates Caucasian.

“EAS” indicates East Asian.

“HIS” indicates Hispanic/Latino.

Variance explained of prediction models in the UKBB Model Selection set

The performance of each model in UKBB Model Selection set is shown in Fig 2. For the regression model with only clinical factors, the R2 of the FNK-BMD and SPN-BMD were 25.388% with a 95% confidence interval (CI) of [25.386%, 25.390%] and 45.313% (95% CI [45.312%, 45.314%]), respectively. For the prediction models including both clinical factors and genetic variants, the performance of the models improved moderately when we set the P-value threshold from 5×10−8 to 5×10−6. However, when we continued to increase the threshold, the performance of the models deteriorated due to overfitting. The prediction models with the best performance were selected as the prediction models for the follow-up analysis. For FNK-BMD, the prediction model was trained with LASSO when the threshold was set at 5×10−6 which resulted in an R2 of 27.70% (95% CI [27.56%, 27.84%]), indicating that the prediction model could explain 27.7% of the variance in FNK-BMD in the UKBB British Model Selection set. Compared with the regression model with only clinical factors, the LASSO model for FNK-BMD improved by 2.3% in R2. For SPN-BMD, the prediction model was trained with LASSO when set the threshold as 5×10−7, which resulted in an R2 of 48.28% (95% CI [48.23%, 48.34%]), indicating that the prediction model could explain 48.28% of the variance in SPN-BMD. Compared with the regression model with only clinical factors, the LASSO model for SPN-BMD improved by 3% in R2.

10.1371/journal.pmed.1004451.g002 Fig 2 The performance of each model in UKBB British Model Selection set.

The R2 of each prediction model for FNK-BMD (A) and SPN-BMD (B) in the UKBB British Model Selection set, respectively. For prediction models including both clinical factors and genetic variants, the performance of the models improved moderately when we set the P-value threshold from 5×10−8 to 5×10−6. However, when we continued to increase the threshold, the performance of the models deteriorated due to overfitting. The prediction models with the best performance were selected as the prediction models for the follow-up analysis. For FNK-BMD, the prediction model was trained with LASSO when the threshold was set at 5×10−6. For SPN-BMD, the prediction model was trained with LASSO when set the threshold as 5×10−7. The vertical short lines represent the upper and lower bounds of the 95% CI. FNK-BMD: Bone mineral density at femoral neck. SPN-BMD: Bone mineral density at lumbar spine. R2: The coefficient of determination. LR: Linear regression. PRS: polygenic risk score. LASSO: Regression with Least absolute shrinkage and selection operator. CNN: Convolutional neural network.

Performance of prediction models in the UKBB testing data sets

We applied the prediction models with the highest R2 to predict the BMD values and calculated the R2 and PCC between the predicted and actual BMD values in each test set. As shown in Table 2 and Fig 3, the models for predicting FNK-BMD and SPN-BMD performed similarly in the UKBB British Test set and the UKBB other White set. The model for predicting FNK-BMD explained 24.03% (95% CI [23.89%, 24.17%]) and 24.06% (95% CI [23.83%, 24.29%]) of variance in the UKBB British Test set and the UKBB other White set, respectively. The model for predicting SPN-BMD explained 44.79% (95% CI [44.74%, 44.84%]) and 44.17% (95% CI [44.08%, 44.25%]) of variance, respectively. However, the model performance dropped significantly in the UKBB African ancestry set, East Asian ancestry set, and other ancestry set. Especially among the UKBB African ancestry set, the R2 for FNK-BMD and SPN-BMD were −0.089 (95% CI [−0.173, −0.005]) and −0.333 (95% CI [−0.446, −0.221]), respectively, indicating that the models cannot capture any BMD variability. As suggested by some previous research, the performance decline might result at least partially from decreased similarity in genetic profile [35]. Additionally, for each type of models, we selected the one with the highest R2 from the UKBB Model selection set for testing across all test data sets. The detailed outcomes are presented in S3 and S4 Tables. Similar to the LASSO-based models, these models demonstrated better performance in the in the UKBB British Test set and the UKBB other White set, while the performance varied significantly across other populations.

10.1371/journal.pmed.1004451.g003 Fig 3 The predictive performance of each model in UKBB Testing sets.

(A) The R2 between the predicted and actual BMDs in the UKBB British testing set, the UKBB other White set, and the UKBB other Ancestry set. (B) The PCC between the predicted and actual BMDs in the UKBB British testing set, the UKBB other White set, and the UKBB other Ancestry set. The models for predicting FNK-BMD and SPN-BMD performed similarly in the UKBB British Test set and the UKBB other White set. However, the model performance drops significantly in the UKBB other Ancestry set. The vertical short lines represent the upper and lower bounds of the 95% CI. FNK-BMD: Bone mineral density at femoral neck. SPN-BMD: Bone mineral density at lumbar spine. R2: The coefficient of determination. PCC: Pearson correlation coefficient.

10.1371/journal.pmed.1004451.t002 Table 2 Performance of the BMD prediction models in the testing sets.

Data set	Avg
GD	FNK-BMD	SPN-BMD	
R2	Lower bound	Upper bound	PCC	Lower bound	Upper bound	R2	Lower bound	Upper bound	PCC	Lower bound	Upper bound	
UKBB	Britsh
Testing	0.026	0.240	0.239	0.242	0.493	0.492	0.494	0.448	0.447	0.448	0.670	0.670	0.670	
Other
White	0.043	0.241	0.238	0.243	0.501	0.500	0.502	0.442	0.441	0.443	0.669	0.668	0.669	
EAS	0.274	0.117	0.088	0.147	0.505	0.498	0.513	0.215	0.197	0.234	0.606	0.600	0.611	
AFR	0.352	−0.089	−0.173	−0.005	0.411	0.392	0.430	−0.333	−0.446	−0.221	0.408	0.395	0.421	
Other
Ancestry	0.140	0.175	0.166	0.185	0.465	0.462	0.468	0.310	0.302	0.319	0.598	0.594	0.602	
Independent
cohorts	LOS
AFR	0.346	−0.420	−0.660	−0.180	0.431	0.414	0.447	−1.830	−2.145	−1.515	0.220	0.214	0.227	
LOS
CAU	0.047	−0.250	−0.624	0.125	0.564	0.559	0.568	−1.641	−2.046	−1.236	0.302	0.295	0.309	
COS
EAS	0.274	−1.866	−2.553	−1.178	0.264	0.262	0.266	−4.471	−4.990	−3.952	0.161	0.154	0.168	
KCOS
CAU	0.033	−0.423	−1.178	0.332	0.600	0.598	0.603	−1.202	−1.552	−0.853	0.340	0.331	0.348	
MrOS
AFR	0.329	−0.985	−1.705	−0.266	0.340	0.322	0.357	−0.907	−1.066	−0.748	0.058	0.035	0.081	
MrOS
CAU	0.042	−0.762	−1.640	0.116	0.377	0.374	0.380	−0.657	−0.878	−0.436	0.207	0.198	0.215	
MrOS
EAS	0.278	−1.131	−2.148	−0.114	0.264	0.256	0.272	−0.725	−0.946	−0.505	0.202	0.188	0.215	
MrOS
HIS	0.120	−0.957	−1.849	−0.065	0.163	0.154	0.171	−0.804	−1.045	−0.562	0.161	0.141	0.181	
CHS
AFR	0.236	−3.136	−3.720	−2.551	0.412	0.401	0.423	−3.937	−4.583	−3.291	0.320	0.308	0.332	
CHS
CAU	0.065	−4.892	−5.613	−4.170	0.437	0.434	0.440	−3.772	−4.374	−3.170	0.282	0.278	0.286	
WHI
AFR	0.326	−1.000	−1.840	−0.161	0.435	0.415	0.455	−1.278	−1.535	−1.022	0.248	0.237	0.259	
WHI
HIS	0.120	−0.827	−1.817	0.163	0.542	0.537	0.547	−1.005	−1.307	−0.702	0.322	0.312	0.332	
FNK-BMD: Bone mineral density at femoral neck.

SPN-BMD: Bone mineral density at lumbar spine.

AvgGD: The average Genetic distance of the individuals in the testing set to the training set.

R2: The coefficient of determination.

PCC: Pearson correlation coefficient.

Lower bound: Lower bound of the 95% confidence interval.

Upper bound: Upper bound of the 95% confidence interval.

“UKBB” indicates the UK biobank.

“LOS” indicates the Louisiana Osteoporosis Study.

“KCOS” indicates the Kansas City Osteoporosis Study.

“COS” indicates the China Osteoporosis Study.

“MrOS” indicates the Osteoporotic Fractures in Men Study.

“WHI” indicates the Women’s Health Initiative Clinical Trial and Observational Study.

“CHS” indicates the Cardiovascular Health Study.

“AFR” indicates African American.

“CAU” indicates Caucasian.

“EAS” indicates East Asian.

“HIS” indicates Hispanic/Latino.

Association between predicted BMDs and fragility fracture in the UKBB Fracture case-control set

To examine the relationship between the predicted BMDs and fragility fracture risk, we first conducted a T test to compare the mean difference of the predicted BMDs between the fragility fracture cases and the controls in the UKBB fracture case-control set. As expected, the fragility fracture cases had significantly lower mean predicted FNK-BMD and SPN-BMD than the controls (T test P < 0.001) (Fig 4A). Furthermore, we stratified the population into different groups based on the predicted BMD values and found that a lower predicted BMD was associated with a higher prevalence of fragility fracture (Fig 4B and 4C). For example, among individuals with a predicted FNK-BMD below the fifth percentile, 1,713 out of 14,360 (11.93%) experienced a fragility fracture within the next 10 years. Similarly, among those with a predicted SPN-BMD below the fifth percentile, 1,557 out of 14,360 (10.84%) experienced a fragility fracture. In contrast, the prevalence of fragility fractures was significantly lower among individuals with a predicted BMD above the 95th percentile, with only 3.84% (551 out of 14,360) for FNK-BMD and 3.93% (564 out of 14,360) for SPN-BMD. These results supported that both the predicted FNK-BMD and SPN-BMD had a negative correlation with fracture risk.

10.1371/journal.pmed.1004451.g004 Fig 4 Association between predicted BMDs and fragility fracture in UKBB Fracture case-control set.

(A) T test showed that the fragility fracture cases had significantly lower mean predicted FNK-BMD and SPN-BMD than the controls in the UKBB fracture case-control set. (B, C) By stratifying the population into different groups based on the predicted BMD values, we found that a lower predicted BMD was associated with a higher prevalence of fragility fracture. FNK-BMD: Bone mineral density at femoral neck. SPN-BMD: Bone mineral density at lumbar spine.

The predictive power of the predicted BMDs for fragility fracture risk

To further evaluate the predictive power of the predicted BMDs for fragility fracture risk, we first conducted a univariate survival analysis with the Kaplan–Meier model to estimate the 10-year cumulative incidence. The log-Rank test showed a significant difference in the cumulative incidence of fragility fracture among different BMD groups (P < 0.001; Fig 5A and 5B). Next, we integrated the predicted BMDs and clinical factors to perform a multivariate survival analysis with cox regression to predict the fracture risk of the participants in the next 10 years. As expected, the predicted FNK-BMD and SPN-BMD were both significant in the cox model (P < 0.001). The HRs of the predicted FNK-BMD and SPN-BMD were 0.83 (95% CI [0.79, 0.88], corresponding to a 1.44% difference in 10-year absolute risk per standard deviation of BMD) and 0.72 (95% CI [0.68, 0.76], corresponding to a 1.64% difference in 10-year absolute risk per standard deviation of BMD), respectively, which means that for every increase of one standard deviation in BMD, the fracture risk will decrease by 17% and 28%, respectively (Fig 5D). Compared with using only the clinical risk factors, combining the predicted BMDs with the clinical risk factors significantly improved the risk prediction for fragility fracture (likelihood ratio test p-values <0.001; Fig 5C).

10.1371/journal.pmed.1004451.g005 Fig 5 The strength of predicted BMDs as the predictors of fragility fracture.

(A, B) The univariate survival analysis showed a significant difference in the cumulative incidence of fragility fracture among different BMD groups (censored at 10 years). (C) Compared with using only the clinical risk factors, combining the predicted BMDs with the clinical risk factors significantly improved the risk prediction for fragility fracture. (D) The predicted FNK-BMD and SPN-BMD were both significant in the cox model. For every increase of one standard deviation in BMD, the fracture risk will decrease by 17% and 28%, respectively. FNK-BMD: Bone mineral density at femoral neck. SPN-BMD: Bone mineral density at lumbar spine.

Prediction models varied across the independent data sets

The performance of BMD predictive models significantly deteriorated across all the independent data sets, as shown in Table 2. For FNK-BMD, the model performed best in the LOS_CAU, with an R2 of −0.250 (95% CI [−0.624, 0.125]). For SPN-BMD, the model performed best in the MrOS_CAU, with an R2 of −0.657 (95% CI [−0.878, −0.436]). However, the R2 were less than 0 for both FNK-BMD and SPN-BMD prediction models in all the independent data sets, indicating that the models could not capture any variance of BMD.

Next, we calculated the odds ratios (OR) using the logistic regression to assess the association between the predicted BMDs and fragility fracture risk. As shown in Table 3, the OR between predicted FNK-BMD and fragility fracture was 0.575 (95% CI [0.414, 0.797]) (P-value = 0.001) in the KCOS_CAU set, indicating that for every one standard deviation increase in BMD, the risk of fracture is reduced by 42.5%. The OR between predicted SPN-BMD and fragility fracture was 0.741 (95% CI 0.589–0.930) (P-value = 0.010) in the KCOS_CAU set, indicating that for every one standard deviation increase in BMD, the risk of fracture is reduced by 25.9%. The OR value between predicted SPN-BMD and fragility fracture was 0.839 (95% CI [0.736, 0.957]) (P-value = 0.009) in the MrOS_CAU set, indicating that for every one standard deviation increase in BMD, the risk of fracture was reduced by 16.1%. However, the predicted BMDs did not show a significant association with fracture in the other data sets.

10.1371/journal.pmed.1004451.t003 Table 3 The association between the predicted BMDs and fragility fracture risk in independent cohorts.

Trait	Data set	Fracture, N (%)	OR	Lower bound	Upper bound	P	
FNK-BMD	LOS_CAU	925 (32.31)	1.087	0.823	1.437	0.557	
LOS_AFR	335 (15.98)	0.897	0.744	1.082	0.254	
KCOS_CAU	967 (42.58)	0.575	0.414	0.797	0.001*	
COS_EAS	124 (7.9)	3.069	0.622	15.264	0.169	
MrOS_CAU	684 (14.91)	0.962	0.779	1.188	0.717	
MrOS_AFR	11 (6.08)	1.870	0.630	6.217	0.279	
MrOS_EAS	20 (12.12)	2.149	0.808	5.858	0.127	
MrOS_HIS	12 (10.81)	0.861	0.274	2.996	0.803	
WHI_AFR	164 (24.44)	1.056	0.808	1.386	0.690	
WHI_HIS	108 (27.48)	0.961	0.576	1.616	0.879	
CHS_CAU	75 (17.36)	3.252	0.944	11.368	0.063	
CHS_AFR	38 (20)	1.820	0.813	4.259	0.154	
SPN-BMD	LOS_CAU	925 (32.31)	0.911	0.737	1.125	0.385	
LOS_AFR	335 (15.98)	1.009	0.838	1.216	0.922	
KCOS_CAU	967 (42.58)	0.741	0.589	0.930	0.010*	
COS_EAS	124 (7.9)	1.183	0.439	3.123	0.737	
MrOS_CAU	684 (14.91)	0.839	0.736	0.957	0.009*	
MrOS_AFR	11 (6.08)	0.602	0.256	1.347	0.223	
MrOS_EAS	20 (12.12)	1.462	0.735	2.936	0.275	
MrOS_HIS	12 (10.81)	0.776	0.259	2.309	0.646	
WHI_AFR	164 (24.44)	1.048	0.832	1.320	0.691	
WHI_HIS	108 (27.48)	0.742	0.496	1.103	0.142	
CHS_CAU	75 (17.36)	0.814	0.335	1.957	0.646	
CHS_AFR	38 (20)	0.521	0.228	1.175	0.116	
FNK-BMD: Bone mineral density at femoral neck.

SPN-BMD: Bone mineral density at lumbar spine.

OR: Odds ratio.

Lower bound: Lower bound of the 95% confidence interval.

Upper bound: Upper bound of the 95% confidence interval.

“LOS” indicates the Louisiana Osteoporosis Study.

“KCOS” indicates the Kansas City Osteoporosis Study.

“COS” indicates the China Osteoporosis Study.

“MrOS” indicates the Osteoporotic Fractures in Men Study.

“WHI” indicates the Women’s Health Initiative Clinical Trial and Observational Study.

“CHS” indicates the Cardiovascular Health Study.

“AFR” indicates African American.

“CAU” indicates Caucasian.

“EAS” indicates East Asian.

“HIS” indicates Hispanic/Latino.

As suggested in a recent study, the decline in performance can be partially explained by GD [35]. In our investigation, we observed a statistically significant but modest correlation between the residuals of predicted and true BMD values (standardized by the standard deviation of true BMD values) and individual-level GD for both FNK-BMD (PCC = 0.04, P-value <0.001) and SPN-BMD (PCC = 0.18, P-value <0.001). Apart from differences in genetic architecture across cohorts, other factors influencing transferability include genotype–environment interactions and population-specific causal variants. To illustrate, we conducted a backward stepwise regression to develop predictive models for covariates related to BMD in each data set. We noted not only variations in the coefficients of covariates across cohorts but also considerable diversity in the combinations of covariates included in the regression models (see S5 and S6 Tables). For example, while age proved to be a significant predictor for FNK-BMD in most data sets, it was not selected in the COS_EAS set, all MrOS sets, and WHI_HIS set. This discrepancy suggests that the contribution of clinical factors to BMD variation differs across cohorts, leading to variations in the performance of prediction models in different cohorts. This consideration is crucial when applying prediction models to other cohorts, emphasizing the need for adjustments based on local population data.

Discussion

In summary, we built and tested a series of prediction models for FNK-BMD and SPN-BMD separately in 32,999 independent participants with DXA-BMD measurements from UKBB, and finally the LASSO regression was selected as the best prediction models. We observed that integrating the clinical risk factors and genetic variants could slightly improve the predictive performance in the European white population from UKBB. Predicted BMDs for 287,183 European individuals without DXA-BMD measurements were strongly associated with fragility fracture risk. However, the models’ performance decreased significantly on 5 UKBB test sets and 12 independent cohorts of diverse ancestries (totaling over 15,000 individuals).

Unlike previous studies that used only GWAS-significant SNPs, we included a broader range of SNPs, improving the variance explained in BMD. For example, a previous study including 62 SNPs associated explained only 1.3% (approximately 42% when adding gender, age, and weight) of the variance in FNK-BMD and 1.6% (approximately 35% when adding gender and weight) of the variance in SPN-BMD [16], while our models explained up to 48% for SPN-BMD. Interestingly, we observed an obvious difference in the explanation in variance of FNK-BMD and SPN-BMD by the prediction model. The prediction model could explain approximately 28% of the variance in FNK-BMD while approximately 48% of the variance in SPN-BMD. This may be due to several reasons: First, the SPN and FNK have different biological and structural characteristics that might influence the accuracy of BMD predictions. Second, the relative contributions of genetic and clinical factors vary across skeletal sites [54]. For example, FNK-BMD is more sensitive to the clinical effects such as physical activity and nutritional intake, while SPN-BMD is more sensitive to body weight [11,55–57]. In our study, we determined that the most suitable P-value thresholds for FNK-BMD and SPN-BMD were 5×10−6 and 5×10−7, encompassing 836 and 276 SNPs, respectively. However, it is crucial to note that the optimal threshold is not universally constant; rather, it depends on the specific trait and the sample size of the training data. For example, a prior study developed a PRS for eBMD involving 341,449 individuals in the training set, and they incorporating 21,717 SNPs with a P-value threshold of 5×10−4 [17]. Given that obtaining DXA-BMD measurements is more challenging and costly compared to eBMD, our training set was limited to a sample size of 17,964. Naturally, there are likely numerous BMD-associated SNPs have been overlooked. As more DXA-BMD data accumulates and computational power improves, we believe that future studies have the potential to expand the inclusion of selected SNPs, thereby capturing a broader spectrum of BMD variation.

Furthermore, many existing PRS researches only accounted for the heritable component of a trait and ignores the significant role of environmental and lifestyle factors in disease etiology, further limiting their overall prediction power. In contrast, we considered both genetic and clinical factors such as sex, age, height, and weight, which have been proven to be closely associated with BMD [11]. By highlighting the significance of clinical factors, our research provides a more comprehensive understanding of BMD prediction and the assessment of osteoporosis risk.

Another contribution of our study to existing research is that we underscore the imperative for training models on diverse population. As has been widely noted, when a genetic prediction approach is applied to populations of different ancestries, its predictive performance may vary, sometimes to a large extent [58]. Our study further highlights that, the prediction approaches vary in predictive performance not only among populations of different ethnics, but also among populations of similar ethnic but different geographic background. This conclusion was also proposed in a previous PRS study for coronary artery disease (CAD) [59]. It was reported that GD is an important contributor to decay of predictive performance, which corresponds to decreased similarity in genetic profile (assessed by relatedness, linkage disequilibrium, and/or minor allele frequency differences, fixation index (Fst) and so on) between testing individuals and training data [35]. However, GD cannot fully explain the reduced predictive performance, as a decrease in the predictive efficacy was also observed in data sets from other European populations with different geographical backgrounds. Additionally, the absence of certain variants can also affect the accuracy of predictive models. Such rare variants are particularly noticeable because their frequency in the population is too low to be detected. For example, the SNP rs554533790, located in an intergenic region, appears in the 1000 Genomes data set with a minor allele frequency of less than 0.005. In our study, rs554533790 has a minor allele frequency of approximately 0.0004 in the UKBB CAU populations, and zero in the East EAS and AFR populations. However, this variant was not detected in other independent data sets. Although rare, rs554533790 has a high weight in our FNK-BMD predictive model, even higher than age, highlighting its significance. Therefore, the absence of this SNP in other independent data sets will inevitably reduce the model’s predictive performance. However, as a rare variant in an intergenic region, the function of rs554533790 remains largely unknown. Further functional or mechanistic studies are needed to elucidate its role in bone. Some clinical factors also contribute to performance degradation, such as age and weight, which were identified as crucial elements in our FNK- and SPN-BMD prediction models. The most striking explanation for these variations is the disparate age distributions and sex ratios present in the data sets. Moreover, differences in diet, physical activity levels, sun exposure, and other lifestyle and environmental factors across populations can result in inconsistent impacts of these clinical factors on BMD, even among individuals from the similar ethnic background. These variances may represent another contributing factor to the diminished predictive capability of our model in different populations. Therefore, every single population-specific risk prediction model needs to be derived on their very own training data set or at least verified for application on the target population, even if the models trained in other populations of the same ethnicity are available.

Strengths of our study include the large sample size from UKBB and comparison of various types of prediction models that integrated of both genetic and clinical factors, which enable us to further enhance the performance of BMD predictions. However, this study has some limitations: First, the distribution of clinical factors varied across the UKBB data sets and the independent cohorts. Second, the included genetic factors included in our model is mainly based on GWAS results, which can only detect SNVs associated with phenotypes, but not other types of variants or gene expression regulation. Besides SNVs, other types of genomic variants, such as copy number variants (CNVs) and genetic rearrangements (GRs) may also affect BMD [60,61]. These variants may alter gene expression or function by causing gene duplication, deletion, translocation, or inversion. BMD is also affected by the complex gene expression regulation, such as epigenetic modification, transcription factor, and noncoding RNA [11,62]. These factors may regulate gene activity and interaction in different cell types, developmental stages, and environmental conditions. Third, the selection of fragility fracture cases utilizes self-report, and hospital inpatient data for the definition of fragility fracture cases, which may lead to some incorrect or inaccurate classification.

The most significant aspect of our research is demonstrating that this genomic prediction method can be used to forecast the risk of fragility fractures in advance. Though our model did not capture all BMD variance, it is still helpful for assessing fracture risk. We found that the predicted BMDs could improve the performance of fracture prediction over and above that of clinical risk factors alone, such as height, weight, smoking, drinking, and exercise. Currently, DXA-BMD measurement of the spine and hip is the gold standard imaging test for diagnosing osteoporosis and assessing fragility fracture risk [1]. However, due to its high cost, large size of the equipment and requirement for professional operation, DXA measurement is not suitable for large-scale population screening, especially in medically underdeveloped areas. With the development of whole genome sequencing technology, the cost of sequencing a human genome has been significantly reduced, which means that gene testing technology has become more accessible to the public, and can provide strong support for various fields, such as personalized medicine, precision medicine, and preventive medicine [62]. For areas where DXA measurement is inconvenient, obtaining genetic information may be more convenient and cost-effective. Moreover, building prediction models based on genetic information may be particularly useful in this situation, because the one-time cost of genotyping can be used not only for evaluation of osteoporosis but also for prediction of other complex diseases, such as atrial fibrillation, CAD, type 2 diabetes, breast cancer, colorectal cancer, and prostate cancer [21,63–67]. It is imperative to highlight that the predictive models should serve as tools to support, not replace, the clinical judgment of physicians and the discernment of individuals. Decisions regarding clinical examinations or proactive interventions should not be made on the predictive outcomes alone. Vigilance against disease risks must be maintained, even if the model indicates a lower probability of illness. A realistic expectation for our approach is to identify individuals who have a higher risk of osteoporosis or fractures compared to the general populace, thereby encouraging them to recognize potential health threats, embrace positive lifestyle changes, and pursue timely diagnosis and preventive measures.

Of course, before these genomic prediction models can be formally applied in clinical settings, further research is needed. This includes collecting data from diverse populations to build prediction models that benefit a wider range of individuals, developing more complex models that incorporate a broader spectrum of variant types to enhance predictive accuracy, and using interpretable models and mechanistic studies to identify potential therapeutic targets.

In conclusion, our study shows that incorporating genetic factors can enhance DXA-BMD prediction beyond clinical factors alone. Adjusting SNP inclusion thresholds (e.g., 5×10−6 or 5×10−7) instead of only using GWAS-significant SNPs (P < 5×10−8) can improve model performance, depending on the trait and sample size. The clinical utility of BMD prediction models can help prioritize individuals at high risk in early stages. This early identification encourages individuals to take their risk seriously, participate actively in necessary early screenings, and proactively improve their lifestyle habits. However, this is contingent upon having sufficient data from populations with similar backgrounds to construct accurate predictive models. Our study also emphasizes the importance of prediction model training in diverse populations, only across populations of different ethnicities but also among populations with similar ethnic backgrounds but different geographic origins.

Supporting information

S1 Table Used ICD10 and self-reported codes for extracting fracture and RA cases.

RA, rheumatoid arthritis.

(XLSX)

S2 Table Input features of prediction models.

LR, linear regression; PRS, polygenic risk score; LASSO, regression with least absolute shrinkage and selection operator; CNN, convolutional neural network; FNK-BMD, bone mineral density at femoral neck; SPN-BMD, bone mineral density at lumbar spine.

(XLSX)

S3 Table Performance of all 4 FNK-BMD prediction models in the testing data sets.

LR, linear regression; PRS, polygenic risk score; LASSO, regression with least absolute shrinkage and selection operator; CNN, convolutional neural network; R2, the coefficient of determination; PCC, Pearson correlation coefficient.

(XLSX)

S4 Table Performance of all 4 SPN-BMD prediction models in the testing data sets.

LR, linear regression; PRS, polygenic risk score; LASSO, regression with least absolute shrinkage and selection operator; CNN, convolutional neural network; R2, the coefficient of determination; PCC, Pearson correlation coefficient.

(XLSX)

S5 Table Coefficients of covariates for FNK-BMD in each cohort.

Coef: Coefficient of the covariates after stepwise regression. Pval: P-values of the covariates after stepwise regression. ‘-‘ represents that the covariate was removed from the stepwise regression (or not available in this data set).

(XLSX)

S6 Table Coefficients of covariates for SPN-BMD in each cohort.

Coef: Coefficient of the covariates after stepwise regression. Pval: P-values of the covariates after stepwise regression. ‘-‘ represents that the covariate was removed from the stepwise regression (or not available in this data set).

(XLSX)

S1 Fig The architecture diagram of the CNN model.

The convolutional layers are 1-dimensional (1 × 4) with pooling layers.

(TIF)

S2 Fig Selection of λ for LASSO models.

The R2 (coefficient of determination) were calculated within the UKBB Model Selection set.

(TIF)

S3 Fig Selection of drop rate for CNN models.

The R2 (coefficient of determination) were calculated within the UKBB Model Selection set.

(TIF)

S4 Fig The diagram of the models and the processing of model training, selection, and evaluation.

PRS, polygenic risk score; LASSO, regression with least absolute shrinkage and selection operator; R2, the coefficient of determination; PCC, Pearson correlation coefficient.

(TIF)

S5 Fig The first 2 genetic PCs of all individuals and the GD to the training data.

PC, principal component of genetic profile; GD, individual-level genetic distance.

(TIF)

S1 Appendix Supplementary methods.

(DOCX)

This research has been conducted using the UK Biobank Resource under project number 63047. We appreciate the generosity of all cohort volunteers. This work was supported in part by the High Performance Computing Center of Central South University.

Abbreviations

BMD bone mineral density

BMI body mass index

CAD coronary artery disease

CHS Cardiovascular Health Study

CI confidence interval

CNN convolutional neural network

CNV copy number variant

COS China Osteoporosis Study

DXA dual-energy X-ray absorptiometry

FNK femoral neck

GD genetic distance

GR genetic rearrangement

GWAS genome-wide association study

HR hazard ratio

HRC Haplotype Reference Consortium

KCOS Kansas City Osteoporosis Study

LASSO least absolute shrinkage and selection operator

LD linkage disequilibrium

LOS Louisiana Osteoporosis Study

LR linear regression

MBGD mini-batch gradient descent

MSE mean squared error

OR odds ratio

PC principal component

PCA principal component analysis

PCC Pearson correlation coefficient

PRS polygenic risk score

QUS quantitative ultrasound

RA rheumatoid arthritis

ReLU rectified linear unit

SNP single-nucleotide polymorphism

SPN lumbar spine

UKBB UK Biobank

10.1371/journal.pmed.1004451.r001
Decision Letter 0
Dodd Philippa C. Senior Editor
© 2024 Philippa C. Dodd
2024
Philippa C. Dodd
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Submission Version0
26 Mar 2024

Dear Dr Liu,

Thank you for submitting your manuscript entitled "The performance of genetic-enhanced DXA-BMD predicting models trained in UK biobank varies across diverse ethnic and geographical populations" for consideration by PLOS Medicine.

Your manuscript has now been evaluated by the PLOS Medicine editorial staff and I am writing to let you know that we would like to send your submission out for external peer review.

However, before we can send your manuscript to reviewers, we need you to complete your submission by providing the metadata that is required for full assessment. To this end, please login to Editorial Manager where you will find the paper in the 'Submissions Needing Revisions' folder on your homepage. Please click 'Revise Submission' from the Action Links and complete all additional questions in the submission questionnaire.

Please re-submit your manuscript on or after April 2nd 2024.

Login to Editorial Manager here: https://www.editorialmanager.com/pmedicine

Once your full submission is complete, your paper will undergo a series of checks in preparation for peer review. Once your manuscript has passed all checks it will be sent out for review.

Feel free to email us at plosmedicine@plos.org if you have any queries relating to your submission.

Kind regards,

Pippa

Philippa C. Dodd, MBBS MRCP PhD

PLOS Medicine

pdodd@plos.org

10.1371/journal.pmed.1004451.r002
Decision Letter 1
Dodd Philippa C. Senior Editor
© 2024 Philippa C. Dodd
2024
Philippa C. Dodd
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Submission Version1
17 May 2024

Dear Dr. Liu,

Many thanks for submitting your manuscript "The performance of genetic-enhanced DXA-BMD predicting models trained in UK biobank varies across diverse ethnic and geographical populations, PMEDICINE-D-24-00964R1” to PLOS Medicine. The paper has been reviewed by two subject experts and a statistician; their comments are included below and can also be accessed here:

[LINK]

As you will see, the reviewers were positive about the paper but, they raised a number of questions about specific study details and the methodological approach. After discussing the paper with the editorial team and an academic editor with relevant expertise, I’m pleased to invite you to revise the paper in response to the reviewers’ comments. We plan to send the revised paper to some of all of the original reviewers*, and of course we cannot provide any guarantees at this stage regarding publication.

When you upload your revision, please include a point-by-point response that addresses all of the reviewer and editorial points, indicating the changes made in the manuscript and either an excerpt of the revised text or the location (eg: page and line number) where each change can be found. Please submit a clean version of the paper as the main article file and a version with changes marked should as a marked-up manuscript. Please also check the guidelines for revised papers at http://journals.plos.org/plosmedicine/s/revising-your-manuscript for any that apply to your paper.

We ask that you submit your revision by June 7th 2024. However, if this deadline is not feasible, please contact me by email, and we can discuss a suitable alternative.

Please don’t hesitate to contact me directly with any questions (pdodd@plos.org). If you reply directly to this message, please be sure to ‘Reply All’ so your message comes directly to my inbox.

Kind regards,

Pippa

Philippa Dodd MBBS MRCP PhD

PLOS Medicine

plosmedicine.org

pdodd@plos.org

*Please note: If your article is accepted, you may have the opportunity to make the peer review history publicly available. The record will include editor decision letters (with reviews) and your responses to reviewer comments. If eligible, we will contact you to opt in or out.

-----------------------------------------------------------

Editorial comments:

1) The editorial team are in agreement that your manuscript is very interesting and well presented. Overall, opinions were somewhat split because of the limited direct clinical application. This was echoed by the Academic Editor (please see below). Please consider this point as part of your revisions and ensure to include clear discussion, not only of the potential usefulness of such a model, but the benefits and barriers to its implementation in clinical practice.

2) Data Availability

PLOS Medicine requires that the de-identified data underlying the specific results in a published article be made available, without restrictions on access, in a public repository or as Supporting Information at the time of article publication, provided it is legal and ethical to do so. Please see the policy at http://journals.plos.org/plosmedicine/s/data-availability and FAQs at

http://journals.plos.org/plosmedicine/s/data-availability#loc-faqs-for-data-policy

The Data Availability Statement (DAS) requires revision. For each data source used in your study:

a) If the data are freely or publicly available, note this and state the location of the data: within the paper, in Supporting Information files, or in a public repository (include the DOI or accession number).

b) If the data are owned by a third party but freely available upon request, please note this and state the owner of the data set and contact information for data requests (web or email address). Note that a study author cannot be the contact person for the data.

c) If the data are not freely available, please describe briefly the ethical, legal, or contractual restriction that prevents you from sharing it. Please also include an appropriate contact (web or email address) for inquiries (again, this cannot be a study author).

3) Of all authors who submit modelling studies we ask that the following points are included in the main manuscript. Please review the list below and ensure that each item is included:

* Please provide a diagram that shows the model structure, including how the disease natural history is represented, the process and determinants of disease acquisition, and how the putative intervention could affect the system.

* Please provide a complete list of model parameters, including clear and precise descriptions of [the meaning of each parameter, together with the values or ranges for each, with justification or the primary source cited, and important caveats about the use of these values noted].

* Please provide a clear statement about how the model was fitted to the data [including goodness-of-fit measure, the numerical algorithm used, which parameter varied, constraints imposed on parameter values, and starting conditions].

* For uncertainty analyses, please state the sources of uncertainties quantified and not quantified [can include parameter, data, and model structure].

* Please provide sensitivity analyses to identify which parameter values are most important in the model. Uncertainty estimates seek to derive a range of credible results on the basis of an exploration of the range of reasonable parameter values. The choice of method should be presented and justified.

* Please discuss the scientific rationale for this choice of model structure and identify points where this choice could influence conclusions drawn. Please also describe the strength of the scientific basis underlying the key model assumptions.

Items are derived from Geoffrey P Garnett, Simon Cousens, Timothy B Hallett, Richard Steketee, Neff Walker. Mathematical models in the evaluation of health programmes. (2011) Lancet DOI:10.1016/S0140-6736(10)61505-X

4) Statistical reporting

Throughout, including tables and figures, please quantify the main results with 95% CIs and p values.

When reporting p values please report as <0.001 and where higher as p=0.002, for example. If not reporting p values, for the purpose of transparent data reporting, please clearly state the reasons why not. When reporting 95% CIs please separate upper and lower bounds with commas instead of hyphens as the latter can be confused with reporting of negative values.

Please include the actual amounts and/or absolute risk(s) of relevant outcomes (including NNT or NNH where appropriate), not just relative risks or correlation coefficients. (example for absolute risks: PMID: 28399126).

5) Abstract layout

Please ensure you structure your abstract using the PLOS Medicine headings (Background, Methods and Findings, Conclusions). Please combine the Methods and Findings sections into one section, “Methods and findings”.

6) Author summary

At this stage, we ask that you include a short, non-technical Author Summary of your research to make findings accessible to a wide audience that includes both scientists and non-scientists. The authors summary should consist of 2-3 succinct bullet points under each of the following headings:

• Why Was This Study Done? Authors should reflect on what was known about the topic before the research was published and why the research was needed.

• What Did the Researchers Do and Find? Authors should briefly describe the study design that was used and the study’s major findings. Do include the headline numbers from the study, such as the sample size and key findings.

• What Do These Findings Mean? Authors should reflect on the new knowledge generated by the research and the implications for practice, research, policy, or public health. Authors should also consider how the interpretation of the study’s findings may be affected by the study limitations. In the final bullet point of ‘What Do These Findings Mean?’, please describe the main limitations of the study in non-technical language.

The Author Summary should immediately follow the Abstract in your revised manuscript. This text is subject to editorial change and should be distinct from the scientific abstract. Please see our author guidelines for more information: https://journals.plos.org/plosmedicine/s/revising-your-manuscript#loc-author-summary

7) Introduction layout

Please address past research and explain the need for and potential importance of your study. Indicate whether your study is novel and how you determined that. If there has been a systematic review of the evidence related to your study (or you have conducted one), please refer to and reference that review and indicate whether it supports the need for your study.

8) Discussion layout

Please present and organize the Discussion as follows: a short, clear summary of the article's findings; what the study adds to existing research and where and why the results may differ from previous research; strengths and limitations of the study; implications and next steps for research, clinical practice, and/or public policy; one-paragraph conclusion.

-----------------------------------------------------------

Comments from the reviewers:

Reviewer #1: The authors present a very interesting manuscript on the use of predictive modelling combining not just genetic data, but also clinical data. This exercise is highly warranted and presents a more viable and cost-effective approach to evaluating the risk of osteoporosis and fracture susceptibility.

A few comments are included below:

- How were clinical risk factors chosen for BMD and fracture risk?

- Was LS BMD available for L1-L4 or L2-L4?

- Could the low predictive modelling be due to the utilisation of GWAS data and inclusion of common variants?

- Was there any particular clinical risk factor(s) included in the analysis that was driving the association? Perhaps even SNP(s) or Indels? In the case of the latter, in which genes did these fall?

- Is there a reason why predictive modelling in the UKBB Training sets seemed to have worked better for spine compared to FN BMD?

- In the case of WGS, were variants annotated? If yes, which tool was used? Was allele balance accounted for? What was the coverage cut-off used for the variants?

Reviewer #2: "The performance of genetic-enhanced DXA-BMD predicting models trained in UK biobank varies across diverse ethnic and geographical populations" reports the outcome of developing various prediction models, towards early detection of osteoporosis in the neck and spine. The focus on genetic data is in response to DXA-based bone mineral density measurement (BMD) being relatively inaccessible, and possibly inappropriate for general screening due to proportion of negative cases. The models were developed on largely white individuals between 40-69 years from the UKBB dataset, and validated on a number of independent cohorts of varying ethnicity distributions. It was found that while prediction was improved by the inclusion of genetic data, performance on non-white ethnicity cohorts was significantly poorer than on the white demographic as used to develop the prediction models.

This paper presents justified support for the integration of suitable genetic data in prediction models for osteoporosis screening, and raises timely concerns about such models possibly not being directly applicable to demographics outside that which were used in their development. A number of issues might be considered:

1. In Line 154, it is stated that "Then, we selected the model with best performance and assessed the predictive performance in three UKBB test sets and 12 independent cohorts". It might be clarified whether this "model with best performance" is separately selected for the FNK and SPN classes, or otherwise.

2. Related to the above, the best model (from the four) might be indicated for the main results in Table 2. Moreover, the full results for all four models on both tasks might be reported if possible.

3. In Line 171, it is stated that family relationship inference was performed using KING software. The reliability of such inference might be briefly stated.

4. In Line 172, it is stated that only individuals with no relative 3rd degree or closer were retained, to ensure sample independence. It might be clarified as to what was done when pairs (or groups) of individuals with close relatives were found. Were all of them excluded, or were one of such pairs (or groups) retained?

5. In the Phenotype measurements and quality control section, the treatment of missing variables (if any; i.e. imputation, exclusion, etc.) might be briefly stated, if relevant.

6. In Line 245, the 1/1000 percentile is referred to. It might be clarified as to how this threshold was chosen. Also, it might be confirmed as to whether this is 1/1000 (i.e. 0.1th percentile), or indeed the 0.001th percentile.

7. In Line 296, it is stated that the most significant SNP was chosen within each LD window. Was there any reason to restrict each LD window to a single SNP?

8. In Line 317, a CNN model is described (also in S1 Fig). It might be clarified as to the input dimensions of the model - is it 1D or 2D, and if the latter, what would be the width & height dimensions of the input?

9. A number of minor grammatical/phrasing issues might be considered, e.g.

(Line 231) "in other ethnic population" -> "populations"

(Line 235) "The rest participants" -> "The rest of the participants"

(Line 294) "To reduce of the input features" -> "To reduce the input features"

etc.

Reviewer #3: In this study, the authors developed a predictive model for dual-energy X-ray absorptiometry (DXA)bone mineral density (BMD) at femoral neck (FNK) and lumbar spine (SPN), based on genetic and clinical factors within a Caucasian cohort from the UK Biobank (UKBB). The model was then tested across other ancestry groups within the UKBB and several independent cohorts. The findings revealed that the predictive model performed well within the UKBB Caucasian population, but there was a notable variance in performance across other populations and independent cohorts. The study demonstrated that genetic factors could improve the performance of DXA-BMD prediction beyond that of clinical factors alone. Moreover, the predicted BMDs significantly improved the prediction of fracture risk. This study provides new insights into predicting osteoporosis and fracture risk, aiding clinicians in better understanding and preventing fragile fractures, thereby improving patient care and quality of life. Additionally, testing the BMD predictive model's performance across different ethnic backgrounds underscores the importance of training models on diverse population datasets, which is scientifically substantiated for public health measures targeting prevention. I have several comments for authors to consider improve the manuscript:

1. To underscore the potential clinical and research implications of this study, it would be beneficial to expand upon the significance of the findings in the Conclusion section. This could include discussing how their findings might influence clinical practices or directions for futures tudies.

2. The authors used linear regressions to demonstrate that the impact of various clinical factors on BMD prediction differs across cohorts, even within the same ethnicity (Line 489-497). A more detailed exploration of the underlying reasons for these differences would be informative in the Discussion.

3. The study grouped all non-Caucasian ethnic populations within the UKBB into a into a single dataset (Line 229-241). For a more nuanced analysis, it would be advantageous todelineate these populations byspecific ethnicities. This would enhance the comparability of the model's performance with other ethnic groups in independent cohorts.

4. The inclusion of additional variables such as glucocorticoid use and rheumatoid arthritis status in the Cox proportional hazards regression analysis (Lines 359-362) should be justified. Clarifying the rationale for these choices would strengthen the methodological transparency of the study.

5. There are some detail problems in some figures and tables, such as a spelling mistake in Table 2: 'Average_DS' should be corrected to 'Average_GD'; and the full names of abbreviations like LR, PRS, etc., in Fig 2 should be clarified in the legend. A thorough review of all visual elements is recommended to ensure accuracy and clarity.

Any attachments provided with reviews can be seen via the following link:

[LINK]

-----------------------------------------------------------

Comments from the Academic Editor:

I agree with offering major revision. However, I understand the reservations of the editorial team. From a clinician's perspective, I am not sure these findings will have direct clinical implications. To date, there is no routine screening of genetic factors for common diseases such as osteoporosis, since the cost-effectiveness of such procedure is unknown. It would be helpful for the authors to comment in this respect.

-----------------------------------------------------------

General editorial requests:

1. Please upload any figures associated with your paper as individual TIF or EPS files with 300dpi resolution at resubmission; please read our figure guidelines for more information on our requirements: http://journals.plos.org/plosmedicine/s/figures. While revising your submission, please upload your figure files to the PACE digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email us at PLOSMedicine@plos.org.

2. Please ensure that the paper adheres to the PLOS Data Availability Policy (see http://journals.plos.org/plosmedicine/s/data-availability), which requires that all data underlying the study's findings be provided in a repository or as Supporting Information. For data residing with a third party, authors are required to provide instructions with contact information for obtaining the data. PLOS journals do not allow statements supported by "data not shown" or "unpublished results." For such statements, authors must provide supporting data or cite public sources that include it.

3. We ask every co-author listed on the manuscript to fill in a contributing author statement, making sure to declare all competing interests. If any of the co-authors have not filled in the statement, we will remind them to do so when the paper is revised. If all statements are not completed in a timely fashion this could hold up the re-review process. If new competing interests are declared later in the revision process, this may also hold up the submission. Should there be a problem getting one of your co-authors to fill in a statement we will be in contact. YOU MUST NOT ADD OR REMOVE AUTHORS UNLESS YOU HAVE ALERTED THE EDITOR HANDLING THE MANUSCRIPT TO THE CHANGE AND THEY SPECIFICALLY HAVE AGREED TO IT. You can see our competing interests policy here: http://journals.plos.org/plosmedicine/s/competing-interests.

-----------------------------------------------------------

To submit your revised manuscript please use the following link:

https://www.editorialmanager.com/pmedicine/

Your article can be found in the "Submissions Needing Revision" folder.

10.1371/journal.pmed.1004451.r003
Author response to Decision Letter 1
Submission Version2
26 Jun 2024

Attachment Submitted filename: Response.docx

10.1371/journal.pmed.1004451.r004
Decision Letter 2
Dodd Philippa C. Senior Editor
© 2024 Philippa C. Dodd
2024
Philippa C. Dodd
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Submission Version2
12 Jul 2024

Dear Dr. Liu,

Thank you very much for re-submitting your manuscript "The performance of genetic-enhanced DXA-BMD predicting models trained in UK biobank varies across diverse ethnic and geographical populations" (PMEDICINE-D-24-00964R2) for review by PLOS Medicine.

I have discussed the paper with my colleagues and the academic editor and it was also seen again by 3 reviewers. I am pleased to say that provided the remaining editorial and production issues are dealt with we are planning to accept the paper for publication in the journal.

The remaining issues that need to be addressed are listed at the end of this email. Any accompanying reviewer attachments can be seen via the link below. Please take these into account before resubmitting your manuscript:

[LINK]

***Please note while forming your response, if your article is accepted, you may have the opportunity to make the peer review history publicly available. The record will include editor decision letters (with reviews) and your responses to reviewer comments. If eligible, we will contact you to opt in or out.***

In revising the manuscript for further consideration here, please ensure you address the specific points made by each reviewer and the editors. In your rebuttal letter you should indicate your response to the reviewers' and editors' comments and the changes you have made in the manuscript. Please submit a clean version of the paper as the main article file. A version with changes marked must also be uploaded as a marked up manuscript file.

Please also check the guidelines for revised papers at http://journals.plos.org/plosmedicine/s/revising-your-manuscript for any that apply to your paper. If you haven't already, we ask that you provide a short, non-technical Author Summary of your research to make findings accessible to a wide audience that includes both scientists and non-scientists. The Author Summary should immediately follow the Abstract in your revised manuscript. This text is subject to editorial change and should be distinct from the scientific abstract.

We expect to receive your revised manuscript within 1 week. Please email us (plosmedicine@plos.org) if you have any questions or concerns.

We ask every co-author listed on the manuscript to fill in a contributing author statement. If any of the co-authors have not filled in the statement, we will remind them to do so when the paper is revised. If all statements are not completed in a timely fashion this could hold up the re-review process. Should there be a problem getting one of your co-authors to fill in a statement we will be in contact. YOU MUST NOT ADD OR REMOVE AUTHORS UNLESS YOU HAVE ALERTED THE EDITOR HANDLING THE MANUSCRIPT TO THE CHANGE AND THEY SPECIFICALLY HAVE AGREED TO IT.

Please ensure that the paper adheres to the PLOS Data Availability Policy (see http://journals.plos.org/plosmedicine/s/data-availability), which requires that all data underlying the study's findings be provided in a repository or as Supporting Information. For data residing with a third party, authors are required to provide instructions with contact information for obtaining the data. PLOS journals do not allow statements supported by "data not shown" or "unpublished results." For such statements, authors must provide supporting data or cite public sources that include it.

To enhance the reproducibility of your results, we recommend that you deposit your laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols

Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript.

Please note, when your manuscript is accepted, an uncorrected proof of your manuscript will be published online ahead of the final version, unless you've already opted out via the online submission form. If, for any reason, you do not want an earlier version of your manuscript published online or are unsure if you have already indicated as such, please let the journal staff know immediately at plosmedicine@plos.org.

If you have any questions in the meantime, please contact me or the journal staff on plosmedicine@plos.org.  

We look forward to receiving the revised manuscript by Jul 19 2024 11:59PM.   

Kind regards,

Pippa

Philippa Dodd, MBBS MRCP PhD

Senior Editor 

PLOS Medicine

plosmedicine.org

pdodd@plos.org

------------------------------------------------------------

Requests from Editors:

GENERAL

Thank you for your responses to previous editor and reviewer requests, please see below for further comments which we require you address in full.

The editorial comments pertain largely to specific content and formatting requirements. Some items may not apply and others may have already been incorporated appropriately but please review the complete list and amend as necessary.

The main manuscript is very long thus makes for rather a challenging read at times. It would benefit from being made more concise throughout, particularly the methods and the discussion sub-sections. Specific comments are detailed below.

DATA AVAILABILITY STATEMENT

Please note that a study author cannot be a contact for data inquiries. The Data Availability Statement (DAS) requires revision. In respect of the ‘raw whole genome array or sequencing (WGS) data’:

a) If the data are freely or publicly available, note this and state the location of the data: within the paper, in Supporting Information files, or in a public repository (include the DOI or accession number).

b) If the data are owned by a third party but freely available upon request, please note this and state the owner of the data set and contact information for data requests (web or email address). Note that a study author cannot be the contact person for the data.

c) If the data are not freely available, please describe briefly the ethical, legal, or contractual restriction that prevents you from sharing it. Please also include an appropriate contact (web or email address) for inquiries (again, this cannot be a study author).

TITLE

Please revise your title according to PLOS Medicine's style. Your title must be nondeclarative and not a question. It should begin with main concept if possible. "Effect of" should be used only if causality can be inferred, i.e., for an RCT. Please place the study design ("A randomized controlled trial," "A retrospective study," "A modelling study," etc.) in the subtitle (ie, after a colon).

ABSTRACT

Please structure your abstract using the PLOS Medicine headings (Background, Methods and Findings, Conclusions).

Please combine the Methods and Findings sections into one section, “Methods and findings”.

Abstract Background: Please provide context of why the study is important. The final sentence should clearly state the study question.

Abstract Background: Please provide context of why the study is important. The final sentence should clearly state the study question.

Abstract Methods and Findings:

Please ensure that all numbers presented in the abstract are present and identical to numbers presented in the main manuscript text.

Please include the study design, population and setting, number of participants, years during which the study took place, length of follow up, and main outcome measures.

Please quantify the main results with 95% CIs and p values.

Please include the important dependent variables that are adjusted for in the analyses.

Please include the actual amounts and/or absolute risk(s) of relevant outcomes (including NNT or NNH where appropriate), not just relative risks or correlation coefficients. (example for absolute risks: PMID: 28399126).

Please include a summary of adverse events if these were assessed in the study.

In the last sentence of the Abstract Methods and Findings section, please describe the main limitation(s) of the study's methodology.

Abstract Conclusions:

Please address the study implications without overreaching what can be concluded from the data; the phrase "In this study, we observed ..." may be useful.

Please interpret the study based on the results presented in the abstract, emphasizing what is new without overstating your conclusions.

Please avoid vague statements such as "these results have major implications for policy/clinical care". Mention only specific implications substantiated by the results.

Please avoid assertions of primacy ("We report for the first time....")

STATISTICAL REPORTING

Previously we asked, “Please include the actual amounts and/or absolute risk(s) of relevant outcomes (including NNT or NNH where appropriate), not just relative risks or correlation coefficients. (example for absolute risks: PMID: 28399126).” We did not see any amendments related to this comment or any specific response rebuttal. Throughout, including the abstract, please report measurements of actual risk, this is a prerequisite to publication.

AUTHOR SUMMARY

Thank you for including an author summary which requires some revision. The authors summary should consist of 2-3 succinct bullet points under each of the following headings:

• Why Was This Study Done? Authors should reflect on what was known about the topic before the research was published and why the research was needed.

• What Did the Researchers Do and Find? Authors should briefly describe the study design that was used and the study’s major findings. Do include the headline numbers from the study, such as the sample size and key findings.

• What Do These Findings Mean? Authors should reflect on the new knowledge generated by the research and the implications for practice, research, policy, or public health. Authors should also consider how the interpretation of the study’s findings may be affected by the study limitations. In the final bullet point of ‘What Do These Findings Mean?’, please describe the main limitations of the study in non-technical language.

Please amend as outlined above.

INTRODUCTION

Please conclude the Introduction with a clear description of the study question or hypothesis.

METHODS and RESULTS

Please report the number of [patients, samples, etc] and dates of recruitment, and account for all methods used in your study.

Please define "lost to follow-up" as used in this study. Other reasons for exclusion should be defined.

Please define the length of follow up (eg, in mean, SD, and range).

Please provide the actual numbers of events for the outcomes, not just summary statistics or ORs.

Please present numerators and denominators used to derive percentages.

Please ensure to indicate where analyses are adjusted and which factors are adjusted for.

As for the abstract, please ensure to quantify the main results with 95% CIs and p values.

When a p value is given, please specify the statistical test used to determine it.

The methods section is very long and it may help to improve reader accessibility if some of the details could be moved to the supporting information files.

Line 242 – Study cohorts and data preprocessing – this section is particularly long and contains information regarding datasets and statistical methods. Suggest separating this information into two different sections (databases/cohorts and statistical methods) to improve reader accessibility.

Line 529 – suggest ‘Characteristics of the study population’ as an alternative sub-heading.

TABLES and FIGURES

Please see here for guidelines on submitting and citing figures https://journals.plos.org/plosmedicine/s/figures#loc-how-to-submit-figures-and-captions

Please provide titles and legends for all tables and figures (including those in Supporting Information files).

Please ensure that each table and figure is affiliated to a caption which clearly describes the figure content without the need to refer to the text.

Please ensure that all abbreviations are defined in the caption or an appropriate footnote, including those used to report statistical information.

Please ensure to include the meaning of any dots/lines/bars.

Please consider avoiding the use of red and green in order to make figures more accessible to those with colour blindness.

To help facilitate transparent data reporting, where adjusted analyses are presented please also present the unadjusted analyses for comparison. In a caption or footnote please detail all factors adjusted for.

DISCUSSION

Previously we asked, ‘Please present and organize the Discussion as follows: a short, clear summary of the article's findings; what the study adds to existing research and where and why the results may differ from previous research; strengths and limitations of the study; implications and next steps for research, clinical practice, and/or public policy; one-paragraph conclusion.’

Reading through, the discussion is very long and it is easy to get lost. Please revise for brevity and improve accessibility ensuring that the above structure is followed and the items easily identifiable. Pleas avoid the use of subheadings.

REFERENCES

For in-text reference callouts please place citations in square parentheses separate by commas. For example, [1,3,6] or [1-3]. Please check and amend throughout all sub-sections of the manuscript and supporting files.

In the bibliography please ensure that you list up to but no more than 6 author names followed by et al.

For all web references please ensure you include an, ‘Accessed [date].’

Journal name abbreviations should be those listed in the National Center for Biotechnology Information (NCBI) databases.

Line 881 – please remove the data availability statement from the main manuscript and include only in the manuscript submission form when you resubmit the manuscript. It will be compiled as metadata at the time of publication.

SUPPORTING INFORMATION

In the published article, supporting information files are accessed only through a hyperlink attached to the captions. For this reason, you must list captions at the end of your manuscript file. You may include a caption within the supporting information file itself, as long as that caption is also provided in the manuscript file. Do not submit a separate caption file.

As the supporting information files are contained with a single file:

Please label the file as ‘S1 Supporting Information’.

Please apply alphabetical labelling to each table and figure contained within the S1 file. For example, ‘Fig A’ to ‘Fig Z’ and ‘Table A’ to ‘Table Z’.

Plain text does not need to be labelled and can just be given a title as necessary. For example, ‘Statistical Analysis Plan’.

Please cite tables/figures as ‘Fig A in S1 Supporting Information’ and/or ‘Table A in S1 Supporting Information’, for example.

Please cite plain text as, ‘Statistical Analysis Plan in S1 Supporting Information’, for example.

Alternatively, you may upload each file in the supporting information as individual files and label tables as ‘S1 Table’ (so on) and figures as ‘S1 Fig’ (and so on).

Any additional documents (protocols or appendices) can be labelled as ‘S1 Appendix’, for example.

Please cite items as exactly as labelled.

Further guidance can be found here https://journals.plos.org/plosmedicine/s/supporting-information

SOCIAL MEDIA

To help us extend the reach of your research, please detail any X (formerly Twitter) handles you wish to be included when we tweet this paper (including your own, your coauthors’, your institution, funder, or lab) in the manuscript submission form when you re-submit the manuscript.

Comments from Reviewers:

Reviewer #1: Comments have been addressed

Perhaps some additional information can be provided regarding the variant rs554533790 (which genes it is close to) - if at all possible.

Reviewer #2: We thank the authors for largely addressing our previous concerns. On the point where one individual from inferred kinship relationships being retained, it might be detailed as to how the retained individual was selected - was it essentially random (i.e. first encountered individual retained, then any of their relations excluded)?

Reviewer #3: The authors have addressed all of my comments. Thanks!

Any attachments provided with reviews can be seen via the following link:

[LINK]

10.1371/journal.pmed.1004451.r005
Author response to Decision Letter 2
Submission Version3
20 Jul 2024

Attachment Submitted filename: Response_R2.docx

10.1371/journal.pmed.1004451.r006
Decision Letter 3
Dodd Philippa C. Senior Editor
© 2024 Philippa C. Dodd
2024
Philippa C. Dodd
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Submission Version3
23 Jul 2024

Dear Dr Liu, 

On behalf of my colleagues and the Academic Editor, Professor Christelle Nguyen, I am pleased to inform you that we have agreed to publish your manuscript "Variability in performance of genetic-enhanced DXA-BMD prediction models trained on the UK Biobank across ethnic and geographical populations: A genetic risk prediction study" (PMEDICINE-D-24-00964R3) in PLOS Medicine.

Prior to publication, when completing the required formatting changes (detailed below), please make the following amendments:

1) Title – please revise the title to read as follows: “Variability in performance of genetic-enhanced DXA-BMD prediction models across diverse ethnic and geographic populations: A risk prediction study”.

2) Author summary – line 85 – please define ‘BMD’ and ‘(DXA)’.

3) Tables – please include the full names of the cohorts either in column 1 or defined in the footnote below.

Before your manuscript can be formally accepted you will need to complete some formatting changes, which you will receive in a follow up email. Please be aware that it may take several days for you to receive this email; during this time no action is required by you. Once you have received these formatting requests, please note that your manuscript will not be scheduled for publication until you have made the required changes.

In the meantime, please log into Editorial Manager at http://www.editorialmanager.com/pmedicine/, click the "Update My Information" link at the top of the page, and update your user information to ensure an efficient production process. 

PRESS

We frequently collaborate with press offices. If your institution or institutions have a press office, please notify them about your upcoming paper at this point, to enable them to help maximise its impact. If the press office is planning to promote your findings, we would be grateful if they could coordinate with medicinepress@plos.org. If you have not yet opted out of the early version process, we ask that you notify us immediately of any press plans so that we may do so on your behalf.

We also ask that you take this opportunity to read our Embargo Policy regarding the discussion, promotion and media coverage of work that is yet to be published by PLOS. As your manuscript is not yet published, it is bound by the conditions of our Embargo Policy. Please be aware that this policy is in place both to ensure that any press coverage of your article is fully substantiated and to provide a direct link between such coverage and the published work. For full details of our Embargo Policy, please visit http://www.plos.org/about/media-inquiries/embargo-policy/.

To enhance the reproducibility of your results, we recommend that you deposit your laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols

Thank you again for submitting to PLOS Medicine, it has been a pleasure handling your manuscript. We look forward to publishing your paper. 

Kind regards,

Pippa 

Philippa Dodd, MBBS MRCP PhD 

Senior Editor 

PLOS Medicine

pdodd@plos.org
==== Refs
References

1 Kanis JA . Diagnosis of osteoporosis and assessment of fracture risk. Lancet. 2002;359 (9321 ):1929–1936. doi: 10.1016/S0140-6736(02)08761-5 12057569
2 Kanis JA , Oden A , Johnell O , Jonsson B , de Laet C , Dawson A . The burden of osteoporotic fractures: a method for setting intervention thresholds. Osteoporos Int. 2001;12 (5 ):417–427. doi: 10.1007/s001980170112 11444092
3 Rachner TD , Khosla S , Hofbauer LC . Osteoporosis: now and the future. Lancet. 2011;377 (9773 ):1276–1287. doi: 10.1016/S0140-6736(10)62349-5 21450337
4 Gu Q , Koenig L , Mather RC 3rd , Tongue J . Surgery for hip fracture yields societal benefits that exceed the direct medical costs. Clin Orthop Relat Res. 2014;472 (11 ):3536–3546. doi: 10.1007/s11999-014-3820-6 25091223
5 Nih Consensus Development Panel on Osteoporosis Prevention D, Therapy. Osteoporosis prevention, diagnosis, and therapy. JAMA. 2001;285 (6 ):785–795. doi: 10.1001/jama.285.6.785 11176917
6 Ensrud KE , Crandall CJ . Osteoporosis. Ann Intern Med. 2017;167 (3 ):ITC17–ITC32. doi: 10.7326/AITC201708010 28761958
7 Choksi P , Jepsen KJ , Clines GA . The challenges of diagnosing osteoporosis and the limitations of currently available tools. Clin Diabetes Endocrinol. 2018;4 :12. doi: 10.1186/s40842-018-0062-7 29862042
8 Compston J , Cooper A , Cooper C , Gittoes N , Gregson C , Harvey N , et al . UK clinical guideline for the prevention and treatment of osteoporosis. Arch Osteoporos. 2017;12 (1 ):43. doi: 10.1007/s11657-017-0324-5 28425085
9 Turner DA , Khioe RFS , Shepstone L , Lenaghan E , Cooper C , Gittoes N , et al . The Cost-Effectiveness of Screening in the Community to Reduce Osteoporotic Fractures in Older Women in the UK: Economic Evaluation of the SCOOP Study. J Bone Miner Res. 2018;33 (5 ):845–851. doi: 10.1002/jbmr.3381 29470854
10 Kelsey JL . Risk factors for osteoporosis and associated fractures. Public Health Rep. 1989;104 Suppl (Suppl ):14–20. 2517695
11 Yang TL , Shen H , Liu A , Dong SS , Zhang L , Deng FY , et al . A road map for understanding molecular and genetic determinants of osteoporosis. Nat Rev Endocrinol. 2020;16 (2 ):91–103. doi: 10.1038/s41574-019-0282-7 31792439
12 Arden NK , Baker J , Hogg C , Baan K , Spector TD . The heritability of bone mineral density, ultrasound of the calcaneus and hip axis length: a study of postmenopausal twins. J Bone Miner Res. 1996;11 (4 ):530–534. doi: 10.1002/jbmr.5650110414 8992884
13 Duncan EL , Brown MA . Clinical review 2: Genetic determinants of bone density and fracture risk—state of the art and future directions. J Clin Endocrinol Metab. 2010;95 (6 ):2576–87. doi: 10.1210/jc.2009-2406 20375209
14 Lewis CM , Vassos E . Polygenic risk scores: from research tools to clinical instruments. Genome Med. 2020;12 (1 ):44. doi: 10.1186/s13073-020-00742-5 32423490
15 Lee SH , Lee SW , Ahn SH , Kim T , Lim KH , Kim BJ , et al . Multiple gene polymorphisms can improve prediction of nonvertebral fracture in postmenopausal women. J Bone Miner Res. 2013;28 (10 ):2156–2164. doi: 10.1002/jbmr.1955 23572424
16 Ho-Le TP , Center JR , Eisman JA , Nguyen HT , Nguyen TV . Prediction of Bone Mineral Density and Fragility Fracture by Genetic Profiling. J Bone Miner Res. 2017;32 (2 ):285–293. doi: 10.1002/jbmr.2998 27649491
17 Lu T , Forgetta V , Keller-Baruch J , Nethander M , Bennett D , Forest M , et al . Improved prediction of fracture risk leveraging a genome-wide polygenic risk score. Genome Med. 2021;13 (1 ):16. doi: 10.1186/s13073-021-00838-6 33536041
18 Gonnelli S , Cepollaro C , Gennari L , Montagnani A , Caffarelli C , Merlotti D , et al . Quantitative ultrasound and dual-energy X-ray absorptiometry in the prediction of fragility fracture in men. Osteoporos Int. 2005;16 (8 ):963–968. doi: 10.1007/s00198-004-1771-6 15599495
19 Hsu Y-H , Xu X , Jeong S . Genetic Determinants and Pharmacogenetics of Osteoporosis and Osteoporotic Fracture. In: Leder BZ , Wein MN , editors. Osteoporosis: Pathophysiology and Clinical Management. Cham: Springer International Publishing; 2020. p. 485–506.
20 Sud A , Horton RH , Hingorani AD , Tzoulaki I , Turnbull C , Houlston RS , et al . Realistic expectations are key to realising the benefits of polygenic scores. BMJ. 2023;380 :e073149. doi: 10.1136/bmj-2022-073149 36854461
21 Schwarzerova J , Hurta M , Barton V , Lexa M , Walther D , Provaznik V , et al . A perspective on genetic and polygenic risk scores-advances and limitations and overview of associated tools. Brief Bioinform. 2024;25 (3 ). doi: 10.1093/bib/bbae240 38770718
22 Kanis JA , Borgstrom F , De Laet C , Johansson H , Johnell O , Jonsson B , et al . Assessment of fracture risk. Osteoporos Int. 2005;16 (6 ):581–589. doi: 10.1007/s00198-004-1780-5 15616758
23 Bycroft C , Freeman C , Petkova D , Band G , Elliott LT , Sharp K , et al . The UK Biobank resource with deep phenotyping and genomic data. Nature. 2018;562 (7726 ):203–209. doi: 10.1038/s41586-018-0579-z 30305743
24 Greenbaum J , Su KJ , Zhang X , Liu Y , Liu A , Zhao LJ , et al . A multiethnic whole genome sequencing study to identify novel loci for bone mineral density. Hum Mol Genet. 2022;31 (7 ):1067–1081. doi: 10.1093/hmg/ddab305 34673960
25 Orwoll E , Blank JB , Barrett-Connor E , Cauley J , Cummings S , Ensrud K , et al . Design and baseline characteristics of the osteoporotic fractures in men (MrOS) study—a large observational study of the determinants of fracture in older men. Contemp Clin Trials. 2005;26 (5 ):569–85. doi: 10.1016/j.cct.2005.05.006 16084776
26 Design of the Women’s Health Initiative clinical trial and observational study. The Women’s Health Initiative Study Group. Control Clin Trials. 1998;19 (1 ):61–109.9492970
27 Fried LP , Borhani NO , Enright P , Furberg CD , Gardin JM , Kronmal RA , et al . The Cardiovascular Health Study: design and rationale. Ann Epidemiol. 1991;1 (3 ):263–276. doi: 10.1016/1047-2797(91)90005-w 1669507
28 Sudlow C , Gallacher J , Allen N , Beral V , Burton P , Danesh J , et al . UK biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS Med. 2015;12 (3 ):e1001779. doi: 10.1371/journal.pmed.1001779 25826379
29 Zhang L , Choi HJ , Estrada K , Leo PJ , Li J , Pei YF , et al . Multistage genome-wide association meta-analyses identified two new loci for bone mineral density. Hum Mol Genet. 2014;23 (7 ):1923–1933. doi: 10.1093/hmg/ddt575 24249740
30 Manichaikul A , Mychaleckyj JC , Rich SS , Daly K , Sale M , Chen WM . Robust relationship inference in genome-wide association studies. Bioinformatics. 2010;26 (22 ):2867–2873. doi: 10.1093/bioinformatics/btq559 20926424
31 McCarthy S , Das S , Kretzschmar W , Delaneau O , Wood AR , Teumer A , et al . A reference panel of 64,976 haplotypes for genotype imputation. Nat Genet. 2016;48 (10 ):1279–1283. doi: 10.1038/ng.3643 27548312
32 Consortium UK , Walter K , Min JL , Huang J , Crooks L , Memari Y , et al . The UK10K project identifies rare variants in health and disease. Nature. 2015;526 (7571 ):82–90. doi: 10.1038/nature14962 26367797
33 Genomes Project C , Abecasis GR , Auton A , Brooks LD , DePristo MA , Durbin RM , et al . An integrated map of genetic variation from 1,092 human genomes. Nature. 2012;491 (7422 ):56–65. doi: 10.1038/nature11632 23128226
34 Das S , Forer L , Schonherr S , Sidore C , Locke AE , Kwong A , et al . Next-generation genotype imputation service and methods. Nat Genet. 2016;48 (10 ):1284–1287. doi: 10.1038/ng.3656 27571263
35 Ding Y , Hou K , Xu Z , Pimplaskar A , Petter E , Boulier K , et al . Polygenic scoring accuracy varies across the genetic ancestry continuum. Nature. 2023;618 (7966 ):774–781. doi: 10.1038/s41586-023-06079-4 37198491
36 Abraham G , Qiu Y , Inouye M . FlashPCA2: principal component analysis of Biobank-scale genotype datasets. Bioinformatics. 2017;33 (17 ):2776–2778. doi: 10.1093/bioinformatics/btx299 28475694
37 Loh PR , Kichaev G , Gazal S , Schoech AP , Price AL . Mixed-model association for biobank-scale datasets. Nat Genet. 2018;50 (7 ):906–908. doi: 10.1038/s41588-018-0144-6 29892013
38 Choi SW , O’Reilly PF . PRSice-2: Polygenic Risk Score software for biobank-scale data. Gigascience. 2019;8 (7 ). doi: 10.1093/gigascience/giz082 31307061
39 Clark K , Leung YY , Lee WP , Voight B , Wang LS . Polygenic Risk Scores in Alzheimer’s Disease Genetics: Methodology, Applications, Inclusion, and Diversity. J Alzheimers Dis. 2022;89 (1 ):1–12. doi: 10.3233/JAD-220025 35848019
40 SEAL HL . Studies in the History of Probability and Statistics. XV The historical development of the Gauss linear model. Biometrika. 1967;54 (1–2 ):1–24.4860564
41 Tibshirani R. Regression Shrinkage and Selection via The Lasso: A Retrospective. J R Stat Soc Series B Stat Methodol. 2011;73 (3 ):273–282.
42 Alzubaidi L , Zhang J , Humaidi AJ , Al-Dujaili A , Duan Y , Al-Shamma O , et al . Review of deep learning: concepts, CNN architectures, challenges, applications, future directions. J Big Data. 2021;8 (1 ):53. doi: 10.1186/s40537-021-00444-8 33816053
43 Ding B , Qian H , Zhou J , editors. Activation functions and their characteristics in deep neural networks. 2018 Chinese Control And Decision Conference (CCDC); 2018 9–11 June 2018 .
44 Mean Squared Error. The Concise Encyclopedia of Statistics. New York, NY: Springer New York; 2008. p. 337–9.
45 Ruder S. An overview of gradient descent optimization algorithms. ArXiv. 2016;abs/1609.04747.
46 Kingma D , Ba J . Adam: A Method for Stochastic Optimization. Computer Science. 2014.
47 Yu H . Bootstrap. In: Liu L , Özsu MT , editors. Encyclopedia of Database Systems. New York, NY: Springer New York; 2018. p. 335–6.
48 Wysham KD , Baker JF , Shoback DM . Osteoporosis and fractures in rheumatoid arthritis. Curr Opin Rheumatol. 2021;33 (3 ):270–276. doi: 10.1097/BOR.0000000000000789 33651725
49 Buttgereit F , Palmowski A , Bond M , Adami G , Dejaco C . Osteoporosis and fracture risk are multifactorial in patients with inflammatory rheumatic diseases. Nat Rev Rheumatol. 2024. doi: 10.1038/s41584-024-01120-w 38831028
50 Kanis JA , Johnell O , Oden A , Johansson H , McCloskey E . FRAX and the assessment of fracture probability in men and women from the UK. Osteoporos Int. 2008;19 (4 ):385–397. doi: 10.1007/s00198-007-0543-5 18292978
51 Du H , Li L , Bennett D , Guo Y , Turnbull I , Yang L , et al . Fresh fruit consumption in relation to incident diabetes and diabetic vascular complications: A 7-y prospective study of 0.5 million Chinese adults. PLoS Med. 2017;14 (4 ):e1002279. doi: 10.1371/journal.pmed.1002279 28399126
52 Krainc T , Fuentes A . Genetic ancestry in precision medicine is reshaping the race debate. Proc Natl Acad Sci U S A. 2022;119 (12 ):e2203033119. doi: 10.1073/pnas.2203033119 35294278
53 Mathieson I , Scally A . What is ancestry? PLoS Genet. 2020;16 (3 ):e1008624. doi: 10.1371/journal.pgen.1008624 32150538
54 Kemp JP , Medina-Gomez C , Estrada K , St Pourcain B , Heppe DH , Warrington NM , et al . Phenotypic dissection of bone mineral density reveals skeletal site specificity and facilitates the identification of novel loci in the genetic regulation of bone mass attainment. PLoS Genet. 2014;10 (6 ):e1004423. doi: 10.1371/journal.pgen.1004423 24945404
55 Piroska M , Tarnoki DL , Szabo H , Jokkel Z , Meszaros S , Horvath C , et al . Strong Genetic Effects on Bone Mineral Density in Multiple Locations with Two Different Techniques: Results from a Cross-Sectional Twin Study. Medicina (Kaunas). 2021;57 (3 ).
56 Videman T , Levalahti E , Battie MC , Simonen R , Vanninen E , Kaprio J . Heritability of BMD of femoral neck and lumbar spine: a multivariate twin study of Finnish men. J Bone Miner Res. 2007;22 (9 ):1455–1462. doi: 10.1359/jbmr.070606 17547536
57 Vuillemin A , Guillemin F , Jouanny P , Denis G , Jeandel C . Differential influence of physical activity on lumbar spine and femoral neck bone mineral density in the elderly population. J Gerontol A Biol Sci Med Sci. 2001;56 (6 ):B248–B253. doi: 10.1093/gerona/56.6.b248 11382786
58 Martin AR , Kanai M , Kamatani Y , Okada Y , Neale BM , Daly MJ . Clinical use of current polygenic risk scores may exacerbate health disparities. Nat Genet. 2019;51 (4 ):584–591. doi: 10.1038/s41588-019-0379-x 30926966
59 Gola D , Erdmann J , Lall K , Magi R , Muller-Myhsok B , Schunkert H , et al . Population Bias in Polygenic Risk Prediction Models for Coronary Artery Disease. Circ Genom Precis Med. 2020;13 (6 ):e002932. doi: 10.1161/CIRCGEN.120.002932 33170024
60 Costantini A , Skarp S , Kampe A , Makitie RE , Pettersson M , Mannikko M , et al . Rare Copy Number Variants in Array-Based Comparative Genomic Hybridization in Early-Onset Skeletal Fragility. Front Endocrinol (Lausanne). 2018;9 :380. doi: 10.3389/fendo.2018.00380 30042735
61 Kinning E , McMillan M , Shepherd S , Helfrich M , Hof RV , Adams C , et al . An Unbalanced Rearrangement of Chromosomes 4:20 is Associated with Childhood Osteoporosis and Reduced Caspase-3 Levels. J Pediatr Genet. 2016;5 (3 ):167–173. doi: 10.1055/s-0036-1584359 27617159
62 Aghajanian P , Hall S , Wongworawat MD , Mohan S . The Roles and Mechanisms of Actions of Vitamin C in Bone: New Developments. J Bone Miner Res. 2015;30 (11 ):1945–1955. doi: 10.1002/jbmr.2709 26358868
63 Marston NA , Garfinkel AC , Kamanu FK , Melloni GM , Roselli C , Jarolim P , et al . A polygenic risk score predicts atrial fibrillation in cardiovascular disease. Eur Heart J. 2023;44 (3 ):221–231. doi: 10.1093/eurheartj/ehac460 35980763
64 Urbut SM , Yeung MW , Khurshid S , Cho SMJ , Schuermans A , German J , et al . MSGene: a multistate model using genetic risk and the electronic health record applied to lifetime risk of coronary artery disease. Nat Commun. 2024;15 (1 ):4884. doi: 10.1038/s41467-024-49296-9 38849421
65 King A , Wu L , Deng HW , Shen H , Wu C . Polygenic risk score improves the accuracy of a clinical risk score for coronary artery disease. BMC Med. 2022;20 (1 ):385. doi: 10.1186/s12916-022-02583-y 36336692
66 Mandla R , Schroeder P , Porneala B , Florez JC , Meigs JB , Mercader JM , et al . Polygenic scores for longitudinal prediction of incident type 2 diabetes in an ancestrally and medically diverse primary care physician network: a patient cohort study. Genome Med. 2024;16 (1 ):63. doi: 10.1186/s13073-024-01337-0 38671457
67 Dannhauser FC , Taylor LC , Tung JSL , Usher-Smith JA . The acceptability and clinical impact of using polygenic scores for risk-estimation of common cancers in primary care: a systematic review. J Community Genet. 2024. doi: 10.1007/s12687-024-00709-8 38769249
