
==== Front
eClinicalMedicine
EClinicalMedicine
eClinicalMedicine
2589-5370
Elsevier

S2589-5370(24)00375-4
10.1016/j.eclinm.2024.102796
102796
Articles
Point-based risk score for the risk stratification and prediction of hepatocellular carcinoma: a population-based random survival forest modeling study
Liu Zhenqiu ab
Yuan Huangbo ab
Suo Chen bcd
Zhao Renjia ab
Jin Li ab
Zhang Xuehong efh
Zhang Tiejun tjzhang@shmu.edu.cn
bcdh∗∗
Chen Xingdong xingdongchen@fudan.edu.cn
abgh∗
a State Key Laboratory of Genetic Engineering, Human Phenome Institute, and School of Life Sciences, Fudan University, Shanghai, China
b Fudan University Taizhou Institute of Health Sciences, Taizhou, China
c Key Laboratory of Public Health Safety, Fudan University, Ministry of Education, Shanghai, China
d Department of Epidemiology, School of Public Health, Fudan University, Shanghai, China
e Channing Division of Network Medicine, Department of Medicine, Brigham and Women's Hospital and Harvard Medical School, Boston, MA, USA
f Yale University School of Nursing, Orange, CT, USA
g National Clinical Research Center for Aging and Medicine, Huashan Hospital, Fudan University, China
∗ Corresponding author. State Key Laboratory of Genetic Engineering, Human Phenome Institute, and School of Life Sciences, Fudan University, Shanghai, China. xingdongchen@fudan.edu.cn
∗∗ Corresponding author. Department of Epidemiology, School of Public Health, Fudan University, Shanghai, China. tjzhang@shmu.edu.cn
h Equally contributed senior authors.

22 8 2024
9 2024
22 8 2024
75 1027963 6 2024
3 8 2024
6 8 2024
© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Summary

Background

The precise associations between common clinical biomarkers and hepatocellular carcinoma (HCC) risk remain unclear but hold valuable insights for HCC risk stratification and prediction.

Methods

We examined the linear and nonlinear associations between the baseline levels of 32 circulating biomarkers and HCC risk in the England cohort of UK Biobank (UKBB) (n = 397,702). The participants were enrolled between 2006 and 2010 and followed up to 31st October 2022. The primary outcome is incident HCC cases. We then employed random survival forests (RSF) to select the top ten most informative biomarkers, considering their association with HCC, and developed a point-based risk score to predict HCC. The performance of the risk score was evaluated in three validation sets including UKBB Scotland and Wales cohort (n = 52,721), UKBB non-White-British cohort (n = 29,315), and the Taizhou Longitudinal Study in China (n = 17,269).

Findings

Twenty-five biomarkers were significantly associated with HCC risk, either linearly or nonlinearly. Based on the RSF model selected biomarkers, our point-based risk score showed a concordance index of 0.866 in the England cohort and varied between 0.814 and 0.849 in the three validation sets. HCC incidence rates ranged from 0.95 to 30.82 per 100,000 from the lowest to the highest quintiles of the risk score in the England cohort. Individuals in the highest risk quintile had a 32–73 times greater risk of HCC compared to those in the lowest quintile. Moreover, over 70% of HCC cases were detected in individuals within the top risk score quintile across all cohorts.

Interpretation

Our simple risk score enables the identification of high-risk individuals of HCC in the general population. However, including some biomarkers, such as insulin-like growth factor 1, not routinely measured in clinical practice may increase the model's complexity, highlighting the need for more accessible biomarkers that can maintain or improve the predictive accuracy of the risk score.

Funding

This work was supported by the 10.13039/501100001809 National Natural Science Foundation of China (grant numbers: 82204125 ) and the Science and Technology Support Program of Taizhou (TS202224 ).

Keywords

HCC
Common clinical biomarkers
Cohort study
Nonlinear correlation
Point-based scoring system
==== Body
pmc Research in context

Evidence before this study

We searched PubMed on 1st May 2024, using the following key terms “hepatocellular carcinoma OR HCC”, “biomarkers OR risk markers”, and “risk prediction OR prediction model OR risk stratification”. Previous epidemiological studies have reported associations of HCC (or liver cancer) with both hepatic and some extrahepatic biomarkers. However, some inconsistent findings remained. There were several prediction models for HCC derived from patient-based cohort such as the aMAP score. These clinical risk scores were initially developed for patients with chronic liver diseases rather than for the general population.

Added value of this study

We conducted an extensive examination of associations between the baseline levels of 32 circulating biomarkers and HCC risk. By providing a comprehensive assessment of the association between 32 clinical common biomarkers and HCC, incorporating both linear and nonlinear correlations, our research contributes to a deeper understanding of HCC risk factors. Moreover, based on the observed associations, we developed a simple point-based risk score for HCC risk stratification and prediction.

Implications of all the available evidence

The development of a simple point-based risk score, as opposed to a complex model, holds promise for HCC risk stratification in the general population. This approach extends beyond patients with chronic liver disease, enabling timely interventions and improved patient outcomes. However, the inclusion of some biomarkers not routinely measured in daily clinical practice may to some extent increase the complexity of our model, underscoring the need to identify alternative, more accessible biomarkers that can maintain or improve the predictive accuracy of the risk score.

Introduction

Hepatocellular carcinoma (HCC) accounts for over 85% of primary liver cancer cases.1 Observational studies have reported a number of associations between circulating biomarkers and HCC risk,2,3 suggesting that HCC is not only associated with hepatic biomarkers but also with non-hepatic biomarkers, such as metabolic and inflammatory factors. However, the observed associations were sometimes inconsistent in directions.4, 5, 6 This inconsistency may be due to incomplete adjustment for confounders and a small sample size. Circulating levels of most biomarkers are also subject to diurnal variation and are easily influenced by many factors, such as alcohol drinking.7 Estimating associations based on single measurements of biomarkers may therefore be high in variance. Moreover, previous studies have predominantly focused on linear relationships between biomarkers and HCC, with less consideration given to potential nonlinear correlations.8, 9, 10 Therefore, a comprehensive analysis of the associations between clinical biomarkers and HCC based on a large sample size and considering both linear and nonlinear patterns is necessary.

A dissection for the associations between common clinical biomarkers and HCC may also be informative for the risk stratification and prediction of HCC.11 To date, several composite biomarkers and scores have been developed for the prediction and early detection of HCC.11,12 For instance, the GALAD score, which integrates sex, age, and serum levels of α-fetoprotein (AFP), AFP isoform L3 (AFP-L3), and des-gamma-carboxy prothrombin, is used for the early detection of HCC in patients with nonalcoholic steatohepatitis.13 Similarly, the aMAP score that derives from a regression model and involves age, sex, albumin-bilirubin and platelets has a high potential in predicting HCC development in patients with chronic viral hepatitis.14,15 However, these clinical risk models were initially developed for patients with chronic liver diseases rather than for the general population,16,17 thus partly limiting their utility in community where the liver disease status are often unknown. Moreover, an elaborate model with excellent predictive accuracy may have a low positive predictive value, indicating a limited efficiency for detecting early disease.18 Thus, risk stratification is a more feasible target than risk prediction for disease with very low incidence rate like HCC in the general population.

In the current study, using data from the UK Biobank (UKBB) cohort, we examined the linear and nonlinear associations of 32 clinical biomarkers with the risk of HCC. We then developed and validated a point-based risk score for HCC based on these clinical biomarkers. Our findings reveal novel associations between common clinical biomarkers and HCC and are conducive to assessing HCC risk using these biomarkers in clinical and public health practices.

Methods

Study design and participant

The study schema was shown in Fig. 1. Briefly, we firstly examined the linear and nonlinear associations between circulating biomarkers and HCC in the England cohort of the UKBB. In this analysis, we excluded participants who lacked blood biochemistry data (n = 1567) and those with a history of cancer and any type of liver diseases that identified by International Classification of Disease (ICD) codes (n = 20,230; Table S1). A total of 397,702 participants were finally included for this analysis. We then developed a point-based scoring system for HCC predication based on the clinical biomarkers in the England cohort. The point-based risk score was validated in the Scotland and Wales cohort (n = 52,721). We also validated the risk score in other two cohorts with different ethnicities (i.e., the UKBB non-White-British cohort [n = 29,315] and a sub-cohort of the Taizhou Longitudinal Study (TZL)19 that containing 17,269 Chinese people).Fig. 1 The study flowchart.

Clinical biomarkers

UKBB embarked on a project to measure a wide range of blood biochemical markers (n = 31) in biological samples collected at baseline in nearly all participants and also in the samples provided by approximately 20,000 participants who returned for a repeat assessment in 2012–2013.20 The analysis methods that applied to measure the circulating levels of biomarkers and the quality control process have been previously described in detail elsewhere.20 In this study, we aimed to develop a point-based risk score for HCC risk stratification using biochemical markers from the UKBB. A total of 31 biomarkers were available, of which 28, measured in more than 75% of participants, were included. Additionally, we calculated three biomarkers based on the available data (estimated glomerular filtration rate [eGFR], aspartate aminotransferase to alanine aminotransferase ratio [AST2ALT], and non-albumin protein). The eGFR measurement is an indicator of renal function and was calculated by the CKD-EPI equation in this study.21 We defined non-albumin protein levels as the difference between the levels of total protein and albumin. Additionally, we have also considered platelet count (PLT) due to its established role as a surrogate marker of liver fibrosis and a predictor of HCC.22 In total, 32 biomarkers were examined for their associations with HCC risk. The levels of biomarkers measured at repeat assessment visit were also collected and were used for correcting regression dilution bias. The proportion of missing values ranged from 4.8% to 25.0% across the biomarkers (Figure S1). For missing value of the biomarkers, a random forest method was applied for imputation. The random forest imputation was performed using the following parameters: 50 trees (estimators), with a maximum depth of trees set to ensure that each leaf contains no fewer than 5 observations. Additionally, the number of variables tried at each split was set to the square root of the total number of variables.

Covariates

We retrieved age at recruitment, sex, average household income, body-mass index (BMI), education deprivation score, smoking status, alcohol drinking frequency, and physical activity level. Smoking status was classified as never, former, and current smokers. Average annual household income was categorized into five groups: <£18,000, £18,000–30,999, £31,000–51,999, £52,000–100,000, and >£100,000. We used the number of days per week of moderate physical activity that lasted for ≥10 min as a surrogate to measure the physical activity level. Three groups were formed according to physical activity level: 0–1 day/week, 2–4 days/week, and 5–7 days/week. Alcohol drinking frequency was classified into three groups: never or special occasions only, ≤ twice a week, and three times or more a week. The missing proportions of these covariates were <0.5%, and therefore, no imputation was performed.

Outcome

Incident cancer cases and cancer cases recorded first in death certificates within the UKBB cohort were identified through linkage to national cancer and death registries. We used the ICD-10 code of C22.0 to identify incident HCC (Table S1). At the first diagnosis of any cancer, patients were censored for HCC.

Statistics

Linear association assessment

The levels of all biomarkers were log-transformed to approximate a normal distribution and then modelled on continuous scale (per 1-standard deviation [SD] increment) in the Cox proportional hazard (CPH) models. Age at recruitment, sex, average household income, BMI, education deprivation score, smoking status, alcohol drinking frequency, and physical activity level were adjusted in the CPH model. All covariates except age, BMI and education deprivation score were entered into the CPH models as categorical variables. For fasting glucose, glycosylated hemoglobin (HbA1c), low-density lipoprotein (LDL), apolipoprotein B (ApoB), and total cholesterol, we further adjusted for the medication information (i.e., whether regularly taking insulin or lipid-lowering medicine) in the CPH models. Hazard ratios (HRs) and 95% confidence intervals (CIs) were estimated to quantify these associations. We used the proportional hazards test to determine the time-varying effects of each covariate and biomarker. Additionally, we incorporated restricted cubic splines (RCS) in the model to evaluate the nonlinear relationships between continuous covariates (age, BMI, and education deprivation score) and HCC risk. Although no time-varying effects were detected for any covariates, we identified a significant nonlinear association between BMI and HCC risk (P for nonlinearity <0.0001). To address this, we included RCS with four knots for BMI in the CPH model. We also retrieved the infection status of hepatitis B and C viruses (HBV/HCV) and assessed whether the circulating levels of biomarkers varied by HBV/HCV.

Correction for regression dilution bias

Owing to the combined effects of measurement error and within-person variability, single baseline measurements of biomarkers do not mirror the “usual” medium term levels, leading to underestimation of the real associations of disease rates with the “usual” levels of such risk factors.23 HRs were therefore corrected for this “regression dilution bias” by dividing the log HR associated with baseline levels (and its standard error) by ρ, the correlation coefficient of the biomarker values measured in individuals at different times (Table S2).24

Nonlinear association assessment

We used RCS with 4 knots at 5%, 35%, 65%, and 95% centiles, as suggested by Harrell FE,25 to flexibly model the association between clinical biomarkers and HCC. We tested for potential non-linearity by using a likelihood ratio test comparing the model with only a linear term against the model with linear and cubic spline terms. For biomarker showing significant nonlinear association with HCC, we also assessed its association with HCC using the CPH models in subgroups that stratified by median value of the biomarker.

Random survival forest modelling to select informative biomarkers

We employed random survival forest model (RSF, number of trees = 500, number of variables to split at each node = 6) to evaluate the accuracy of biomarkers for predicting HCC in the England cohort.26 Only the biomarkers showing a significant association with HCC, as detected by either the CPH model or the RCS model, were included in this analysis. For biomarkers that exhibited high mutual correlation (Pearson's correlation >0.75) (Figure S2), we retained the biomarker with the lower missing rate or higher variance.27 A total of 21 biomarkers and age and sex were finally included in the RSF models. The biomarkers that show nonlinear association with HCC were modelled using RCS with 4 knots. We selected the informative variables based on minimal depth metric28 and evaluated the Harrell's C-index of different combinations of variables in predicting HCC (Figures S3 and S4). The C-index increased from 0.868 with the top five variables to 0.883 with the top ten variables, and then slightly increased to 0.884 with the top 20 variables. When all 23 variables were included, the C-index reached 0.889. Therefore, we decided to use the top ten variables to balance simplicity and accuracy.

Development and validation of the point-based risk score

In the England cohort, we developed a point-based risk score based on the top ten variables selected by the RSF model.29 For each selected variable, we categorized it into several subgroups with the consideration of its association with HCC and the HCC cases distribution across subgroups. Using the CPH model, we calculated the coefficients for each subgroup, assigning a risk point of 0 for the reference group and subgroups with a statistically non-significant coefficient (P > 0.05). For subgroups with a statistically significant coefficient, we assigned a risk point that was the integer ceiling of the coefficient (e.g., 2 for coefficients in the range of 1.01–1.99). Subgroups with the same risk point were integrated together, and the risk score for each participant was calculated by summing the risk points of all variables. We then calculated the point-based risk score for participants in the Scotland and Wales cohort (n = 52,721, with 63 HCC cases) and compared its predictive accuracy with the aMAP score and three non-invasive fibrosis tests (i.e., FIB-4, APRI, and Forns) (Table S3). Harrell's C-index and AUROC were used to measure the predictive accuracy, and we used the DeLong method to test the difference in AUROC values. We also conducted calibration analyses to evaluate the agreement between predicted probabilities and observed outcomes.

To assess the generalizability of our results, we also evaluated the predictive ability of the risk score in UKBB participants with non-White-British ethnic backgrounds (n = 29,315, with 21 HCC cases) who were excluded from our main analysis. Furthermore, we assessed the predictive performance of the risk score in a subpopulation of the TZL.19 A total of 18,139 Chinese individuals that enrolled between 2011 and 2014 were included and were followed-up to the end of 2020 (median follow-up time: 8.0 years). After excluding participants with baseline diagnoses of liver diseases and cancers, we finally included 17,269 participants, among which 69 incident HCC cases occurred (Table S4).

All statistical analyses were performed using R program (R core team, v4.1.1; specifically, mice [v3.16.0] was used for missing value imputation; survival [v3.4-0] and rms [v6.2-0] package was used for the CPH models and RCS models, respectively; randomForestSRC [v3.1.1] was used for the RSF model; pROC [v1.18.0] was used for the analysis of ROC curves). To avoid type one error inflation because of multiplicity, we used the Benjamini-Hochberg false-discovery-rate (FDR) method to adjust P values. Differences were considered statistically significant at an FDR value of <0.05.

Ethics

The UK Biobank received ethical approval from the research ethics committee (REC reference for UK Biobank 11/NW/0382) and participants provided written informed consent. The TZL was approved by the ethics committee of Taizhou Institute of Health Science (No. B013).

Role of the funding source

The funder had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Results

Basic characteristics

Of the 397,702 participants in the England cohort, 54.1% were females and the mean (SD) age at enrollment was 57.2 (8.1) years (Table 1). During the median follow-up of 13.7 years (interquartile range: 13.0–14.2 years) to 31st October 2022, 312 incident HCC cases occurred. HBV/HCV status was tested in only 9423 participants, of which 2.7% were positive. The circulating levels of all biomarkers, except non-albumin protein and total protein, were comparable between the subgroups by HBV/HCV status (Table S5).Table 1 Baseline characteristics across the discovery and validation sets in the UK Biobank.

Characteristics	England cohort	Scotland and wales cohort	Non-White-British cohort	P value	
Sample size (No.)	397,702	52,721	29,315		
HCC (NO.)	312	63	21	0.0069	
HCC (incidence rate, per 100,000)	5.7	8.1	5.3		
Female	215,266 (54.1)	28,982 (55.0)	15,596 (53.2)	<0.0001	
Age, mean (SD), year	57.2 (8.1)	56.7 (8.0)	53.3 (8.3)	<0.0001	
Average total household income before tax (£)		<0.0001	
 Less than 18,000	74,669 (18.8)	9871 (18.7)	6927 (23.6)		
 18,000–30,999	143,745 (36.1)	17,996 (34.1)	13,424 (45.8)		
 31,000–51,999	89,416 (22.5)	12,651 (24.0)	4699 (16.0)		
 52,000–100,000	70,770 (17.8)	9875 (18.7)	3316 (11.3)		
 Greater than 100,000	19,102 (4.8)	2328 (4.4)	949 (3.2)		
Alcohol intake frequency				<0.0001	
 Never or special occasions only	68,501 (17.2)	9640 (18.3)	14,856 (50.7)		
 ≤Twice a week	148,334 (37.3)	20,925 (39.7)	9112 (31.1)		
 Three times or more a week	180,867 (45.5)	22,156 (42.0)	5347 (18.2)		
Smoking status				<0.0001	
 Never	215,538 (54.2)	29,299 (55.6)	20,323 (69.3)		
 Former	141,359 (35.5)	17,146 (32.5)	5548 (18.9)		
 Current	40,805 (10.3)	6276 (11.9)	3444 (11.7)		
Physical activity				<0.0001	
 0–1 day/week	78,081 (19.6)	11,252 (21.3)	5374 (18.9)		
 2–4 days/week	169,211 (42.5)	22,873 (43.4)	13,091 (46.0)		
 5–7 days/week	150,410 (37.8)	18,596 (35.3)	10,011 (35.2)		
BMI categories				<0.0001	
 <18.5	2014 (0.5)	265 (0.5)	177 (0.6)		
 18.5–24.9	130,718 (32.9)	16,328 (31.0)	8425 (28.7)		
 25.0–29.9	170,176 (42.8)	22,622 (42.9)	12,855 (43.9)		
 ≥30	94,794 (23.8)	13,506 (25.6)	7858 (26.8)		
aMAP score group				<0.0001	
 Low (<50)	266,924 (67.1)	37,187 (70.5)	21,590 (73.6)		
 Intermediate (50–60)	123,726 (31.1)	14,808 (28.1)	7294 (24.9)		
 High (>60)	7052 (1.8)	726 (1.4)	441 (1.5)		
Biomarkersa					
 Platelet count (109/L)	245.8 (1.3)	253.3 (1.3)	241.7 (1.3)	<0.0001	
 Alanine aminotransferase (U/L)	23.2 (13.5)	24.4 (14.3)	23.3 (13.9)	<0.0001	
 Gamma glutamyltransferase (U/L)	36.1 (38.9)	38.0 (42.9)	37.0 (37.7)	<0.0001	
 Apolipoprotein B (g/L)	1.0 (0.2)	1.0 (0.2)	1.0 (0.2)	<0.0001	
 Sex hormone binding globulin (nmol/L)	50.9 (25.6)	51.0 (26.0)	45.6 (23.9)	<0.0001	
 Insulin growth factor 1 (nmol/L)	21.4 (5.5)	21.2 (5.5)	21.6 (5.7)	<0.0001	
 Albumin (g/L)	45.3 (2.4)	45.1 (2.4)	45.0 (2.6)	<0.0001	
 Direct bilirubin (μmol/L)	1.8 (0.7)	1.8 (0.7)	1.8 (0.8)	<0.0001	
 Glycosylated hemoglobin (mmol/mol)	35.9 (6.2)	35.9 (6.5)	38.5 (9.1)	<0.0001	
SD, standard deviation; IQR, inter-quantile range.

Values are numbers (percentage at column) unless stated otherwise.

a Values are shown in mean values (standard deviation).

Associations between biomarker levels and HCC

No significant time-varying effect on the risk of HCC was detected for any biomarkers (Table S6). Of the 32 examined biomarkers, we found 25 biomarkers showed significant linear or nonlinear association with HCC (Table S7): linear positive association for C-reactive protein (CRP), cystatin C, direct and total bilirubin, total protein, and non-albumin protein (Fig. 2); linear negative association for albumin, ApoB, cholesterol, LDL, phosphate, and vitamin D (Fig. 2); nonlinear positive association for alkaline phosphatase (ALP), ALT, AST, gamma glutamyltransferase (GGT), HbA1c, glucose, sex hormone binding globulin (SHBG), eGFR, and testosterone (Fig. 3); and nonlinear negative association for PLT, HDL-C, urea, and insulin-like growth factor 1 (IGF-1) (Fig. 3). Furthermore, our study revealed that the associations between clinical biomarkers and HCC risk were consistent across subgroups based on sex, age, and BMI, albeit the nuances in HR estimates (Figures S5–S7; Tables S8–S10). In addition, we conducted a sensitivity analysis to compare the HR of each biomarker based on the random forest-imputed data and the data without any imputation. The results showed a high correlation between the HR estimates, demonstrating the robustness of our findings (Figure S8).Fig. 2 The linear association of age and 12 clinical biomarkers with the risk of hepatocellular carcinoma. The hazard ratios (HR) of biomarkers were calculated from the Cox regression models and represented the risk per one standard deviation increment in the log-transformed biomarker levels. The HR of age represented the risk of HCC per 5-year increment of age. The steel blue shadow denotes the 95% confidence intervals.

Fig. 3 The nonlinear association of 13 clinical biomarkers with risk of hepatocellular carcinoma. The hazard ratios (HR) were calculated by the Cox regression models and represented the risk per one standard deviation increment in the log-transformed biomarker levels. The steel blue shadow denotes the 95% confidence intervals. The vertical dash lines denote the median values of biomarkers.

Development and validation of the point-based risk-scoring system for HCC

The baseline characteristics were significantly varied across the discovery and validation cohorts (Table 1). We selected the top ten variables for predicting HCC, including age, sex, ApoB, PLT, direct bilirubin, IGF-1, ALT, SHBG, GGT, and HbA1c (Figure S3). Upon integrating these top ten variables, the CPH model demonstrated a C-index of 0.883 (standard error [se] = 0.012) for predicting HCC (Figure S4). We developed a point-risk scoring system for HCC based on the top ten variables, as depicted in Fig. 4A and Table S11. The resulting point-based risk score ranged from 0 to 26 and demonstrated a C-index of 0.866 (se = 0.012) for predicting HCC using the CPH model. The calibrations of the CPH model were acceptable in both derivation and validation cohorts, as indicated by the Hosmer–Lemeshow statistics and P values (Fig. 4B). Furthermore, we observed that 71.5% (223/312) of all HCC cases were enriched in the top quintile of the risk score (Fig. 5A), with an HCC incidence rate of 30.82/100,000.Fig. 4 The point-based risk scoring-system for predicting hepatocellular carcinoma. Panel A shows the point-based risk score for each biomarker. The integer 0–5 denotes the risk point for different ranges of biomarker levels. Panel B shows the calibration of the point-based risk score in the discovery and validation sets (cumulative risk of HCC at the end of follow-up were shown). ALT, alanine aminotransferase; ApoB, apolipoprotein B; PLT, platelet count; SHBG, sex hormone binding globulin; GGT, gamma glutamyltransferase; IGF-1, insulin growth factor 1; DBi, direct bilirubin; HbA1c, glycosylated hemoglobin.

Fig. 5 The risk stratification and prediction abilities of the point-based risk score for hepatocellular carcinoma. Panels A and B show the 5-year cumulative risk of HCC in participants with different levels of the risk score in the England (A) and Scotland and Wales (B) cohorts, respectively. Panels C–F show the comparisons of the predictive ability of our risk score with the aMAP score and non-invasive fibrosis tests in the England (C), Scotland and Wales (D), UKBB non-White-British (E), and the Taizhou Longitudinal Study (F), respectively. The ROC curves for our risk score appear smoother compared to the stepwise appearance of the established scores. This difference arises because our point-based risk score is discretely distributed as integers between 0 and 26, while the established scores are continuously distributed over a broader range, resulting in finer steps. The differences in the number of steps between cohorts are due to variations in the distribution of each score across different populations. Larger sample sizes result in more unique score values, affecting the granularity and appearance of the ROC curves. The numbers shown in panel C-F were AUROC values and their 95% confidence interval. ∗∗∗, ∗, and # denotes P value < 0.0001, <0.05, and >0.05, respectively.

The CPH model based on the risk score performed well in the Scotland and Wales cohort, yielding a C-index of 0.848 (se = 0.029). Among all HCC cases, 71.4% (45/63) were identified in the top quintile of the risk score (Fig. 5B), with an HCC incidence rate of 40.91/100,000. Compared to those in the bottom quintile, individuals in the top quintile of the risk score had a more than 73-fold higher risk of developing HCC. The risk score also showed good discriminative ability for HCC, with an AUROC value of 0.863 (95% CI 0.838–0.887) and 0.850 (95% CI 0.791–0.907) in the England cohort and the Scotland and Wales cohort, respectively, which were significantly higher than that of the aMAP score and the three non-invasive fibrosis tests (Fig. 5C and D).

In the UKBB non-White-British cohort, 71.4% (15/21) of the total HCC cases were found in the top quintile of the risk score (HRQ5 vs Q1 = 38.4, 95% CI 5.0–293.9), which achieved an AUROC value of 0.848 (95% CI 0.756–0.940) (Fig. 5E) and a C-index of 0.849 (se = 0.048) for predicting HCC. In the TZL study, 72.5% (50/69) of the total HCC cases were enriched in the top quintile of the risk score (HRQ5 vs Q1 = 41.8, 95% CI 18.5–72.8). The risk score achieved an AUROC value of 0.821 (95% CI 0.744–0.948) (Fig. 5F) and a C-index of 0.814 (se = 0.033) for predicting HCC. The risk score outperformed the aMAP score, FIB-4, and APRI in these two cohorts, while its performance was comparable to that of the Forns index. The sensitivities, specificities, positive predictive values (PPVs), and negative predictive values (NPVs) of all the risk scores across derivation and validation cohorts are shown in Table S12. Briefly, while our risk score demonstrated comparable sensitivity to other risk scores in most cohorts, its PPV was higher.

To enhance the clinical utility of our risk score, we excluded IGF-1 and SHBG as they may not be routinely accessible in daily clinical practice. The simplified risk score yielded C-index and AUROC values ranging from 0.803 and 0.800 in the TZL cohort to 0.850 and 0.842 in the England cohort, respectively. However, the performance of the simplified risk score was inferior to that of the original risk score (Table S13).

Discussion

In this prospective cohort study, we conducted a comprehensive examination of the association between common clinical biomarkers and the risk of HCC. Based on these observed associations, we developed a point-based risk score that incorporates eight biomarkers, age, and sex to stratify and predict the risk of HCC. We found that in both the discovery and validation cohorts, more than 70% of total HCC cases were enriched in the top quintile of the risk score. Moreover, the risk score alone showed an acceptable accuracy for predicting HCC in all validation cohorts.

Our analyses yielded 25 significant linear and nonlinear associations between the biomarker levels and HCC risk, which are consistent with previous findings in some cases.4,30, 31, 32, 33, 34, 35 As anticipated, liver enzymes demonstrated the most significant association with the risk of HCC. Interestingly, we observed that these associations were nonlinear, with positive associations only seen in people with circulating levels above the median value. Previous epidemiological studies also reported nonlinear correlations between serum liver enzymes levels and all-cause mortality risk.20,36 This nonlinear pattern was also observed for other biomarkers including HbA1c, glucose, eGFR, and SHBG. Previous studies have assessed the linear association between certain of these biomarkers and HCC. For example, Li et al. reported a significant linear trend in HCC incidence with increasing HbA1c among patients with type 2 diabetes.37

Our study also uncovered potential curvilinear correlations that have not been well investigated in previous studies, such as the nonlinear positive correlation between circulating SHBG and HCC.32,38 SHBG is a hepatokine and its elevated circulating level was reported to be inversely associated with hepatic steatosis, whereas it was positively associated with liver fibrosis,39,40 connoting that SHBG may be a marker of liver damage. The mechanism underlying the correlation between SHBG and liver damage has been suggested to be through IGF-1,38 which down regulates SHBG production and is inversely related to HCC risk, as reported in our study and elsewhere.32

We also found significant linear and nonlinear associations between several renal biomarkers and HCC risk. Specifically, we observed positive associations between cystatin C and non-albumin protein levels and HCC risk, with a 1-SD increase resulting in a 26% and 2.23-fold higher risk, respectively. On the other hand, we observed that elevating serum levels of phosphate were associated with a decreased risk of HCC. These findings suggest that renal function may reflect the process of liver tumorigenesis, despite these associations being rarely reported in prior studies, and the underlying mechanisms remaining poorly understood.

We also noted a significant negative association of serum vitamin D levels with HCC risk, which is consistent with a previous meta-analysis.35 Likewise, we found that three cardiovascular biomarkers, ApoB, cholesterol, and LDL, which were highly correlated with each other, exhibited a nearly equal degree of inverse association with HCC.4,5 In an experimental study, Qin et al. reported that high serum levels of cholesterol increase antitumor functions of natural killer cells and reduce the growth of liver tumours in mice.41 LDL might have a similar function due to its high correlation with cholesterol at both phenotypic and genetic levels. Furthermore, in the fasting state, serum cholesterol and LDL are mainly derived from very low-density lipoprotein (VLDL), which is mainly bound to ApoB and then secreted from the liver.42 Previous studies have demonstrated that the allele reducing the liver fat content was significantly associated with an increasing level of circulating LDL, and conferred protective effects against hepatocyte injury.43,44 It is noteworthy that our study also revealed a negative association between HDL-C levels and the risk of HCC, but only when the levels were below their respective median values. This finding is interesting because it suggests that higher levels of HDL-C may not necessarily be better in terms of HCC risk.4,45

To facilitate the clinical application of these common clinical biomarkers, we applied a simple method to summarize the relationships between these biomarkers and HCC and developed a point-based risk-scoring system to stratify and predict the HCC risk. This system is more practical than complex prediction models since it allows for a quick assessment of a patient's risk without relying on computers or other electronic devices.29 Although the good prediction capabilities, it is important to acknowledge that our risk score may not be as comprehensive as previous models,17,46 as it does not consider certain HCC-related factors such as AFP levels and HBV/HCV infections. However, AFP may not be an optimal indicator for HCC in the general population and is not a routine testing biomarker in primary care.47 Although HBV/HCV infection is the major contributor to HCC development, the prevalence of HBV/HCV infections is <1% in the UK.48, 49, 50 This means that incorporating the HBV/HCV infection status in the risk model may not be efficient in a population with extremely low prevalence of HBV/HCV.

Unlike previous predictive models designed for specific target populations, such as HBV carriers, our point-based risk-scoring system has been developed for the general population, thereby expanding its applicability. More importantly, the risk score alone has exhibited excellent accuracy in predicting HCC in both the discovery and validation cohorts, outperforming the aMAP score and non-invasive fibrosis tests (e.g., FIB-4). The aMAP score, which was derived from patients with viral (mainly) and non-viral hepatitis, is an important tool for stratifying HCC risk in treated and observed patients with known liver disease. However, this score may be suboptimal for stratifying HCC risk in the general population, where the majority of subjects are unaware of their disease and therefore remain untreated. In contrast, our primary aim was to develop a risk score that incorporates common laboratory parameters, which are more suited for risk stratification in the general population. However, it is worth mentioning that even within the top quintile of the risk score, the 10-year probability of developing HCC remains low (∼0.2%) in populations with a low incidence of HCC, such as the UKBB cohort. In populations with relatively higher incidence rates, such as the TZL population, this probability rises to approximately 1.5%. This limitation underscores the challenges of detecting HCC in a general population where the baseline incidence is relatively low.51 Using this risk score to predict future HCC risk in the general population may not be cost-effective due to the low positive predictive value. Nonetheless, the risk score can serve as a primary screening tool to identify individuals at high risk of developing HCC, as more than 70% of total HCC cases could be found within the top 20% of the risk score. Subsequent targeted screening programs, such as abdominal ultrasound, can then be implemented to achieve high cost-effectiveness in these high-risk populations. Indeed, clinical risk stratification is a key strategy used to identify low- and high-risk subjects to optimize the management.

While our study focuses on developing a risk score for the general population, we acknowledge that clinically at-risk populations, such as patients with cirrhosis, have already been well established as high-risk groups for HCC.52 The primary unmet need in HCC risk stratification lies in differentiating higher-risk individuals within these already at-risk populations,53 as outlined in professional guidelines.22,54 Our findings provide a foundation for understanding the broader associations between common clinical biomarkers and HCC risk, which can be refined for specific clinical settings. Importantly, our simple point-based risk score also enables the identification of potential high-risk individuals in the general population who have not been diagnosed with chronic liver diseases. Early identification of these individuals could lead to timely interventions and improved outcomes. Future research should validate and adapt our point-based risk score within clinically enriched at-risk populations to enhance its clinical utility and address this critical need in HCC risk stratification.

Our study has limitations. First, biomarker levels were measured only once at baseline in all participants, and it is possible that these measurements may not reflect exposure levels across time, resulting in biased association estimates. However, we corrected the “regression dilution bias” using repeated measurements, and thus partly limited the effects of measurement error and within-person variability. Moreover, most biomarkers showed a moderate to high correlation between the values measured at different times, suggesting that the circulating levels of most biomarkers were stable over time in the general population. Second, although the large sample size of the discovery cohort, only a limited number of HCC cases occurred during follow-up, which mainly due to the low risk of HCC in the UK population and also to the insufficient follow-up time.55 Thus, a longer follow-up time is warranted to identify more HCC cases, thereby increasing the statistical power. Moreover, our risk score did not include lifestyle risk factors like alcohol consumption. We focused on biomarkers that can be accurately measured in clinical settings to ensure reliability and reproducibility. Lifestyle factors in the UKBB were self-reported, varying across cohorts and subject to information bias, thus lacking the accuracy of clinical biomarkers. This exclusion may limit the comprehensiveness of our risk score. Future research should consider integrating more accurate measures of environmental exposures, such as biochemical markers, to enhance HCC risk stratification. Finally, our risk score included some biomarkers, such as IGF-1 and SHBG, that may not be routinely measured in daily clinical practice. These biomarkers were selected based on their strong associations with HCC risk, and excluding them significantly decreased the predictive performance of the risk score. However, we acknowledge that the inclusion of costly biomarkers may limit the practical utility of our risk score in routine screening. Future research should focus on identifying alternative, more accessible biomarkers that maintain the predictive accuracy of the risk score. We also compared the predictive performance of our risk score with established systems. While our risk score demonstrated slightly higher PPVs, its comparable sensitivities and specificities indicate that it may be more suitable for HCC risk stratification rather than direct prediction in clinical practice.

In conclusion, our study offers a comprehensive evaluation of the associations between common clinical biomarkers and HCC risk, highlighting the significance of considering nonlinear relationships and biomarker levels within specific ranges. We also propose a straightforward and practical risk model for HCC based on these biomarkers. Clinical validation and further testing are necessary to assess the value of this scoring system in aiding evidence-based clinical decision making and stratification of HCC risk in the general population.

Contributors

ZL conceived the study design. ZL and HY performed statistical analysis and data visualization. ZL and CS wrote the manuscript. HY and RZ scrutinized the statistics. XC, LJ, TZ, and XZ supervised this study. All authors provided critical revisions of the draft and approved the submitted draft. The corresponding author attests that all listed authors meet authorship criteria and that no others meeting the criteria have been omitted. XC is the guarantor. All authors read and approved the final version of the manuscript. The underlying data has been verified by ZL and HY.

Data sharing statement

The UK Biobank data are available from the UK Biobank upon request (www.ukbiobank.ac.uk/). The data of the Taizhou Longitudinal Study used in this study are available on reasonable request from the corresponding author (X.C.).

Declaration of interests

All authors declare no conflict of interests.

Appendix A Supplementary data

Supplementary Material

Acknowledgements

This study was conducted using the UKBB resource under application number 63726. We sincerely appreciate the great works of the UK Biobank collaborators.

Appendix A Supplementary data related to this article can be found at https://doi.org/10.1016/j.eclinm.2024.102796.
==== Refs
References

1 Kulik L. El-Serag H.B. Epidemiology and management of hepatocellular carcinoma Gastroenterology 156 2019 477 491.e1 30367835
2 Aleksandrova K. Boeing H. Nöthlings U. Inflammatory and metabolic biomarkers and risk of liver and biliary tract cancer Hepatology 60 2014 858 871 24443059
3 Fedirko V. Duarte-Salles T. Bamia C. Prediagnostic circulating vitamin D levels and risk of hepatocellular carcinoma in European populations: a nested case-control study Hepatology 60 2014 1222 1230 24644045
4 Zhao L. Deng C. Lin Z. Dietary fats, serum cholesterol and liver cancer risk: a systematic review and meta-analysis of prospective studies Cancers (Basel) 13 2021 1580 33808094
5 Åberg F. Helenius-Hietala J. Puukka P. Interaction between alcohol consumption and metabolic syndrome in predicting severe liver disease in the general population Hepatology 67 2018 2141 2149 29164643
6 Ahn J. Lim U. Weinstein S.J. Prediagnostic total and high-density lipoprotein cholesterol and risk of cancer Cancer Epidemiol Biomarkers Prev 18 2009 2814 2821 19887581
7 Córdoba J. O'Riordan K. Dupuis J. Diurnal variation of serum alanine transaminase activity in chronic liver disease Hepatology 28 1998 1724 1725 9890798
8 Pang Y. Kartsonaki C. Turnbull I. Diabetes, plasma glucose, and incidence of fatty liver, cirrhosis, and liver cancer: a prospective study of 0.5 million people Hepatology 68 2018 1308 1318 29734463
9 Song M. Liu T. Liu H. Association between metabolic syndrome, C-reactive protein, and the risk of primary liver cancer: a large prospective study BMC Cancer 22 2022 853 35927639
10 McGlynn K.A. Sahasrabuddhe V.V. Campbell P.T. Reproductive factors, exogenous hormone use and risk of hepatocellular carcinoma among US women: results from the liver cancer pooling project Br J Cancer 112 2015 1266 1272 25742475
11 Shahini E. Pasculli G. Solimando A.G. Updating the clinical application of blood biomarkers and their algorithms in the diagnosis and surveillance of hepatocellular carcinoma: a critical review Int J Mol Sci 24 2023 4286 36901717
12 Tayob N. Kanwal F. Alsarraj A. The performance of AFP, AFP-3, DCP as biomarkers for detection of hepatocellular carcinoma (HCC): a phase 3 biomarker study in the United States Clin Gastroenterol Hepatol 21 2023 415 423.e4 35124267
13 Best J. Bechmann L.P. Sowa J.P. GALAD score detects early hepatocellular carcinoma in an international cohort of patients with nonalcoholic steatohepatitis Clin Gastroenterol Hepatol 18 2020 728 735.e4 31712073
14 Fan R. Papatheodoridis G. Sun J. aMAP risk score predicts hepatocellular carcinoma development in patients with chronic hepatitis J Hepatol 73 2020 1368 1378 32707225
15 Johnson P.J. Innes H. Hughes D.M. Evaluation of the aMAP score for hepatocellular carcinoma surveillance: a realistic opportunity to risk stratify Br J Cancer 127 2022 1263 1269 35798825
16 Moon A.M. Ioannou G.N. Moving away from a one-size-fits-all approach to hepatocellular carcinoma surveillance Am J Gastroenterol 117 2022 1409 1411 35973179
17 Kim H.S. Yu X. Kramer J. Comparative performance of risk prediction models for hepatitis B-related hepatocellular carcinoma in the United States J Hepatol 76 2022 294 301 34563579
18 Monaghan T.F. Rahman S.N. Agudelo C.W. Foundational statistical principles in medical research: sensitivity, specificity, positive predictive value, and negative predictive value Medicina (Kaunas) 57 2021 503 34065637
19 Chen X. Gole J. Gore A. Non-invasive early detection of cancer four years before conventional diagnosis using a blood test Nat Commun 11 2020 3475 32694610
20 Kunutsor S.K. Apekey T.A. Seddoh D. Liver enzymes and risk of all-cause mortality in general populations: a systematic review and meta-analysis Int J Epidemiol 43 2014 187 201 24585856
21 Levey A.S. Stevens L.A. Schmid C.H. A new equation to estimate glomerular filtration rate Ann Intern Med 150 2009 604 612 19414839
22 Singal A.G. Sanduzzi-Zamparelli M. Nahon P. International Liver Cancer Association (ILCA) white paper on hepatocellular carcinoma risk stratification and surveillance J Hepatol 79 2023 226 239 36854345
23 Clarke R. Shipley M. Lewington S. Underestimation of risk associations due to regression dilution in long-term follow-up of prospective studies Am J Epidemiol 150 1999 341 353 10453810
24 Clarke R. Emberson J.R. Breeze E. Biomarkers of inflammation predict both vascular and non-vascular mortality in older men Eur Heart J 29 2008 800 809 18303034
25 Harrell F. Regression modeling strategies 2015 Springer Cham
26 Ishwaran H. Kogalur U. Blackstone E. Random survival forests Ann Appl Stat 2 2008 841 860
27 Kuhn M. Johnson K. Applied predictive modeling 2013 Springer
28 Ishwaran H. Kogalur U.B. Gorodeski E.Z. High-dimensional variable selection for survival data J Am Stat Assoc 105 2010 205 217
29 Austin P.C. Lee D.S. D'Agostino R.B. Developing points-based risk-scoring systems in the presence of competing risks Stat Med 35 2016 4056 4072 27197622
30 Cho Y. Cho E.J. Yoo J.J. Association between lipid profiles and the incidence of hepatocellular carcinoma: a nationwide population-based study Cancers (Basel) 13 2021 1599 33808412
31 Kunutsor S.K. Apekey T.A. Van Hemelrijck M. Gamma glutamyltransferase, alanine aminotransferase and risk of cancer: systematic review and meta-analysis Int J Cancer 136 2015 1162 1170 25043373
32 Lukanova A. Becker S. Hüsing A. Prediagnostic plasma testosterone, sex hormone-binding globulin, IGF-I and hepatocellular carcinoma: etiological factors or risk markers? Int J Cancer 134 2014 164 173 23801371
33 Kim K. Choi S. Park S.M. Association of fasting serum glucose level and type 2 diabetes with hepatocellular carcinoma in men with chronic hepatitis B infection: a large cohort study Eur J Cancer 102 2018 103 113 30189372
34 Hann H.W. Wan S. Myers R.E. Comprehensive analysis of common serum liver enzymes as prospective predictors of hepatocellular carcinoma in HBV patients PLoS One 7 2012 e47687
35 Zhang Y. Jiang X. Li X. Serum vitamin D levels and risk of liver cancer: a systematic review and dose-response meta-analysis of cohort studies Nutr Cancer 73 2021 1 9
36 Koehler E.M. Sanna D. Hansen B.E. Serum liver enzymes are associated with all-cause mortality in an elderly population Liver Int 34 2014 296 304 24219360
37 Li C.I. Chen H.J. Lai H.C. Hyperglycemia and chronic liver diseases on risk of hepatocellular carcinoma in Chinese patients with type 2 diabetes--national cohort of Taiwan diabetes study Int J Cancer 136 2015 2668 2679 25387451
38 Petrick J.L. Florio A.A. Zhang X. Associations between prediagnostic concentrations of circulating sex steroid hormones and liver cancer among postmenopausal women Hepatology 72 2020 535 547 31808181
39 Sarkar M. VanWagner L.B. Terry J.G. Sex hormone-binding globulin levels in young men are associated with nonalcoholic fatty liver disease in midlife Am J Gastroenterol 114 2019 758 763 30730350
40 Fujihara Y. Hamanoue N. Yano H. High sex hormone-binding globulin concentration is a risk factor for high fibrosis-4 index in middle-aged Japanese men Endocr J 66 2019 637 645 31068503
41 Qin W.H. Yang Z.S. Li M. High serum levels of cholesterol increase antitumor functions of nature killer cells and reduce growth of liver tumors in mice Gastroenterology 158 2020 1713 1727 31972238
42 Olofsson S.O. Borèn J. Apolipoprotein B: a clinically important apolipoprotein which assembles atherogenic lipoproteins and promotes the development of atherosclerosis J Intern Med 258 2005 395 410 16238675
43 Jamialahmadi O. Mancina R.M. Ciociola E. Exome-wide association study on alanine aminotransferase identifies sequence variants in the GPAM and APOE associated with fatty liver disease Gastroenterology 160 2021 1634 1646.e7 33347879
44 Bennet A.M. Di Angelantonio E. Ye Z. Association of apolipoprotein E genotypes with lipid levels and coronary risk JAMA 298 2007 1300 1311 17878422
45 Wang J.B. Abnet C.C. Chen W. Association between serum 25(OH) vitamin D, incident liver cancer and chronic liver disease mortality in the Linxian nutrition intervention trials: a nested case-control study Br J Cancer 109 2013 1997 2004 24008664
46 Kim H.Y. Lampertico P. Nam J.Y. An artificial intelligence model to predict hepatocellular carcinoma risk in Korean and Caucasian patients with chronic hepatitis B J Hepatol 76 2022 311 318 34606915
47 Trevisani F. Garuti F. Neri A. Alpha-fetoprotein for diagnosis, prognosis, and transplant selection Semin Liver Dis 39 2019 163 177 30849784
48 Polaris Observatory Collaborators Global prevalence, treatment, and prevention of hepatitis B virus infection in 2016: a modelling study Lancet Gastroenterol Hepatol 3 2018 383 403 29599078
49 Polaris Observatory HCV Collaborators Global change in hepatitis C virus prevalence and cascade of care between 2015 and 2020: a modelling study Lancet Gastroenterol Hepatol 7 2022 396 415 35180382
50 Liu Z. Jiang Y. Yuan H. The trends in incidence of primary liver cancer caused by specific etiologies: results from the global burden of disease study 2016 and implications for liver cancer prevention J Hepatol 70 2019 674 683 30543829
51 Shelton J. Zotow E. Smith L. 25 year trends in cancer incidence and mortality among adults aged 35-69 years in the UK, 1993-2018: retrospective secondary analysis BMJ 384 2024 e076962
52 Singal A.G. Zhang E. Narasimman M. HCC surveillance improves early detection, curative treatment receipt, and survival in patients with cirrhosis: a meta-analysis J Hepatol 77 2022 128 139 35139400
53 El-Serag H. Kanwal F. Ning J. Serum biomarker signature is predictive of the risk of hepatocellular cancer in patients with cirrhosis Gut 73 2024 1000 1007
54 European Association for the Study of the Liver EASL clinical practice guidelines: management of hepatocellular carcinoma J Hepatol 69 2018 182 236 29628281
55 Liu Z. Song C. Suo C. Alcohol consumption and hepatocellular carcinoma: novel insights from a prospective cohort study and nonlinear Mendelian randomization analysis BMC Med 20 2022 413 36303185
