
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

39251780
71517
10.1038/s41598-024-71517-w
Article
Validation of neuron activation patterns for artificial intelligence models in oculomics
An Songyang Songyang.an@auckland.ac.nz

12
Squirrell David 2
1 https://ror.org/03b94tp07 grid.9654.e 0000 0004 0372 3343 School of Optometry and Vision Science, The University of Auckland, 85 Park Rd, Grafton, Auckland, 1023 New Zealand
2 Toku Eyes Limited NZ, Auckland, New Zealand
9 9 2024
9 9 2024
2024
14 2094014 6 2024
28 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Recent advancements in artificial intelligence (AI) have prompted researchers to expand into the field of oculomics; the association between the retina and systemic health. Unlike conventional AI models trained on well-recognized retinal features, the retinal phenotypes that most oculomics models use are more subtle. Consequently, applying conventional tools, such as saliency maps, to understand how oculomics models arrive at their inference is problematic and open to bias. We hypothesized that neuron activation patterns (NAPs) could be an alternative way to interpret oculomics models, but currently, most existing implementations focus on failure diagnosis. In this study, we designed a novel NAP framework to interpret an oculomics model. We then applied our framework to an AI model predicting systolic blood pressure from fundus images in the United Kingdom Biobank dataset. We found that the NAP generated from our framework was correlated to the clinically relevant endpoint of cardiovascular risk. Our NAP was also able to discern two biologically distinct groups among participants who were assigned the same predicted systolic blood pressure. These results demonstrate the feasibility of our proposed NAP framework for gaining deeper insights into the functioning of oculomics models. Further work is required to validate these results on external datasets.

Keywords

Explainable artificial intelligence
Fundus image
Systolic blood pressure
Neuron activation pattern
Subject terms

Machine learning
Biomarkers
Biomedical engineering
http://dx.doi.org/10.13039/100008354 Callaghan Innovation TEYES2101/PROP-81260-FELLOW-TEYES An Songyang issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Advancements in artificial intelligence (AI), particularly convolutional neural networks (CNNs), have led to oculomics as a new research direction in the field of machine learning for retinal imaging1. Oculomics seeks to explore the broader connections between the retina and systemic health, an area that was difficult for clinicians to explore as the retinal features underpinning these associations were difficult to discern. Studies in oculomics have used AI to explore the associations between the retina and medication conditions such as cardiovascular disease2,3, Alzheimer's disease4, and renal failure5, as well as the associations between the retina and biomarkers of general well-being, such as blood pressure6 and biological aging7.

A limitation of AI when applied to oculomics is the difficulty in explaining their decision process. Widely used explainable AI techniques, specifically saliency maps8–10, have shown significant limitations when applied to these models. The saliency map illustrations published by developers of oculomics models often highlight large and seemingly irrelevant areas of fundus images11–14. Furthermore, saliency maps from different research groups looking at similar oculomics tasks can be visibly different6,7,15. This lack of consistency, as well as studies substantiating the potential for explanations provided by saliency maps to be incomprehensible and untruthful16–18, raises concerns about the utility of these techniques in providing reliable explanations for AI model behavior19.

A key weakness of saliency maps is that their pictorial explanations lack context, and this leaves their interpretation open to bias. This is especially true for oculomics models that use weakly-substantiated retinal features when making their decision. To address this deficiency, we believe that explanations for oculomics models should also encompass quantitative metrics, which are more robust against bias and can provide the missing context. This approach is more practical than exploring entirely new paradigms in explainable AI, which often come with their own limitations. For example, prototype-based approaches20 aim to be self-explainable by implementing mechanisms to break an image down into explainable prototypes and then aggregate the prototypes for the final prediction. However, studies suggest that the prototypes themselves can be uninterpretable21. Multistage architectures22,23 that break complicated AI models down into simpler sub-models have also been used to gain clearer insights into underlying interactions. However, to train the sub-models, the developers must have a pre-understanding of the retinal features required by the overall model. This can be challenging for oculomics tasks where the retinal phenotype is often poorly described.

An approach to generate a quantitative metric that can describe AI model behavior is neuron activation patterns (NAPs)24–26. NAP frameworks try to summarize patterns in neuron activation throughout an AI model and then associate these patterns with the reliability of the final prediction. NAPs are commonly synthesized through unsupervised clustering27, kernel density estimation28, dimensionality reduction approaches such as t-SNE29 or UMAP30, or targeted feature selection31. As NAPs are intended for AI model failure diagnosis, to date, only one other group has published a NAP framework focusing on model interpretation. This group applied NAPs to interpret an AI predicting Alzheimer’s from fundus images32, but their study did not investigate regression models, which are better suited to oculomics tasks using continuous biomarkers. Due to this difference in focus, most NAP frameworks use global average pooling or other equally harsh simplification approaches to simplify the feature map generated by the intermediate stages of AI models26,31,33. However, these approaches may truncate information useful for model interpretation, as feature maps are high-dimensional tensors with channel, width, and height dimensions. Furthermore, the synthesized NAPs are usually clusters or metrics related to the reliability of a model prediction32,34. None of these representations are well-suited to interpreting the continuous predicted output from a regression oculomics model looking at continuous biomarkers.

The aim of this study is to develop and validate a NAP framework that is more suited for interpreting an oculomics model looking at continuous biomarkers, specifically systolic blood pressure (SBP). To achieve this, we developed a novel NAP framework that leverages image similarity metrics, with the resulting NAP being a continuous metric contextualized on the model-predicted outcome. To validate the efficacy of this novel framework, we applied it to CNN predicting SBP from fundus images in the United Kingdom Biobank dataset (UKBB). We then used the outcome of our framework to investigate two clinically-focused hypotheses: (1) The NAP should be correlated with real-world outcomes that can be linked to signs of elevated SBP identified from the fundus, such as cardiovascular disease risk. (2) The NAP should provide clinically relevant insights relating to model prediction behavior that cannot be gleaned from predicted SBP alone.

We believe the novel contributions of this study are as follows:To the best of our knowledge, this is the first study that uses NAPs to investigate a oculomics AI model trained to predict SBP.

To date, the output of most NAP frameworks are clusters or metrics describing the reliability of AI model predictions. In contrast, our framework produces a metric designed specifically to interpret regression models.

The first study to examine the feasibility of image similarity metrics, such as structural similarity, for generating NAPs.

The first study to show that NAPs can identify biologically distinct groups in participants assigned the same predicted outcome from an AI model trained to predict SBP.

Methods

Experiment overview

Database

The UKBB is an open-access research resource containing health information for over half a million participants from the UK who were initially recruited from 2006 to 2010, with follow-up visits occurring until 2022. Ethics approval was obtained from the Northwest Multi-center Research Ethics Committee. All methods were performed in accordance with the relevant guidelines and regulations as per ethics approval and the material transfer agreement signed between our research group and UKBB upon initial data acquisition. During the initial and subsequent assessments of the non-mydriatic, 45° primary field of view, macula-centered fundus images from both the left and right eyes were captured using the TOPCON 3D OCT 1000 Mk 2.

Quality control and data curation

175,788 fundus images from 85,707 participants were obtained from the UKBB. As the raw fundus images had black borders, an internally developed computer vision algorithm was used to crop the image. A preexisting AI-based automated image screening algorithm developed by our group was then used to remove poor-quality images35. This AI model was designed to remove images where a significant portion of the retina was missing, or key retinal landmarks were obscured because of poor or uneven illumination, artifacts, or excessive blurring. Samples of good and poor-quality images are shown in Fig. 1.Fig. 1 Experiment overview. Flow diagram to illustrate the experimental process.

For those participants who had good-quality images, key biometric parameters, such as age, sex, and blood pressure, were retrieved and matched to the fundus image taken during the same UKBB assessment visit. Participants with invalid systolic blood pressure (UKBB field 4080) measurements were removed. After this process, 95,669 good-quality images from 58,606 participants were available for use in this study.

These good-quality images were then divided at 70:15:15 into train, validation, and test splits to develop the DL model. Subsequently, 10,000 images (background set) were sampled from the train and validation splits, and 5000 images (analysis set) were sampled from the test split to develop and validate NAPs (Fig. 1). The demographic details of these subsets are shown below (Table 1).Table 1 Table of demographic details.

Field	70% Train split	15% Validation split	15% Test split	10,000 background set	5000 analysis set	
Number of images	66,968	14,350	14,351	10,000	5000	
Number of participants	47,982	13,463	13,414	9559	4895	
Proportion of men (%)	45.3	44.1	45.4	44.7	43.9	
Age	56.8 (8.3)	56.7 (8.3)	56.5 (8.32)	56.7 (8.28)	56.7 (8.26)	
Body mass index (kg/m2)	27.1 (4.67)	27.2 (4.68)	27.1 (4.68)	27.1 (4.7)	27.2 (4.64)	
Hemoglobin A1C (mmol/mol)	35.8 (6.11)	35.9 (7.17)	35.9 (6.21)	35.8 (6.35)	36 (7.72)	
Cholesterol (mmol/L)	5.70 (1.14)	5.70 (1.13)	5.72 (1.12)	5.71 (1.13)	5.71 (1.14)	
High-density lipoprotein cholesterol (mmol/L)	1.49 (0.389)	1.49 (0.393)	1.49 (0.392)	1.49 (0.385)	1.49 (0.389)	
Systolic blood pressure (mmHg)	137 (18.3)	137 (18.4)	137 (18.4)	137 (18.4)	137 (18.2)	
Diastolic blood pressure (mmHg)	81.5 (10)	81.5 (9.92)	81.7 (10.1)	81.4 (10)	81.5 (9.86)	
Diabetes proportion (%)	4.67%	4.70%	4.59%	4.59%	4.89%	
Summary of datasets used in the study. The 70% Train, 15% Validation, and 15% Test splits were used for training the SBP prediction model. The 10,000 background and 5000 analysis sets were used to develop NAP and are subsamples of the 70% Train and 15% Test splits, respectively. The values are provided as mean and standard deviation in parentheses.

Development of systolic blood pressure prediction model

Using the train, validation, and test splits, a custom CNN based on the EfficientNetV2S36 backbone was trained to predict SBP from fundus images. SBP was chosen as the training target because it is a proven and consistent oculomics task6,14. As the UKBB has two readings for SBP, the average of the two readings was used as the final training target. Information for model architecture and training configuration can be found in Supplementary Information 1. The mean absolute error (MAE) of the model was 11.65 mmHg. This figure was in line with similar studies conducted on the same dataset (MAE of 11.23 mmHg6).

Implementation of neuron activation patterns

Our proposed NAP framework has three stages (Fig. 2):Identification of key stages of CNN architecture that would serve as representative points for monitoring.

Using “similarity averages” to reduce high dimension feature maps into NAPs.

Further simplification of NAPs into “activation pattern scores” to facilitate hypothesis testing.

Fig. 2 Illustration for the high-level process used to implement NAPs. The top element depicts the feature maps at different stages of the CNN, starting from thinner stacks of larger feature maps to deeper stacks of smaller feature maps. The middle element illustrates the concept of converting feature maps to simplified representations (similarity averages). The bottom element illustrates the approach in further simplifying similarity averages into activation pattern scores.

Selection of key points for monitoring

As modern neural networks have multiple layers, monitoring all possible neuron activations is computationally prohibitive. Like previous studies, we defined key points of interest in the neural network based on architectural landmarks, such as pooling layers or the final outputs of a convolutional block, and only examined the activations at those key points. As EfficientNetV2S36 was inherently designed to be multi-staged, we leveraged this design intent and monitored activations at the end of stages 3, 4, 5, and 6.

“Similarity averages” for simplifying feature maps into NAPs

The reduction strategy is illustrated in Fig. 3. For a given feature map, Mx,l∈RC×N×M generated from an input image x at stage l of a CNN, we first identified the 1% most similar feature maps to Mx,l from the 10,000 background set. The background set, Bl:=M1,l,y^1,M2,l,y^2⋯,Mb,l,y^b, can be considered to be a set of feature maps at a specific stage paired with corresponding predicted SBPs. To facilitate the application of image similarity measures, channel-wise min–max scaling was used to scale the N × M 2D matrix from every feature map channel to be within the range of 0–1.Fig. 3 Illustration of the process used to calculate similarity averages. This illustration follows from Fig. 2, with the top element being identical. The bottom left element illustrates the expanded procedures for the “Reduction method” box presented in the middle element of Fig. 2.

Image similarity measures, specifically MS-SSIM37, SSIM38, and the Frobenius norm were then applied as appropriate to quantify the degree of similarity between the feature map and feature map entries in the background set, ‖Mx,l,Mb,l‖. Three different image similarity measures were chosen as different levels of information were captured in the feature maps at the different stages. The feature maps from Stage 3 were larger, more detailed, and more similar in appearance to the input fundus image and thus, to deal with this increased complexity, the more performant MS-SSIM was used to derive similarities. On the other hand, the feature maps from stages 4 and 5 were smaller and less detailed, with activations corresponding to higher-level features. Consequently, the less performant but more computationally efficient SSIM was used to derive similarities. Finally, the smallest feature maps from stage 6 were the least detailed, the simpler Frobenius norm, was deemed appropriate to derive similarities in this layer.

From the set of most similar feature maps, Sx,l⊂Bl, we retrieved the SBP values and calculated a “similarity average”, μx,l (Eq. 1). This similarity average can be understood to be the simplified representation of a feature map, Mx,l. We repeated this process for the l=stage3,stage4,stage5 stages of EfficientNetV2S. The resulting vector of similarity averages, {μx,stage3,μx,stage4,μx,stage5,μx,stage6}, is the NAP synthesized by our framework.1 μx,l=1Sx,l∑M,y^∈Sx,ly^

Summarizing NAPs through an “activation pattern score”

We then derived an “activation pattern score” (Eq. 2) from the vector of similarity averages to facilitate further statistical analysis. Internal experiments indicated that similarity averages from stage 6 had high levels of convergence with the predicted SBP. In contrast, the similarity averages in stage 3 were constrained in a narrow band centered around the population mean for the measured SBP. As such, we decided to define the activation pattern score based on the mean of similarity averages from stages 4 and 5.2 Ax=μx,stage5+μx,stage42

Statistical validation

The activation pattern score was validated across two outcomes on the 5000-analysis set (Fig. 1). As the analysis set was sampled from the hold-out testing set, this ensured it was independent from the 10,000 background set used to define the NAP.

We first investigated the relationship between the activation pattern score and real-world outcomes that are known to be correlated to signs of elevated SBP identified from the fundus (Fig. 4). We chose the pooled cohort equation (PCE) 10-year atherosclerotic cardiovascular disease (ASCVD) risk score39 due to its widespread clinical use and the well-known associations between SBP and increase in ASCVD risk40. To account for the relationship between age, sex and cardiovascular risk, the experiment was performed on age and sex matched groups. As a second outcome (Fig. 5), we then examined whether participants with the same predicted SBP, but different activation pattern scores had differences in biomarkers that are known to be correlated to SBP.Fig. 4 Data quality control for outcome 1 analysis. 763 images from participants with invalid cholesterol were removed as they could not be used calculation of PCE scores. 95 images from participants older than 70 were removed due to small sample size. PCE scores were calculated from remaining participants. Sex and age groups were then constructed for hypothesis testing.

Fig. 5 Data quality control for outcome 1 analysis. 101 images from participants with predicted SBP greater or equal to 160 were removed due to small sample size. The remaining images were divided into predicted SBP bands of 20 mmHg for hypothesis testing.

In both experiments, the statistical significance of the outcomes was validated by comparing the first (0–25% of points) and fourth quartile (75–100% of points) groups. For continuous variables, such as biomarkers, the t-test was used to determine the significance of the difference between the two groups. For categorical variables, such as sex, the chi-squared contingency test was used. As the trained AI model only uses a single fundus image as an input, fundus images from different eyes can result in different activation pattern scores. Accordingly, the tests were performed on a per-image rather than per-participant basis, with the biomarkers from the participant being matched to the corresponding fundus images.

Results

Outcome 1: Difference in activation pattern score for first and fourth quartiles of age and sex-matched PCE score

For both men and women, images from participants in different quartiles of PCE scores had statistically significant differences in activation pattern scores (Table 2). This difference is visually illustrated in Figs. 6 and 7, as the orange (fourth quartile PCE score) and blue traces (first quartile PCE score) of similarity averages are seen to progress in different directions.Table 2 Table comparing neuron activation progression metrics for first and fourth quartiles of PCE scores.

	40 ≤ age < 50	50 ≤ age < 60	60 ≤ age < 70	
First quartile of PCE scores	Fourth quartile of PCE scores	p value	Fourth quartile of PCE scores	First quartile of PCE scores	p value	Fourth quartile of PCE scores	First quartile of PCE scores	p value	
Activation pattern score for men	133.5 (4.1)	136.6 (4.4)	p < 0.001	139.7 (3.6)	137.0 (3.7)	p < 0.001	141.2 (3.3)	139.4 (3.2)	p < 0.001	
Activation pattern score for women	132.5 (3.7)	137.0 (4.7)	p < 0.001	139.5 (4.1)	135.5 (3.7)	p < 0.001	141.8 (3.5)	138.6 (3.1)	p < 0.001	
This table compares the activation pattern scores of participants in first and fourth quartiles of PCE scores across 6 groups. The groups are defined based on sex and age range. Sex is presented in the rows. Age ranges (40 ≤ age < 50, 50 ≤ age < 60, and 60 ≤ age < 70) are presented in the columns.

Significant values are in bold.

Fig. 6 Figure comparing similarity average traces for men’s first and fourth quartiles of PCE scores. This figure illustrates the trace from stage 3 similarity average to predicted SBP for first and fourth quartiles of PCE scores for men. At each stage, the center point represents the group average, and the error bars indicate ± 1 standard deviation from the average. Three plots are shown, corresponding to age brackets of 40 ≤ age < 50, 50 ≤ age < 60, and 60 ≤ age < 70.

Fig. 7 Figure comparing similarity average traces for women’s first and fourth quartiles of PCE scores. This figure illustrates the trace from stage 3 similarity average to predicted SBP for first and fourth quartiles of PCE scores for women. At each stage, the center point represents the group average, and the error bars indicate ± 1 standard deviation from the average. Three plots are shown, corresponding to age brackets of 40 ≤ age < 50, 50 ≤ age < 60, and 60 ≤ age < 70.

Outcome 2: Difference in biomarkers for first and fourth quartiles of prediction-matched activation pattern scores

Across all three predicted SBP bands, fundus images from participants in the first (blue) and fourth (orange) quartiles of activation pattern scores had differently shaped traces (Fig. 8). The distributions of the similarity averages for the first and fourth quartiles were noticeably different in the shallower stages (3–5) but showed convergence at stage 6 and were identical at the predicted SBP. A comparison of the quartile groups (Table 3) revealed that the first quartile of activation pattern scores consistently corresponded to fundus images from a younger cohort with higher diastolic blood pressure (DBP). Conversely, the fourth quartile corresponded with an older cohort with lower DBP but higher Hemoglobin A1C (HbA1C).Fig. 8 Figure illustrating similarity average traces for different activation pattern scores. This figure illustrates the trace from stage 3 similarity average to predicted SBP for first and fourth quartiles groups based on the activation pattern score. At each stage, the center point represents the group average, and the error bars indicate ± 1 standard deviation from the average. Three plots are shown, corresponding to brackets of 100 ≤ predicted SBP < 120, 120 ≤ predicted SBP < 140, and 140 ≤ predicted SBP < 160.

Table 3 Table capturing biomarker comparisons for first and fourth quartiles of activation pattern score.

	100 ≤ Predicted blood pressure < 120	120 ≤ Predicted blood pressure < 140	140 ≤ Predicted blood pressure < 160	
First quartile of activation pattern score	Fourth quartile of activation pattern score	p-value	First quartile of activation pattern score	Fourth quartile of activation pattern score	p-value	First quartile of activation pattern score	Fourth quartile of activation pattern score	p-value	
Age (years)	46.5 (5.7)	52.1 (5.7)	p < 0.001	51.7 (7.4)	59.1 (7.4)	p < 0.001	58.2 (7.5)	62.2 (7.5)	p < 0.001	
Sex assigned at birth (proportion of males)	34%	23%	0.059	47%	43%	0.108	55%	48%	p < 0.05	
Pulse wave arterial stiffness index (m/s)	8.1 (2.4)	8.8 (2.4)	p < 0.05	9.1 (2.7)	9.4 (2.7)	0.096	10.2 (5.9)	10.2 (5.9)	0.974	
Measured systolic blood pressure (mmHg)	119.1 (11.4)	118.1 (11.4)	0.436	132.3 (14.0)	133.4 (14.0)	0.184	149.1 (15.5)	146.7 (15.5)	p < 0.05	
Measured diastolic blood pressure (mmHg)	74.3 (7.8)	72.6 (7.8)	0.059	81.1 (8.4)	78.9 (8.4)	p < 0.001	87.4 (9.5)	84.7 (9.5)	p < 0.001	
HbA1C (mmol/mol)	34.5 (5.9)	35.4 (5.9)	0.297	34.7 (5.0)	37.0 (5.0)	p < 0.001	35.9 (5.6)	37.8 (5.6)	p < 0.05	
This table compares SBP related biomarkers first and fourth quartile groups for the activation pattern score. The table depicts three sets of comparisons across the predicted SBP ranges of 100 ≤ predicted SBP < 120, 120 ≤ predicted SBP < 140, and 140 ≤ predicted SBP < 160. The values displayed in the cells are formatted as mean and then the standard deviation in brackets.

Significant values are in bold.

Discussion

Our results show that the proposed NAP based on similarity averages and activation pattern scores generated a metric that could provide additional insights into the behavior of an oculomics AI model trained to predict SBP. The activation pattern score, a summarized representation of the NAP, exhibited a statistically significant correlation with ASCVD risk defined by the PCE in sex and age-matched participants. Furthermore, participants who were assigned the same predicted SBP by the AI model, but who had different activation pattern scores were identified as being biologically distinct.

As the NAP we have developed is a trace of similarity averages based on predicted SBP, it offered some innate explainability for the underlying SBP model. As shown by Figs. 6, 7 and 8, at stage 3, the similarity averages clustered around a range of 135-145mmHg, which was close to the population mean for measured SBP (137mmHg). This implied that at shallower stages, the similarity search could not find highly similar peers and instead sampled equally from the entire distribution. In the later stages of the CNN, the similarity averages started to converge to the predicted SBP. This suggests feature maps in the later stages were becoming more specialized, with fundus images assigned similar predicted SBPs having more similar feature maps. Previous studies that described a decrease in the randomness of feature maps from the shallower to the deeper layers of neural networks support this hypothesis41. As our framework considered feature maps from the shallower stages of the network, and not only the simplified ones from the last few stages, this could explain why our approach was able to identify biologically distinct groups amongst participants who were assigned the same predicted SBP.

Although promising, further work is needed to validate and refine the methods examined in this study. Firstly, as this study was restricted to the UKBB, there is a need to expand this approach to a wider population to validate the cross-population significance of our results. Secondly, the approach described in this study is an experiment limited to examining the feasibility of NAPs for oculomics, rather than an entirely new approach in explainable AI. Although the trace plots (Figs. 6, 7 and 8) offer interesting insights into model behavior, in themselves they cannot fully explain how an oculomics model arrives at its prediction. As such, we believe that, NAP frameworks could be used in conjunction with existing explainable AI methods, such as saliency maps or conceptual activation vector analysis42, to form a more comprehensive tool for interpreting oculomics AI models.

We identified a number of improvements that could be incorporated into future iterations. While the image similarity metrics used in this study were suitable for model explanation, they may be too slow for real-time failure diagnosis. Processing times could be reduced if a smaller but equally representative background dataset was identified through clustering analysis. The definition for most similar feature maps could also be improved. For the sake of simplicity, we used a static threshold of the top 1%. However, more robust methods, such as triangle thresholding43, could be used to derive a more reliable dynamic threshold. Finally, the standard deviation of the most similar activations could also be integrated into the NAP formulation as this may yield deeper insights.

Conclusion

We found that our proposed NAP framework could identify clinically relevant insights when applied to an AI model trained to predict SBP from fundus images. The activation pattern, a simplified representation of the NAP, was correlated to a participant’s 10-year ASCVD risk as defined by the PCE. The framework was also able to identify that fundus images from participants assigned the same predicted SBP, but different activation pattern scores belonged to biologically distinct cohorts. The first quartile of activation pattern scores represented a younger cohort with higher DBP. The fourth quartile captured an of older cohort with lower DBP and higher HbA1C. Though the approach shows promise, further work, including validation on external datasets and refinements to improve the efficiency and robustness of the technique is required.

Supplementary Information

Supplementary Information.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-024-71517-w.

Acknowledgements

S. An was awarded a R&D Fellowship Grant, TEYES2101/PROP-81260-FELLOW-TEYES by Callaghan Innovation, https://www.callaghaninnovation.govt.nz. The sponsor or funding organization had no role in the design or conduct of this research.

Author contributions

S.A. collected and analyzed the data. D.S. conceived and supervised the project. S.A. wrote the manuscript with assistance from D.S.

Data availability

The scripts that support the findings of the study can be found in the following repository: https://huggingface.co/san727-UOA/nap-for-sbp-oculomics. Please contact the corresponding author at songyang.an@auckland.ac.nz to request data from this study.

Competing interests

D. Squirrell is a co-founder and medical advisor at Toku Eyes Limited NZ. S. An are employees of Toku Eyes Limited NZ. The authors report no other conflicts of interest in this work.

Ethics declarations

This study is a retrospective study of the medical records captured in the UK Biobank. UK Biobank has been granted approval from the North West Multi-Centre Research Ethics Committee as a Research Tissue Bank approval (RTB). This approval from the North West Multi-Center Research Ethics Committee waives informed consent, meaning researchers do not require separate ethical clearance and can operate under RTB approval. The RTB approval was granted in 2011 and was renewed in 2021. As per UK Biobank’s de-identification protocol, the UK Biobank provides de-identified data to researchers in a manner that preserves the anonymity of its participants and, as far as practically possible, does not enable participants to be inadvertently identified. A material transfer agreement between our research group and UK Biobank was finalized on the 28th of March 2022 under application number 86299.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Wagner SK Insights into systemic disease through retinal imaging-based oculomics Transl. Vis. Sci. Technol. 2020 9 6 10.1167/tvst.9.2.6 32704412
Wagner, S. K. et al. Insights into systemic disease through retinal imaging-based oculomics. Transl. Vis. Sci. Technol. 9, 6 (2020).32704412 10.1167/tvst.9.2.6
2. Tseng RMWW Validation of a deep-learning-based retinal biomarker (Reti-CVD) in the prediction of cardiovascular disease: data from UK Biobank BMC Med. 2023 21 28 10.1186/s12916-022-02684-8 36691041
Tseng, R. M. W. W. et al. Validation of a deep-learning-based retinal biomarker (Reti-CVD) in the prediction of cardiovascular disease: data from UK Biobank. BMC Med. 21, 28 (2023).36691041 10.1186/s12916-022-02684-8
3. Vaghefi E Development and validation of a deep-learning model to predict 10-year atherosclerotic cardiovascular disease risk from retinal images using the UK Biobank and EyePACS 10K datasets Cardiovasc. Digit. Health J. 2024 5 59 69 10.1016/j.cvdhj.2023.12.004 38765618
Vaghefi, E. et al. Development and validation of a deep-learning model to predict 10-year atherosclerotic cardiovascular disease risk from retinal images using the UK Biobank and EyePACS 10K datasets. Cardiovasc. Digit. Health J. 5, 59–69 (2024).38765618 10.1016/j.cvdhj.2023.12.004
4. Cheung CY A deep learning model for detection of Alzheimer’s disease based on retinal photographs: A retrospective, multicentre case-control study Lancet Digit. Health 2022 4 e806 e815 10.1016/S2589-7500(22)00169-8 36192349
Cheung, C. Y. et al. A deep learning model for detection of Alzheimer’s disease based on retinal photographs: A retrospective, multicentre case-control study. Lancet Digit. Health 4, e806–e815 (2022).36192349 10.1016/S2589-7500(22)00169-8
5. Joo YS Non-invasive chronic kidney disease risk stratification tool derived from retina-based deep learning and clinical factors NPJ Digit. Med. 2023 6 1 7 10.1038/s41746-023-00860-5 36596833
Joo, Y. S. et al. Non-invasive chronic kidney disease risk stratification tool derived from retina-based deep learning and clinical factors. NPJ Digit. Med. 6, 1–7 (2023).36596833 10.1038/s41746-023-00860-5
6. Poplin R Prediction of cardiovascular risk factors from retinal fundus photographs via deep learning Nat. Biomed. Eng. 2018 2 158 164 10.1038/s41551-018-0195-0 31015713
Poplin, R. et al. Prediction of cardiovascular risk factors from retinal fundus photographs via deep learning. Nat. Biomed. Eng. 2, 158–164 (2018).31015713 10.1038/s41551-018-0195-0
7. Zhu Z Retinal age gap as a predictive biomarker for mortality risk Br. J. Ophthalmol. 2023 107 547 554 10.1136/bjophthalmol-2021-319807 35042683
Zhu, Z. et al. Retinal age gap as a predictive biomarker for mortality risk. Br. J. Ophthalmol. 107, 547–554 (2023).35042683 10.1136/bjophthalmol-2021-319807
8. Chuter B Deep learning identifies high-quality fundus photographs and increases accuracy in automated primary open angle glaucoma detection Transl. Vis. Sci. Technol. 2024 13 23 10.1167/tvst.13.1.23 38285462
Chuter, B. et al. Deep learning identifies high-quality fundus photographs and increases accuracy in automated primary open angle glaucoma detection. Transl. Vis. Sci. Technol. 13, 23 (2024).38285462 10.1167/tvst.13.1.23
9. Dai L A deep learning system for predicting time to progression of diabetic retinopathy Nat. Med. 2024 30 584 594 10.1038/s41591-023-02702-z 38177850
Dai, L. et al. A deep learning system for predicting time to progression of diabetic retinopathy. Nat. Med. 30, 584–594 (2024).38177850 10.1038/s41591-023-02702-z
10. Ju L Hierarchical knowledge guided learning for real-world retinal disease recognition IEEE Trans. Med. Imaging 2024 43 335 350 10.1109/TMI.2023.3302473 37549071
Ju, L. et al. Hierarchical knowledge guided learning for real-world retinal disease recognition. IEEE Trans. Med. Imaging 43, 335–350 (2024).37549071 10.1109/TMI.2023.3302473
11. Zhang K Deep-learning models for the detection and incidence prediction of chronic kidney disease and type 2 diabetes from retinal fundus images Nat. Biomed. Eng. 2021 5 533 545 10.1038/s41551-021-00745-6 34131321
Zhang, K. et al. Deep-learning models for the detection and incidence prediction of chronic kidney disease and type 2 diabetes from retinal fundus images. Nat. Biomed. Eng. 5, 533–545 (2021).34131321 10.1038/s41551-021-00745-6
12. Kim YD Effects of hypertension, diabetes, and smoking on age and sex prediction from retinal fundus images Sci. Rep. 2020 10 4623 10.1038/s41598-020-61519-9 32165702
Kim, Y. D. et al. Effects of hypertension, diabetes, and smoking on age and sex prediction from retinal fundus images. Sci. Rep. 10, 4623 (2020).32165702 10.1038/s41598-020-61519-9
13. Betzler BK Deep learning algorithms to detect diabetic kidney disease from retinal photographs in multiethnic populations with diabetes J. Am. Med. Inform. Assoc. 2023 30 1904 1914 10.1093/jamia/ocad179 37659103
Betzler, B. K. et al. Deep learning algorithms to detect diabetic kidney disease from retinal photographs in multiethnic populations with diabetes. J. Am. Med. Inform. Assoc. 30, 1904–1914 (2023).37659103 10.1093/jamia/ocad179
14. Rim TH Prediction of systemic biomarkers from retinal photographs: Development and validation of deep-learning algorithms Lancet Digit. Health 2020 2 e526 e536 10.1016/S2589-7500(20)30216-8 33328047
Rim, T. H. et al. Prediction of systemic biomarkers from retinal photographs: Development and validation of deep-learning algorithms. Lancet Digit. Health 2, e526–e536 (2020).33328047 10.1016/S2589-7500(20)30216-8
15. Nusinovici S Retinal photograph-based deep learning predicts biological age, and stratifies morbidity and mortality risk Age Ageing 2022 51 afac065 10.1093/ageing/afac065 35363255
Nusinovici, S. et al. Retinal photograph-based deep learning predicts biological age, and stratifies morbidity and mortality risk. Age Ageing 51, afac065 (2022).35363255 10.1093/ageing/afac065
16. Arun N Assessing the trustworthiness of saliency maps for localizing abnormalities in medical imaging Radiol. Artif. Intell. 2021 3 e200267 10.1148/ryai.2021200267 34870212
Arun, N. et al. Assessing the trustworthiness of saliency maps for localizing abnormalities in medical imaging. Radiol. Artif. Intell. 3, e200267 (2021).34870212 10.1148/ryai.2021200267
17. Jin W Li X Fatehi M Hamarneh G Guidelines and evaluation of clinical explainable AI in medical image analysis Med. Image Anal. 2023 84 102684 10.1016/j.media.2022.102684 36516555
Jin, W., Li, X., Fatehi, M. & Hamarneh, G. Guidelines and evaluation of clinical explainable AI in medical image analysis. Med. Image Anal. 84, 102684 (2023).36516555 10.1016/j.media.2022.102684
18. Zhang J Revisiting the trustworthiness of saliency methods in radiology AI Radiol. Artif. Intell. 2023 6 e220221 10.1148/ryai.220221
Zhang, J. et al. Revisiting the trustworthiness of saliency methods in radiology AI. Radiol. Artif. Intell. 6, e220221 (2023).10.1148/ryai.220221
19. Saranya A Subhashini R A systematic review of Explainable Artificial Intelligence models and applications: Recent developments and future trends Decis. Anal. J. 2023 7 100230 10.1016/j.dajour.2023.100230
Saranya, A. & Subhashini, R. A systematic review of Explainable Artificial Intelligence models and applications: Recent developments and future trends. Decis. Anal. J. 7, 100230 (2023).10.1016/j.dajour.2023.100230
20. Carloni, G., Berti, A., Iacconi, C., Pascali, M. A. & Colantonio, S. On the applicability of prototypical part learning in medical images: Breast masses classification using ProtoPNet. In Pattern Recognition, Computer Vision, and Image Processing. ICPR 2022 International Workshops and Challenges (eds. Rousseau, J.-J. & Kapralos, B.) 539–557 (Springer Nature Switzerland, 2023). 10.1007/978-3-031-37660-3_38.
21. Davoodi O Mohammadizadehsamakosh S Komeili M On the interpretability of part-prototype based classifiers: A human centric analysis Sci. Rep. 2023 13 23088 10.1038/s41598-023-49854-z 38155163
Davoodi, O., Mohammadizadehsamakosh, S. & Komeili, M. On the interpretability of part-prototype based classifiers: A human centric analysis. Sci. Rep. 13, 23088 (2023).38155163 10.1038/s41598-023-49854-z
22. Son J An interpretable and interactive deep learning algorithm for a clinically applicable retinal fundus diagnosis system by modelling finding-disease relationship Sci. Rep. 2023 13 5934 10.1038/s41598-023-32518-3 37045856
Son, J. et al. An interpretable and interactive deep learning algorithm for a clinically applicable retinal fundus diagnosis system by modelling finding-disease relationship. Sci. Rep. 13, 5934 (2023).37045856 10.1038/s41598-023-32518-3
23. Hervella ÁS Ramos L Rouco J Novo J Ortega M Explainable artificial intelligence for the automated assessment of the retinal vascular tortuosity Med. Biol. Eng. Comput. 2024 62 865 881 10.1007/s11517-023-02978-w 38060101
Hervella, Á. S., Ramos, L., Rouco, J., Novo, J. & Ortega, M. Explainable artificial intelligence for the automated assessment of the retinal vascular tortuosity. Med. Biol. Eng. Comput. 62, 865–881 (2024).38060101 10.1007/s11517-023-02978-w
24. Cheng, C.-H., Nührenberg, G. & Yasuoka, H. Runtime monitoring neuron activation patterns. In 2019 Design, Automation & Test in Europe Conference & Exhibition (DATE) 300–303 (IEEE, 2019).
25. Geissler, F., Qutub, S., Paulitsch, M. & Pattabiraman, K. A low-cost strategic monitoring approach for scalable and interpretable error detection in deep neural networks. In Computer Safety, Reliability, and Security (eds. Guiochet, J., Tonetta, S. & Bitsch, F.) 75–88 (Springer Nature Switzerland, 2023). 10.1007/978-3-031-40923-3_7.
26. Olber, B., Radlak, K., Popowicz, A., Szczepankiewicz, M. & Chachuła, K. Detection of out-of-distribution samples using binary neuron activation patterns. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 3378–3387 (2023).
27. Ester M Kriegel H-P Sander J Xu X A density-based algorithm for discovering clusters in large spatial databases with noise KDD. 1996 96 226 231
Ester, M., Kriegel, H.-P., Sander, J. & Xu, X. A density-based algorithm for discovering clusters in large spatial databases with noise. KDD. 96, 226–231 (1996).
28. Chen Y-C A tutorial on kernel density estimation and recent advances Biostat. Epidemiol. 2017 1 161 187 10.1080/24709360.2017.1396742
Chen, Y.-C. A tutorial on kernel density estimation and recent advances. Biostat. Epidemiol. 1, 161–187 (2017).10.1080/24709360.2017.1396742
29. van der Maaten L Hinton G Visualizing data using t-SNE J. Mach. Learn. Res. 2008 9 2579 2605
van der Maaten, L. & Hinton, G. Visualizing data using t-SNE. J. Mach. Learn. Res. 9, 2579–2605 (2008).
30. McInnes L Healy J Saul N Großberger L UMAP: Uniform manifold approximation and projection J. Open Source Softw. 2018 3 861 10.21105/joss.00861
McInnes, L., Healy, J., Saul, N. & Großberger, L. UMAP: Uniform manifold approximation and projection. J. Open Source Softw. 3, 861 (2018).10.21105/joss.00861
31. Ma, D. et al. Dr. DNA: Combating silent data corruptions in deep learning using distribution of neuron activations. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, vol. 3 239–252 (Association for Computing Machinery, 2024). 10.1145/3620666.3651349.
32. Yousefzadeh N Neuron-level explainable AI for Alzheimer’s disease assessment from fundus images Sci. Rep. 2024 14 7710 10.1038/s41598-024-58121-8 38565579
Yousefzadeh, N. et al. Neuron-level explainable AI for Alzheimer’s disease assessment from fundus images. Sci. Rep. 14, 7710 (2024).38565579 10.1038/s41598-024-58121-8
33. Tang, K. et al. CORES: Convolutional response-based score for out-of-distribution detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 10916–10925 (2024).
34. Yatbaz, H. Y., Dianati, M., Koufos, K. & Woodman, R. Run-time monitoring of 3D object detection in automated driving systems using early layer neural activation patterns. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops 3522–3531 (2024).
35. Vaghefi E A multi-centre prospective evaluation of THEIATM to detect diabetic retinopathy (DR) and diabetic macular oedema (DMO) in the New Zealand screening program Eye 2022 10.1038/s41433-022-02217-w 36057664
Vaghefi, E. et al. A multi-centre prospective evaluation of THEIATM to detect diabetic retinopathy (DR) and diabetic macular oedema (DMO) in the New Zealand screening program. Eye10.1038/s41433-022-02217-w (2022).36057664 10.1038/s41433-022-02217-w
36. Tan, M. & Le, Q. Efficientnetv2: Smaller models and faster training. In International Conference on Machine Learning 10096–10106 (PMLR, 2021).
37. Wang, Z., Simoncelli, E. P. & Bovik, A. C. Multiscale structural similarity for image quality assessment. In The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003, vol. 2 1398–1402 (2003).
38. Wang Z Bovik AC Sheikh HR Simoncelli EP Image quality assessment: From error visibility to structural similarity IEEE Trans. Image Process. 2004 13 600 612 10.1109/TIP.2003.819861 15376593
Wang, Z., Bovik, A. C., Sheikh, H. R. & Simoncelli, E. P. Image quality assessment: From error visibility to structural similarity. IEEE Trans. Image Process. 13, 600–612 (2004).15376593 10.1109/TIP.2003.819861
39. Yadlowsky S Clinical implications of revised pooled cohort equations for estimating atherosclerotic cardiovascular disease risk Ann. Intern. Med. 2018 169 20 29 10.7326/M17-3011 29868850
Yadlowsky, S. et al. Clinical implications of revised pooled cohort equations for estimating atherosclerotic cardiovascular disease risk. Ann. Intern. Med. 169, 20–29 (2018).29868850 10.7326/M17-3011
40. Vaduganathan M Mensah GA Turco JV Fuster V Roth GA The global burden of cardiovascular diseases and risk: A compass for future health J. Am. Coll. Cardiol. 2022 80 2361 2371 10.1016/j.jacc.2022.11.005 36368511
Vaduganathan, M., Mensah, G. A., Turco, J. V., Fuster, V. & Roth, G. A. The global burden of cardiovascular diseases and risk: A compass for future health. J. Am. Coll. Cardiol. 80, 2361–2371 (2022).36368511 10.1016/j.jacc.2022.11.005
41. Wang L Wang C Li Y Wang R Explaining the behavior of neuron activations in deep neural networks Ad Hoc Netw. 2021 111 102346 10.1016/j.adhoc.2020.102346
Wang, L., Wang, C., Li, Y. & Wang, R. Explaining the behavior of neuron activations in deep neural networks. Ad Hoc Netw. 111, 102346 (2021).10.1016/j.adhoc.2020.102346
42. Kim, B. et al. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV). In Proceedings of the 35th International Conference on Machine Learning 2668–2677 (PMLR, 2018).
43. Rosin PL Unimodal thresholding Pattern Recogn. 2001 34 2083 2096 10.1016/S0031-3203(00)00136-9
Rosin, P. L. Unimodal thresholding. Pattern Recogn. 34, 2083–2096 (2001).10.1016/S0031-3203(00)00136-9
