
==== Front
J Cell Mol Med
J Cell Mol Med
10.1111/(ISSN)1582-4934
JCMM
Journal of Cellular and Molecular Medicine
1582-1838
1582-4934
John Wiley and Sons Inc. Hoboken

10.1111/jcmm.18553
JCMM18553
JCMM-05-2024-163.R1
Original Article
Original Article
Microbe–disease associations prediction by graph regularized non‐negative matrix factorization with L2,1 norm regularization terms
Chen et al.
Chen Ziwei 1 zwchen@bjtu.edu.cn

Zhang Liangzhe https://orcid.org/0009-0005-6618-7765
1
Li Jingyi 1
Chen Hang 1
1 School of Electronic and Information Engineering Beijing Jiaotong University Beijing China
* Correspondence
Ziwei Chen, School of Electronic and Information Engineering, Beijing Jiaotong University, Beijing 100044, China.
Email: zwchen@bjtu.edu.cn

06 9 2024
9 2024
28 17 10.1111/jcmm.v28.17 e1855319 6 2024
15 5 2024
09 7 2024
© 2024 The Author(s). Journal of Cellular and Molecular Medicine published by Foundation for Cellular and Molecular Medicine and John Wiley & Sons Ltd.
https://creativecommons.org/licenses/by/4.0/ This is an open access article under the terms of the http://creativecommons.org/licenses/by/4.0/ License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited.

Abstract

Microbes are involved in a wide range of biological processes and are closely associated with disease. Inferring potential disease‐associated microbes as the biomarkers or drug targets may help prevent, diagnose and treat complex human diseases. However, biological experiments are time‐consuming and expensive. In this study, we introduced a new method called iPALM‐GLMF, which modelled microbe–disease association prediction as a problem of non‐negative matrix factorization with graph dual regularization terms and L2,1 norm regularization terms. The graph dual regularization terms were used to capture potential features in the microbe and disease space, and the L2,1 norm regularization terms were used to ensure the sparsity of the feature matrices obtained from the non‐negative matrix factorization and to improve the interpretability. To solve the model, iPALM‐GLMF used a non‐negative double singular value decomposition to initialize the matrix factorization and adopted an inertial Proximal Alternating Linear Minimization iterative process to obtain the final matrix factorization results. As a result, iPALM‐GLMF performed better than other existing methods in leave‐one‐out cross‐validation and fivefold cross‐validation. In addition, case studies of different diseases demonstrated that iPALM‐GLMF could effectively predict potential microbial‐disease associations. iPALM‐GLMF is publicly available at https://github.com/LiangzheZhang/iPALM‐GLMF.

graph dual regularization L2,1 norm regularization
inertial proximal alternating linearized minimization
microbe–disease association
non‐negative matrix factorization
source-schema-version-number2.0
cover-dateSeptember 2024
details-of-publishers-convertorConverter:WILEY_ML3GV2_TO_JATSPMC version:6.4.8 mode:remove_FC converted:06.09.2024
Chen Z , Zhang L , Li J , Chen H . Microbe–disease associations prediction by graph regularized non‐negative matrix factorization with L2,1 norm regularization terms. J Cell Mol Med. 2024;28 :e18553. doi:10.1111/jcmm.18553

Ziwei Chen and Liangzhe Zhang are joint first authors.
==== Body
pmc1 INTRODUCTION

Microbes are a diverse group of microscopic organisms existing in either single‐cell or multicellular forms, primarily categorized into bacteria, fungi, viruses, archaea and protozoa. 1 They are a ubiquitous and important part of all ecosystems, with their habitats ranging widely, extending even to harsh environments such as polar regions and the ocean depths. 2 Therefore, it is not surprising to find a wealth of symbiotic microbes thriving in the human body, including on the skin, in the gut and throughout the oral cavity. 3 Research indicates that they play significant roles in human health and disease, such as maintaining internal equilibrium, 4 developing the immune system 5 and resisting pathogens. 2 For instance, studies have indicated that the proliferation of pathogenic bacteria in the oral cavity may lead to an inflammatory disease known as periodontitis. 6 The findings demonstrate that periodontitis‐associated microbial communities have highly conserved changes in metabolic and virulence gene expression profiles, whereas healthy samples did not. This suggests that changes in the composition of the oral microbial community may be related to the pathogenesis of periodontitis. 7 Furthermore, there is clinical and histological evidence that topical application of lactic acid can be effective in depigmenting the skin, improving the surface roughness of the skin and reducing mild wrinkles caused by environmental photodamage. 8 Consequently, elucidating the relationship between human diseases and microbes can not only enhance our comprehension of disease pathogenesis, but also provide novel strategies for disease diagnosis and treatment. 9 , 10 , 11 , 12 , 13 Unfortunately, reliance on traditional experimental methods is both laborious and time‐consuming, and it is challenging to fully explore potential microbe–disease associations (MDAs) within a limited timeframe. Consequently, there has been a growing interest in computational models that can predict disease‐associated microbes. 14

In recent years, numerous computational models, including those based on scoring functions, have been developed to predict potential microbe–disease associations. For instance, Chen et al. 15 proposed the first computational model in this field called KATZHMDA, which is based on the KATZ method. In this model, the prediction of potential associations is transformed into an integration based on the number of walks in the network and its own length. It is a valid metric for calculating the probability of potential associations between microbes and diseases. Huang et al. 16 presented the computational model of Path‐Based Human Microbe‐Disease Associations prediction (PBHMDA), which is based on a depth‐first search algorithm for predicting microbes that may be associated with diseases. The model generates a prediction score for each microbe–disease association pair by constructing a heterogeneous network and traversing all connection paths between nodes in the heterogeneous network using a specialized depth‐first search algorithm. Long and Luo 17 proposed a novel computational model for Weighted Meta‐Graph‐based Human Microbe‐Disease Associations prediction (WMGHMDA). The model iteratively implements a pre‐designed weighted meta‐graph search algorithm on a heterogeneous information network, and discovers possible microbe–disease pairs by accumulating the contribution value of the weighted meta‐graph to the microbe–disease pairs as a probability score. Xu et al. 18 developed a novel computational method to discover potential Microbe‐Disease Associations based on the Kronecker Regularized Least Squares (MDAKRLS). The model is designed with Kronecker regularized least squares that have different Kronecker similarities to obtain the prediction scores separately, and the final prediction scores are calculated by integrating the contributions of different similarities. The advantages of these models are that the theory of the algorithms and computational processes involved is relatively easy to understand and the models do not require negative samples for prediction, while the disadvantage is that most models based on scoring functions are not applicable to new diseases.

Researchers have also developed network models for microbe–disease associations prediction. For example, Bao et al. 19 proposed the model of Network Consistency Projection for Human Microbe‐Disease Associations prediction (NCPHMDA), where the model constructs the similarity of nodes in a heterogeneous network to measure the correlation between microbes and diseases, and calculates the consistency projection score to infer latent microbes for diseases. Huang et al. 20 constructed a new model to reveal potential microbial–disease associations by integrating two independent recommendation models: a neighbour‐based prediction model and a graph‐based prediction model. Wu et al. 21 presented a novel computational model employing Random Walking with Restart optimized by Particle Swarm Optimization (PSO) on the heterogeneous interlinked network of Human Microbe‐Disease Associations (PRWHMDA). The model optimizes the random walk and restart of a heterogeneous network of human microbe–disease associations, using a PSO to optimize the random walk parameters and obtain the final association probability vector. Yan et al. 22 proposed a correlation prediction method (BRWMDA) based on similarity and improving bi‐random walk on the disease and microbe networks. The method utilizes network integration and double random walks on the disease and microbe networks. When the maximum number of iterations of both networks is reached, the random walk stops and produces the final correlation probability matrix. The main advantage of these models is that they can fully utilize the topological information in the network. In addition, these models involve fewer parameters, which greatly reduces the difficulty of parameter selection, but the disadvantage is that some network‐based approaches rely heavily on experimentally validated microbe–disease associations, and cannot predict new diseases or microbes in the absence of known association information.

Machine learning techniques especially deep learning have obtained wide applications in bioinformatics due to their better classification performance. 23 , 24 , 25 , 26 , 27 , 28 Concurrently, many computational methods have been developed to help identify the relationship between microbes and diseases. For example, Peng et al. 29 developed a model of Adaptive Boosting for Human Microbe‐Disease Associations prediction (ABHMDA), which reveals microbes associated with a disease by a strong classifier consisting of weak classifiers with their own weights. ABHMDA assigns different weights to multiple weak classifiers to get the final association. Wang et al. developed a semi‐supervised computational model of Laplacian Regularized Least Squares for Human Microbe‐Disease Associations (LRLSHMDA) with good results. Li et al. 30 proposed a novel computational method called BPNNHMDA. The method takes advantage of the fact that the neural network model, including a unique activation function and optimized initial connection weights based on Gaussian interaction profile kernel similarity, which effectively improves the training speed of the model. Hua et al. 31 developed a model of Multi‐View Graph Convolutional Network for Microbe‐Disease Associations prediction (MVGCNMDA), which employs specific data augmentation and multi‐view attentional blocks to reveal microbes associated with diseases. Chen et al. 32 proposed an approach to predict MDAs based on heterogeneous networks and metapath aggregation graph neural networks (MATHNMDA). The model utilizes heterogeneous networks as inputs to the metapath aggregation graph neural network, and employs aggregation and attention mechanisms among metapaths to integrate the semantic information of all the different metapaths, thereby obtaining the final embeddings of microbial nodes and disease nodes.

The discovery of potential microbe–disease associations will undoubtedly be of great help in research to understand disease pathogenesis and develop treatments for human diseases. Since traditional biological experiments are generally time‐consuming and labour‐intensive, efficient and reliable computational prediction methods are urgently needed. In recent, great progress has been made in developing computational models for predicting potential microbe–disease associations. As a machine learning method, the matrix factorization approach has proven to be an effective tool and has been widely used in bioinformatics research. For example, Gönen 33 proposed a kernelized Bayesian matrix factorization with twin kernels method to predict drug‐target interactions. He et al. 34 developed a novel predictive model of Graph Regularized Non‐negative Matrix Factorization for Human Microbe‐Disease Association prediction (GRNMFHMDA).

In this study, we put forward a novel computational model named iPALM‐GLMF, which was a non‐negative matrix factorization model based on graph dual regularization terms and L2,1 norm regularization terms. The graph dual regularization terms were used to integrate the geometric information of the microbe similarity matrix and the disease similarity matrix, and the L2,1 norm regularization terms were used to ensure the sparsity of the matrices obtained from the non‐negative matrix factorization. We then used the non‐negative double singular value decomposition (NNDSVD) 35 to provide valid and interpretable initial component matrices for the matrix factorization and used an inertial proximal alternating linear minimization iterative process, which has been shown to converge to the KKT point, to obtain the final resultant matrix factorization. 36

Overall, our main contributions were summarized as follows: We introduced a novel approach to improve nonnegative matrix factorization by adding graph dual regularization terms and L2,1 norm regularization terms.

By using manifold theory, introducing graph dual regularization terms to efficiently integrate different similarity matrices and preserve manifold features of the data space.

To improve interpretability and mitigate the effects of inherent noise in the microbe and disease feature spaces, L2,1 norm regularization terms was applied to the feature matrices to select the most representative or discriminative sparse features.

NNDSVD was used to initialize the non‐negative matrix factorization and to solve the matrix factorization using a fast convergent inertial proximal alternating linearization minimization algorithm.

Numerous experimental results have shown that iPALM‐GLMF has better performance than other state‐of‐the‐art methods. Experimental results on two data sets, that is, HMDAD and Disbiome, indicate that iPALM‐GLMF model consistently outperforms the other five state‐of‐the‐art methods. Case studies of three common diseases, colorectal cancer, inflammatory bowel disease (IBD) and asthma, further validate the effectiveness of iPALM‐GLMF.

2 MATERIAL

2.1 Human microbe–disease associations

To establish the human microbe–disease interaction network, we retrieved known microbe–disease associations from the Human Microbe‐Disease Association Database (HMDAD) (http://www.cuilab.cn/hmdad). 37 There were 483 experimentally confirmed microbe–disease associations between 39 diseases and 292 microbes. After removing redundant associations, we obtained 450 associations. In addition, Janssens et al. released a new microbe–disease association database called Disbiome (https://disbiome.ugent.be/home), in which 5573 experimentally confirmed human microbe–disease associations were collected from previously published literature and different databases, including 240 diseases and 1098 microbes. 38 In Disbiome, a microbe–disease pair may be recorded multiple times depending on the assay. After filtering out duplicates, we ended up downloading 4351 associations between 218 diseases and 1052 microbes. Overall, the specific statistics of the two microbe–disease association datasets are shown in Table 1.

TABLE 1 The information of two microbe–disease associations datasets.

Dataset	# Microbes	# Diseases	# Association	
HMDAD	292	39	450	
Disbiome	1052	218	4351	

For better description, we formulated microbe–disease associations as a binary matrix A∈ℝm×d with m and d representing the numbers of microbes and diseases, respectively. If there exists an experimentally verified relationship between a microbe mi and a disease dj, Aij equals to 1, otherwise 0.

2.2 Gaussian interaction profile kernel similarity for diseases and microbes

We calculated Gaussian kernel similarity for microbes and diseases based on the hypothesis that microbes related to common diseases are more likely to show same functions. 39 Specifically, since the ith row and the jth column of the adjacency matrix A denote the interactions between microbes mi or disease dj and all microbes or all diseases, we denote IPmi and IPdj as the interaction profiles of microbe mi with disease dj, respectively. The Gaussian kernel similarity between microbes and diseases are defined as follows: (1) GMmimj=exp−γmIPmi−IPmj2,

(2) GDdidj=exp−γdIPdi−IPdj2,

where γm and γd represent the normalized kernel bandwidths and are defined as follows: (3) γm=γm′1m∑i=1mIPmi2,

(4) γd=γd′1d∑i=1dIPdi2,

where γm′ and γd′ are the original bandwidths, and generally both are set to 1.

3 METHODS

In this paper, we presented a novel model iPALM‐GLMF, which modelled the microbe–disease associations prediction problem as a non‐negative factorization problem with graph dual regularization terms and L2,1 norm regularization terms. iPALM‐GLMF took microbe–disease associations matrix A, Gaussian interaction profile kernel similarity of microbes GM and Gaussian interaction profile kernel similarity of diseases GD as inputs, and utilized GM and GD to construct the graph dual regularization terms, and solved the non‐negative matrix factorization problem of A using the graph dual regularization terms and L2,1 norm regularization terms to obtain the feature matrices of microbes and diseases. Finally, the feature matrices were used to predict potential microbe–disease associations. A brief flow chart of the model iPALM‐GLMF is shown in Figure 1.

FIGURE 1 Flow chart of potential microbe–disease association prediction based on the computational model of iPALM‐GLMF.

3.1 Non‐negative matrix factorization

The non‐negative matrix factorization (NMF) is a method for finding two low‐rank non‐negative matrices whose product approximates the original non‐negative matrix well. 40 It incorporates non‐negative constraints to obtain a component‐based representation and enhances the interpretability of the problem accordingly. In microbe–disease associations prediction, the non‐negative matrix factorization (NMF) of associations matrices is widely used to obtain low‐dimensional feature representations of microbes and diseases in matrices space. The general form of NMF is as follows: (5) minA−XYTF2s.t.X≥0,Y≥0.

where X and Y represent the latent feature matrices of microbes and diseases, respectively. k is the rank of X and Y, k≪minm,d, X∈ℝm×k, Y∈ℝd×k. The non‐negativity constraint terms are adopted to ensure non‐negativity of X and Y.

3.2 Graph dual regularized non‐negative matrix factorization

Cai et al 41 proposed a graph regularized non‐negative matrix factorization (GNMF) to find a compact representation that reveals hidden semantics while respecting the intrinsic geometric structure. It has been shown that learning performance can be greatly improved if the information about the flow structure contained in the data is utilized. 42 , 43 In addition, Shang et al. 44 introduced graph dual regularization terms based on data manifolds and feature manifolds.

To obtain the geometric information of microbes and diseases, two K‐nearest neighbour graphs Nm and Nd are constructed for microbes and diseases based on GM and GD, respectively.

For two microbes mi and mj, the weight of the edge between vertices i and j in graph Nm is defined as follows. (6) Nijm=1,j∈NKiandi∈NKj0,j∉NKiandi∉NKj0.5,otherwise,

where NKi denotes the sets of K most similar microbes of microbes mi according to GM. Based on Nm and GM, a sparse matrix GM^ij is computed as follows: (7) GM^ij=NijmGMij,∀i,j.

Here GM^ is a weight matrix representing the microbes neighbour graph. The graph Laplacian of GM^ is Lm=Dm−GM^, where Dm is a diagonal degree matrix with Diim=∑rGM^ir.

Similarly, the weight matrix GD^ corresponding to the diseases neighbour graph is computed as follows: (8) GD^ij=NijdGDij,∀i,j.

The graph Laplacian of GD^ is Ld=Dd−GD^, where Dd is a diagonal degree matrix with Djjd=∑qGD^jq.

The normalized graph Laplacian forms of Lm and Lm are as follows: (9) L~m=Dm−1/2LmDm−1/2,

(10) L~d=Dd−1/2LdDd−1/2.

The optimization model of graph dual regularization terms non‐negative matrix factorization (GDNMF) of the microbe–disease interaction matrix A is formulated as follows: (11) minX,Y12A−XYTF2+λmTrXTL~mX+λdTrYTL~dY.s.t.X≥0,Y≥0

where λm and λd are regularization parameters.

3.3 GDNMF with L2,1 norm regularization terms

Zhang et al. 45 introduced L2,1 norm regularization terms into graph dual regularization non‐negative matrix factorization, which is used to ensure the sparsity of the matrices obtained by the factorization. The optimization model of GDNMF with L2,1 norm regularization terms is formatted as follows: (12) minX,Y12A−XYTF2+λmTrXTL~mX+λdTrYTL~dY+λlX2,1+Y2,1.s.t.X≥0,Y≥0

where λl is a regularization parameter, X2,1 and Y2,1 represent L2,1 norms of matrix X and Y, respectively, and X2,1=∑i∑jxij21/2, Y2,1=∑i∑jyij21/2.

4 ALGORITHM

4.1 Non‐negative double singular value decomposition

Non‐negative double singular value decomposition (NNDSVD) is a method that improves the initialization phase of non‐negative matrix factorization (NMF) by providing valid and interpretable initial component matrices for matrix factorization. 35 Based on the basic property of singular value decomposition, for matrix A, can be expressed as the sum of k leading singular factors A=∑i=1kσiuiviT, where σ is the nonzero singular values of A, and uivii=1k are the corresponding left and right singular vectors.

For a vector or matrix a, a+=max0,a represents nonnegative section of a, a−=max0−a represents nonpositive section of a, a=a+−a−. A=∑i=1kσiuiviT can be transformed to the following form: (13) A=∑i=1kuivi=∑i=1kui+vi++ui−vi−−ui−vi++ui+vi−.

4.2 Proximal alternating linearized minimization

Bolte et al. 46 introduced a Proximal Alternating Linearized Minimization method (PALM), which has global convergence results for nonconvex and nonsmooth semialgebraic problems.

Model (12) can be derived to the following form: (14) minX,Y12A−XYTF2+RX+RY.s.t.X≥0,Y≥0

where RX=λmTrXTL~mX+λlX2,1, RY=λdTrYTL~dY+λlY2,1. The nonnegative constraint of Formula (14) can be transformed to the following form: (15) X≥0→δX=X,X≥0,∞,otherwise,

(16) Y≥0→δY=Y,Y≥0,∞,otherwise.

Model (14) can be derived to the following form: (17) minψX,Y=min12A−XYTF2+RX+RY+δX+δY.

To solve model (17), the Gauss–Seidel method is used. The specific derivations are as follows: (18) Xi+1∈argminXψXYi

(19) Yi+1∈argminYψXi+1,Y

Let GX,Y=12A−XYF2+RX+RY, and substitute Yi into ψX,Y to remove the constant term and get Xi+1∈argminδX+RX+A−XYF2, where GXYi is smooth function. Then the second‐order Taylor series of GXYi at a point Xi is given by: (20) Xi+1∈argminXX−Xi∇XGXiYi+12∇X∇XGXiYiX−XiF2+δX,

where ∇XG is the partial derivative of G with respect to X.

Define the proximal map of f: proxtf=argminfu+12tu−xF2u∈ℝm, where f: ℝm→−∞+∞ is the lower semi‐continuous function to ensure non‐negativity, x is a fixed point, t is a constant. According to the definition of proximal map, the solution of Formula (20) is as follows: (21) Xi+1∈proxc1iδXXi−1c1i∇XGXiYi.

where c1i=∇X∇XGXiYi=YiYTTF.

Then for a sequence XiYii∈ℕ, parameters c1i and c2i, we have (22) Xi+1∈proxc1iδXXi−1c1i∇XGXiYi,Yi+1∈proxc2iδYYi−1c2i∇YGXi+1Yi.

4.3 Inertial terms

Polyak has showed that the inertial term accelerates the convergence of the standard gradient method while the cost of each iteration remains essentially unchanged. 47 A class of proximal methods has been considered by Attouch 48 for maximal monotone operators in the context of second‐order differential equations in time. These methods are called the inertial proximal methods. In PALM, the commonly used optimization scheme is the first‐order gradient descent method, thus the inertial term is used in order to speed up the convergence.

4.4 Inertial proximal alternating linearized minimization

Let G denote the objective function of model (12). Then, model (12) can be expressed as follows: (23) GX,Y=12A−XYTF2+λmTrXTL~mX+λdTrYTL~dY+λlX2,1+Y2,1.

The partial derivatives of the function G with respect to X and Y, respectively, are as follows: (24) ∂G∂X=A−XYTYT+λmL~mX+λl∂X2,1∂X

(25) ∂G∂Y=A−XYTYT+λdL~dY+λl∂Y2,1∂Y

For sequences XiYii∈ℕ, m1im2ii∈ℕ, d1id2ii∈ℕ, parameters c1i, c2i, β1i, β2i, we can get (26) d1i=Xi+α1iXi−Xi−1,m1i=Xi+β1iXi−Xi−1,Xi+1∈proxc1iδXd1i−1c1i∇XGm1iYi.

(27) d2i=Yi+α2iYi−Yi−1,m2i=Yi+β2iYi−Yi−1,Yi+1∈proxc2iδYd2i−1c2i∇YGXi+1m2i.

The detailed steps of iPALM‐GLMF are illustrated in Figure 2. The parameter values used in our model are set based on a previous study. 49

FIGURE 2 Pseudocode for the iPALM‐GLMF algorithm.

5 EXPERIMENTS

5.1 Evaluation metrics

To evaluate the performance of iPALM‐GLMF, we performed two types of cross‐validations, namely global LOOCV and fivefold cross‐validation, in the datasets HMDAD and Disbiome. In global LOOCV, each time we took turns to select a sample from the recorded microbe–disease associations as a test sample and trained our model with the remaining known associations, and finally, the test sample was sorted with all unrecognized microbe–disease pairs. In fivefold cross‐validation, the known microbe–disease associations were randomly and uniformly divided into five parts, and each part was sequentially picked as a test sample, and the remaining four parts were used as training samples. As with global LOOCV, all unknown microbe–disease pairs were treated as candidate samples. To mitigate the possible impact of sample partitioning on the prediction effect, we randomly partitioned the known microbe–disease associations 100 times. In two cross‐validations, we implemented iPALM‐GLMF to obtain a list of scores for all microbe–disease pairs and ranked the scores of each test sample against the scores of the candidate samples. If the test sample ranked before a given threshold, we assumed that the model successfully predicted the association. Notably, we recalculated the microbe (disease) Gaussian interaction profile kernel similarity during each LOOCV and fivefold cross‐validation because the adjacency matrix changed when one or some of the known microbe–disease associations were removed.

5.2 Performance comparison

Based on the cross‐validation results, we used AUPR and AUC values as metrics to evaluate the performance of iPALM‐GLFM. We compared iPALM‐GLFM with the following state‐of‐the‐art methods on the same dataset.

KATZHMDA 15 based on KATZ measure achieves the prediction of potential disease–microbe association through calculating the number and length of paths between two nodes in microbe–disease heterogeneous network.

LRLSHMDA 50 is a semi‐supervised learning calculation model based on Laplacian regularized least squares classification.

NTSHMDA 51 is a random‐walk based predictive model which predicts human microbe–disease associations by integrating network topological similarities.

BiRWHMDA 52 is a method for predicting potential microbe–disease associations by double random walks on heterogeneous networks.

ABHMDA 29 is a model for revealing disease‐associated microbes through a strong classifier composed of weak classifiers with corresponding weights.

Our method was compared with five baselines under fivefold cross‐validation and global LOOCV on two datasets, namely HMDAD and Disbiome. We used AUC values and AUPR values as indicators to evaluate each method. For better visual comparison, the corresponding ROC curves for iPALM‐GLMF, KATZHMDA, LRLSHMDA, NTSHMDA, BiRWHMDA and ABHMDA were shown in Figures 3 and 4.

FIGURE 3 The graphs show the AUCs of iPALM‐GLMF in fivefold cross‐validation (0.9464) and global LOOCV (0.9587) under the HMDAD database, respectively, which outperformed all the aforementioned models (KATZHMDA, LRLSHMDA, NTSHMDA, BiRWHMDA and ABHMDA).

FIGURE 4 The graphs show the AUCs of iPALM‐GLMF in fivefold cross‐validation (0.8660) and global LOOCV (0.8850) under the Disbiome database, respectively, which outperformed all the aforementioned models (KATZHMDA, LRLSHMDA, NTSHMDA, BiRWHMDA and ABHMDA).

On the HMDAD database, iPALM‐GLMF performed best compared to the other five baseline methods, with average AUCs of 0.9464 ± 0.0039 and 0.9587 under fivefold cross‐validation and global LOOCV, respectively. On the Disbiome database, iPALM‐GLMF performed best compared to the other five baseline methods, with average AUCs of 0.8660 ± 0.0015 and 0.8850 under fivefold cross‐validation and global LOOCV, respectively. The results indicated that our method was effective in predicting novel microbe–disease associations.

To further evaluate the validity of our model, the AUPR values under fivefold cross‐validation for the two databases were shown in Figure 5. The average AUPR of our method under HMDAD and Disbiome databases were: 0.8476 and 0.4515, respectively, which were better than the baseline method.

FIGURE 5 Comparison of AUPR values for iPALM‐GLMF and five other methods using fivefold cross‐validation.

The performance of iPALM‐GLMF on HMDAD, Disbiome at fivefold cross‐validation was summarized in Table 2. It could be seen that our method had the best AUC and AUPR on the HMDAD dataset. The main reason may be that the Disbiome was sparser than the HMDAD. The density of HMDAD was 3.95% and the density of Disbiome was 1.90%. Therefore, the iPALM‐GLMF method was better trained on HMDAD than Disbiome.

TABLE 2 Performance of the iPALM‐GLMF method under fivefold cross‐validation on two datasets.

Dataset	AUC	AUPR	
HMDAD	0.9464 ± 0.0039	0.8476	
Disbiome	0.8660 ± 0.0015	0.4515	

5.3 Ablation experiment

In this section, we sought to determine the impact of several techniques on the performance of our proposed iPALM‐GLMF. To this end, we evaluated iPALM‐GLMF, iPALM‐GLMF (without NNDSVD, i.e. SVD was used in the initialization phase of the matrix factorization), iPALM‐GLMF (λm=0, i.e. the graph regularization term for microbe was not used), iPALM‐GLMF (λd=0, i.e. the graph regularization term for disease was not used), iPALM‐GLMF (λl=0, i.e. L2,1 norm regularization term was not used) and PALM‐GRMF (i.e. inertial forces was not used). The results of the above settings were shown in Tables 3 and 4. In fivefold cross‐validation, the iPALM‐GLMF showed better performance than in other settings.

TABLE 3 AUC values of different algorithms under fivefold cross‐validation.

Method	HMDAD	Disbiome	
iPALM‐GLMF	0.9464 (0.0044)	0.8660 (0.0064)	
iPALM‐GLMF (without NNDSVD)	0.8170 (0.0020)	0.8353 (0.0007)	
iPALM‐GLMF (λm=0)	0.9190 (0.0040)	0.8410 (0.0025)	
iPALM‐GLMF (λd=0)	0.9323 (0.0059)	0.8539 (0.0014)	
iPALM‐GLMF (λl=0)	0.9421 (0.0039)	0.8594 (0.0016)	
PALM‐GLMF	0.9412 (0.0017)	0.8557 (0.0023)	
Note: The maximum AUC on each dataset is shown in bold. Standard deviation is shown in parentheses.

TABLE 4 AUPR values of different algorithms under fivefold cross‐validation.

Method	HMDAD	Disbiome	
iPALM‐GLMF	0.8584 (0.0046)	0.4520 (0.0059)	
iPALM‐GLMF (without NNDSVD)	0.5065 (0.0042)	0.2544 (0.0028)	
iPALM‐GLMF (λm=0)	0.8214 (0.0062)	0.4205 (0.0024)	
iPALM‐GLMF (λd=0)	0.8368 (0.0074)	0.4311 (0.0041)	
iPALM‐GLMF (λl=0)	0.8565 (0.0048)	0.4519 (0.0025)	
PALM‐GLMF	0.8574 (0.0051)	0.4513 (0.0061)	
Note: The maximum AUPR on each dataset is shown in bold. Standard deviation is shown in parentheses.

5.4 Case studies

To further evaluate whether iPALM‐GLMF could demonstrate accurate and robust performance, we conducted case studies on two different types of diseases under colorectal cancer, inflammatory bowel disease (IBD) and asthma. These studies were conducted using the HMDAD database.

In the first case study, we performed potential microbe prediction for colorectal cancer and Inflammatory bowel disease (IBD). Specifically, we categorized all unknown samples under the same disease and verified whether the association between the top 10 microbes and the disease under study was validated by relevant literature. Colorectal cancer is the second leading cause of cancer deaths in the United States, and the incidence in young adults is increasing each year. 53 Therefore, there is an urgent need for novel and sensitive biomarkers that can detect colorectal cancer in an effective and timely manner. Researchers have linked many microbes to colon cancer. For example, D. A. Geier and M. R. Geier 54 unearthed a potential link between Clostridium difficile intestinalis infection and colon cancer incidence, finding that adults with Clostridium difficile intestinalis had a significantly increased incidence of colon cancer. Another example is the discovery by Ralser et al. 55 that Helicobacter pylori promotes colorectal carcinogenesis by deregulating intestinal immunity and inducing a mucus‐degrading microbiota signature, based on the evidence they provide suggesting that H. pylori infection is a strong causal facilitator of colorectal carcinogenesis. As shown in Table 5, we implemented iPALM‐GLMF to discover potentially relevant microbe for colorectal cancer and found that 9 out of the top 10 predictions were confirmed by relevant literature.

TABLE 5 The top 10 potential microbes related to colorectal cancer identified by iPALM‐GLMF.

Rank	Microbe	Evidence	
1	Helicobacter pylori	PMID: 38328335	
2	Clostridium difficile	PMID: 38193707	
3	Clostridium coccoides	PMID: 28661219	
4	Clostridia	PMID: 38068869	
5	Lachnospiraceae	PMID: 28988196	
6	Desulfovibrio	Unconfirmed	
7	Prevotellaceae	PMID: 37469407	
8	Bacteroidetes	PMID: 29170280	
9	Porphyromonadaceae	PMID: 37072632	
10	Bacteroides vulgatus	PMID: 38033588	

The main types of inflammatory bowel disease (IBD), which include Crohn's disease and ulcerative colitis, are caused in part by bacteria that may activate the patient's immune system to attack foreign bodies. 56 Once activated, the patient's immune system has difficulty regulating and destroying the gastrointestinal tract, leading to IBD symptoms. Recent research indicates a close correlation between various microorganisms and IBD. For instance, a reduction in members of the phyla Bacteroidetes and Firmicutes has been observed in IBD, particularly across different variants. 57 As shown in Table 6, we implemented iPALM‐GLMF to discover potentially relevant microbe for IBD and found that 10 of the top 10 predictions were confirmed by relevant literature. The high prediction accuracy suggested that our model could be used for real‐life applications.

TABLE 6 The top 10 potential microbes related to inflammatory bowel disease identified by iPALM‐GLMF.

Rank	Microbe	Evidence	
1	Clostridium difficile	PMID: 37894185	
2	Bacteroidetes	PMID: 37894185	
3	Desulfovibrio	PMID: 29462845	
4	Firmicutes	PMID: 25307765	
5	Staphylococcus epidermidis	PMID: 33618750	
6	Clostridium coccoides	PMID: 19235886	
7	Clostridia	PMID: 25307765	
8	Staphylococcus	PMID: 33618750	
9	Enterobacteriaceae	PMID: 30319571	
10	Clostridiales	PMID: 38335423	

In the second case study, we performed relevant microbe prediction for asthma with the goal of evaluating the model's ability to predict associations between unknown microbes and disease in the absence of any known relevant microbes. Specifically, we replaced all microbes associated with a specific disease in the adjacency matrix with zero. After model prediction, we validated the number of microbes sampled in the top 20 ranked diseases confirmed in the relevant literature, as shown in Figure 6, with results demonstrating that HMDAD has included 2 microbes as well as 16 microbes that have been confirmed in the literature, with only 2 that have not yet been validated. In other words, 18 of the top 20 microbes predicted by our model have been confirmed, further demonstrating the validity of iPALM‐GLMF.

FIGURE 6 Prediction results of top‐20 Asthma‐associated microbes.

6 CONCLUSION

Recognizing potential microbe–disease associations not only contributes to disease diagnosis, treatment and prognosis, but also to microbe‐oriented therapies in precision medicine. In this article, we proposed a novel matrix factorization‐based model called iPALM‐GLMF for inferring potential microbe–disease associations. In the model, we combined graph dual regularization terms with L2,1 norm regularization terms, which was used to capture information about the geometric structure in the microbe similarity matrix and the disease similarity matrix and to ensure the sparsity of the matrices obtained from the nonnegative matrix factorization. We then solved the matrix factorizations with graph dual regularization terms and L2,1 norm regularization terms using an inertial proximal alternating linearization minimization algorithm that achieves global convergence. Cross‐validation results demonstrated the superior performance of iPALM‐GLMF over many previous computational methods. Case studies on colorectal cancer, inflammatory bowel disease and asthma also showed that our model could exhibit reliable and accurate predictive performance. It is reasonable to conclude that iPALM‐GLMF will be useful for microbe–disease association prediction.

Although our model showed good performance, there are still some limitations and further improvements are needed in the future. Fewer microbe–disease pairs have been confirmed to be associated, and the similarity information about diseases and microbes is not diverse enough, which seriously affects the predictive performance of the model. In future work, we will collect more correlation as well as similarity information and combine the advantages of the existing models to construct models with stronger predictive ability. In addition, considering the high application value of genetic information, the introduction of host genetic information is bound to greatly increase the prediction performance of the model. 58 , 59 , 60 Moreover, after more and more potential microbe–disease associations have been predicted and confirmed, we can further predict microbe–drug associations based on microbe–disease association information and other related data, which is conducive to providing new strategies for drug design and disease treatment.

AUTHOR CONTRIBUTIONS

Ziwei Chen: Conceptualization (equal); methodology (equal); supervision (equal); writing – review and editing (equal). Liangzhe Zhang: Investigation (equal); methodology (equal); validation (equal); visualization (equal); writing – original draft (equal). Jingyi Li: Formal analysis (equal); methodology (equal); software (equal); validation (equal). Hang Chen: Supervision (equal); writing – review and editing (equal).

CONFLICT OF INTEREST STATEMENT

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

DATA AVAILABILITY STATEMENT

The original contribution from the study is included in the article, further inquiries can be directed to the corresponding author.
==== Refs
REFERENCES

1 Sommer F , Backhed F . The gut microbiota—masters of host development and physiology. Nat Rev Microbiol. 2013;11 (4 ):227‐238. doi:10.1038/nrmicro2974 23435359
2 Mihai P , Koren S , Liu B , Sommer DD . A framework for human microbiome research. Nature. 2012;486 (7402 ):215‐221.22699610
3 Holmes E , Wijeyesekera A , Taylor‐Robinson SD , Nicholson JK . The promise of metabolic phenotyping in gastroenterology and hepatology. Nature Rev Gastroenterol Hepatol. 2015;12 (8 ):458‐471. doi:10.1038/nrgastro.2015.114 26194948
4 Bouskra D , Brézillon C , Bérard M , et al. Lymphoid tissue genesis induced by commensals through NOD1 regulates intestinal homeostasis. Nature. 2008;456 (7221 ):507‐510. doi:10.1038/nature07450 18987631
5 Gollwitzer ES , Saglani S , Trompette A , et al. Lung microbiota promotes tolerance to allergens in neonates via PD‐L1. Nat Med. 2014;20 (6 ):642‐647. doi:10.1038/nm.3568 24813249
6 Colombo APV , Boches SK , Cotton SL , et al. Comparisons of subgingival microbial profiles of refractory periodontitis, severe periodontitis, and periodontal health using the human oral microbe identification microarray. J Periodontol. 2009;80 (9 ):1421‐1432.19722792
7 Jorth P , Turner KH , Gumus P , Nizam N , Buduneli N , Whiteley MJM . Metatranscriptomics of the human oral microbiome during health and disease. MBio. 2014;5 (2 ):e01012‐e01014. doi:10.1128/mbio.01012-14 24692635
8 Huang H‐C , Lee IJ , Huang C , Chang T‐M . Lactic acid bacteria and lactic acid for skin health and melanogenesis inhibition. Curr Pharmaceut Biotechn. 2020;21 (7 ):566‐577.
9 Wang R , Wang T , Zhuo L , et al. Diff‐AMP: tailored designed antimicrobial peptide framework with all‐in‐one generation, identification, prediction and optimization. Brief Bioinform. 2024;25 (2 ):bbae078.38446739
10 Wei J , Li Z , Zhuo L , et al. Enhancing drug‐food interaction prediction with precision representations through multilevel self‐supervised learning. Comput Biol Med. 2024;171 :108104.38335821
11 Zhou Z , du Z , Jiang X , et al. GAM‐MDR: probing miRNA–drug resistance using a graph autoencoder based on random path masking. Brief Funct Genomics. 2024;23 :elae005.
12 Wang J , Zhang L , Sun J , et al. Predicting drug‐induced liver injury using graph attention mechanism and molecular fingerprints. Methods. 2024;221 :18‐26. doi:10.1016/j.ymeth.2023.11.014 38040204
13 Zhu F , Shuai Z , Lu Y , et al. oBABC: a one‐dimensional binary artificial bee colony algorithm for binary optimization. Swarm Evol Comput. 2024;87 :101567. doi:10.1016/j.swevo.2024.101567
14 Zhao Y , Wang C‐C , Chen X . Microbes and complex diseases: from experimental results to computational models. Brief Bioinform. 2021;22 (3 ):bbaa158.32766753
15 Chen X , Huang Y‐A , You Z‐H , Yan G‐Y , Wang X‐SJB . A novel approach based on KATZ measure to predict associations of human microbiota with non‐infectious diseases. Bioinformatics. 2017;33 (5 ):733‐739.28025197
16 Huang Z‐A , Chen X , Zhu Z , et al. PBHMDA: path‐based human microbe‐disease association prediction. Front Microbiol. 2017;8 :233.28275370
17 Long Y , Luo J . WMGHMDA: a novel weighted meta‐graph‐based model for predicting human microbe‐disease association on heterogeneous information network. BMC Bioinform. 2019;20 :1‐18.
18 Xu D , Xu H , Zhang Y , Wang M , Chen W , Gao R . MDAKRLS: predicting human microbe‐disease association based on Kronecker regularized least squares and similarities. J Transl Med. 2021;19 :1‐12.33397399
19 Bao W , Jiang Z , Huang D‐SJB b . Novel human microbe‐disease association prediction using network consistency projection. BMC Bioinform. 2017;18 :173‐181.
20 Huang Y‐A , You Z‐H , Chen X , Huang Z‐A , Zhang S , Yan G‐Y . Prediction of microbe–disease association from the integration of neighbor and graph with collaborative recommendation model. J Transl Med. 2017;15 :1‐11.28049494
21 Wu C , Gao R , Zhang D , Han S , Zhang Y . PRWHMDA: human microbe‐disease association prediction by random walk on the heterogeneous network with PSO. Int J Biol Sci. 2018;14 (8 ):849‐857.29989079
22 Yan C , Duan G , Wu F‐X , Pan Y , Wang J . BRWMDA: predicting microbe‐disease associations based on similarities and bi‐random walk on disease and microbe networks. IEEE/ACM Trans Comput Biol Bioinform. 2019;17 (5 ):1595‐1604.30932846
23 Zhou Z , Zhuo L , Fu X , Zou Q . Joint deep autoencoder and subgraph augmentation for inferring microbial responses to drugs. Brief Bioinform. 2024;25 (1 ):bbad483.
24 Wang T , Li Z , Zhuo L , Chen Y , Fu X , Zou Q . MS‐BACL: enhancing metabolic stability prediction through bond graph augmentation and contrastive learning. Brief Bioinform. 2024;25 (3 ):bbae127.38555479
25 Zhou Z , Zhuo L , Fu X , Lv J , Zou Q , Qi R . Joint masking and self‐supervised strategies for inferring small molecule‐miRNA associations. Mol Ther‐Nucl Acids. 2024;35 (1 ):102103.
26 Chen Z , Zhang L , Sun J , Meng R , Yin S , Zhao Q . DCAMCP: a deep learning model based on capsule network and attention mechanism for molecular carcinogenicity prediction. J Cell Mol Med. 2023;27 (20 ):3117‐3126. doi:10.1111/jcmm.17889 37525507
27 Zhu F , Niu Q , Li X , Zhao Q , Su H , Shuai J . FM‐FCN: a neural network with filtering modules for accurate vital signs extraction. Research (Wash DC). 2024;7 :0361. doi:10.34133/research.0361
28 Gao H , Sun J , Wang Y , et al. Predicting metabolite‐disease associations based on auto‐encoder and non‐negative matrix factorization. Brief Bioinform. 2023;24 (5 ):1‐13. doi:10.1093/bib/bbad259
29 Peng L‐H , Yin J , Zhou L , Liu M‐X , Zhao Y . Human microbe‐disease association prediction based on adaptive boosting. Front Microbiol. 2018;9 :2440.30356751
30 Li H , Wang Y , Zhang Z , et al. Identifying microbe‐disease association based on a novel back‐propagation neural network model. IEEE/ACM Trans Comput Biol Bioinform. 2020;18 (6 ):2502‐2513.
31 Hua M , Yu S , Liu T , Yang X , Wang H . MVGCNMDA: multi‐view graph augmentation convolutional network for uncovering disease‐related microbes. Interdiscip Sci Comput Life Sci. 2022;14 (3 ):669‐682.
32 Chen Y , Lei XJF i M . Metapath aggregated graph neural network and tripartite heterogeneous networks for microbe‐disease prediction. Front Microbiol. 2022;13 :919380.35711758
33 Gönen M . Predicting drug–target interactions from chemical and genomic kernels using Bayesian matrix factorization. Bioinformatics. 2012;28 (18 ):2304‐2310.22730431
34 He B‐S , Peng L‐H , Li Z . Human microbe‐disease association prediction with graph regularized non‐negative matrix factorization. Front Microbiol. 2018;9 :419734.
35 Boutsidis C , Gallopoulos E . SVD based initialization: a head start for nonnegative matrix factorization. Pattern Recogn. 2008;41 (4 ):1350‐1362.
36 Pock T , Sabach S . Inertial proximal alternating linearized minimization (iPALM) for nonconvex and nonsmooth problems. SIAM J Imag Sci. 2016;9 (4 ):1756‐1787.
37 Ma W , Zhang L , Zeng P , et al. An analysis of human microbe–disease associations. Briefings Bioinform. 2017;18 (1 ):85‐97.
38 Janssens Y , Nielandt J , Bronselaer A , et al. Disbiome database: linking the microbiome to disease. BMC Microbiol. 2018;18 :1‐6.29433435
39 Sun F , Sun J , Zhao Q . A deep learning method for predicting metabolite‐disease associations via graph neural network. Brief Bioinform. 2022;23, no. 4 :1‐11. doi:10.1093/bib/bbac266
40 Liu Y , Wang S‐L , Zhang J‐F . Prediction of microbe–disease associations by graph regularized non‐negative matrix factorization. J Comput Biol. 2018;25 (12 ):1385‐1394.
41 Cai D , He X , Han J , Huang TS . Graph regularized nonnegative matrix factorization for data representation. IEEE Trans Pattern Anal Mach Intell. 2010;33 (8 ):1548‐1560.21173440
42 Wang S‐H , Zhao Y , Wang CC , et al. RFEM: a framework for essential microRNA identification in mice based on rotation forest and multiple feature fusion. Comput Biol Med. 2024;171 :108177.38422957
43 Zhang X , Liu M , Li Z , Zhuo L , Fu X , Zou Q . Fusion of multi‐source relationships and topology to infer lncRNA‐protein interactions. Mol Therapy‐Nucl Acids. 2024;35 (2 ):102187.
44 Shang F , Jiao LC , Wang F . Graph dual regularization non‐negative matrix factorization for co‐clustering. Pattern Recogn. 2012;45 (6 ):2237‐2250. doi:10.1016/j.patcog.2011.12.015
45 Zhang J , Xie M . Graph regularized non‐negative matrix factorization with [Formula: see text] norm regularization terms for drug‐target interactions prediction. BMC Bioinform. 2023;24 (1 ):375. doi:10.1186/s12859-023-05496-6
46 Bolte J , Sabach S , Teboulle M . Proximal alternating linearized minimization for nonconvex and nonsmooth problems. Math Program. 2014;146 (1 ):459‐494.
47 Polyak BT . Some methods of speeding up the convergence of iteration methods. USSR Comput Math Math Phys. 1964;4 (5 ):1‐17.
48 Alvarez F , Attouch H . An inertial proximal method for maximal monotone operators via discretization of a nonlinear oscillator with damping. Set‐Valued Anal. 2001;9 :3‐11.
49 Zhang J , Xie M . Graph regularized non‐negative matrix factorization with L 2, 1 norm regularization terms for drug–target interactions prediction. BMC Bioinform. 2023;24 (1 ):375.
50 Wang F , Huang ZA , Chen X , et al. LRLSHMDA: Laplacian regularized least squares for human microbe–disease association prediction. Sci Rep. 2017;7 (1 ):7601.28790448
51 Luo J , Long Y . NTSHMDA: prediction of human microbe‐disease association based on random walk by integrating network topological similarity. IEEE/ACM Trans Comput Biol Bioinform. 2018;17 (4 ):1341‐1351.30489271
52 Zou S , Zhang J , Zhang Z . A novel approach for predicting microbe‐disease associations by bi‐random walk on the heterogeneous network. PLoS ONE. 2017;12 (9 ):e0184394.28880967
53 Siegel RL , Miller KD , Goding Sauer A , et al. Colorectal cancer statistics, 2020. CA Cancer J Clin. 2020;70 (3 ):145‐164.32133645
54 Geier DA , Geier MR . Colon cancer risk following intestinal Clostridioides difficile infection: a longitudinal cohort study. J Clin Med Res. 2023;15 (6 ):310.37434772
55 Ralser A , Dietl A , Jarosch S , et al. Helicobacter pylori promotes colorectal carcinogenesis by deregulating intestinal immunity and inducing a mucus‐degrading microbiota signature. Gut. 2023;72 (7 ):1258‐1270.37015754
56 Chen X , Xie D , Wang L , Zhao Q , You Z‐H , Liu H . BNPMDA: bipartite network projection for MiRNA–disease association prediction. Bioinformatics. 2018;34 (18 ):3178‐3186.29701758
57 Walters WA , Xu Z , Knight R . Meta‐analyses of human gut microbes associated with obesity and IBD. FEBS Lett. 2014;588 (22 ):4223‐4233.25307765
58 Meng R , Yin S , Sun J , Hu H , Zhao Q . scAAGA: single cell data analysis framework using asymmetric autoencoder with gene attention. Comput Biol Med. 2023;165 :107414. doi:10.1016/j.compbiomed.2023.107414 37660567
59 Wang T , Sun J , Zhao Q . Investigating cardiotoxicity related with hERG channel blockers using molecular fingerprints and graph attention mechanism. Comput Biol Med. 2023;153 :106464. doi:10.1016/j.compbiomed.2022.106464 36584603
60 Zhao J , Sun J , Shuai SC , Zhao Q , Shuai J . Predicting potential interactions between lncRNAs and proteins via combined graph auto‐encoder methods. Brief Bioinform. 2023;24 (1 ):1‐9. doi:10.1093/bib/bbac527
