
==== Front
eBioMedicine
EBioMedicine
eBioMedicine
2352-3964
Elsevier

S2352-3964(24)00341-4
10.1016/j.ebiom.2024.105305
105305
Articles
A multi-modal framework improves prediction of tissue-specific gene expression from a surrogate tissue
Xu Yue ab
He Chunfeng ab
Fan Jiayao ab
Zhou Yuan ab
Cheng Chunxiao ab
Meng Ran a
Cui Ya c
Li Wei c
Gamazon Eric R. eric.gamazon@vumc.org
de∗∗
Zhou Dan danzhou@zju.edu.cn
ab∗
a School of Public Health and the Second Affiliated Hospital, Zhejiang University School of Medicine, Hangzhou, China
b The Key Laboratory of Intelligent Preventive Medicine of Zhejiang Province, Hangzhou, Zhejiang, China
c Division of Computational Biomedicine, Department of Biological Chemistry, School of Medicine, University of California, Irvine, CA, 92697, USA
d Vanderbit Genetics Institute, Vanderbilt University Medical Center, Nashville, TN, USA
e Data Science Institute, Vanderbilt University Medical Center, Nashville, TN, USA
∗ Corresponding author. School of Public Health and the Second Affiliated Hospital, Zhejiang University School of Medicine, Hangzhou, China. danzhou@zju.edu.cn
∗∗ Corresponding author. Vanderbit Genetics Institute, Vanderbilt University Medical Center, Nashville, TN, USA. eric.gamazon@vumc.org
23 8 2024
9 2024
23 8 2024
107 10530513 2 2024
8 8 2024
9 8 2024
© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Summary

Background

Tissue-specific analysis of the transcriptome is critical to elucidating the molecular basis of complex traits, but central tissues are often not accessible. We propose a methodology, Multi-mOdal-based framework to bridge the Transcriptome between PEripheral and Central tissues (MOTPEC).

Methods

Multi-modal regulatory elements in peripheral blood are incorporated as features for gene expression prediction in 48 central tissues. To demonstrate the utility, we apply it to the identification of BMI-associated genes and compare the tissue-specific results with those derived directly from surrogate blood.

Findings

MOTPEC models demonstrate superior performance compared with both baseline models in blood and existing models across the 48 central tissues. We identify a set of BMI-associated genes using the central tissue MOTPEC-predicted transcriptome data. The MOTPEC-based differential gene expression (DGE) analysis of BMI in the central tissues (including brain caudate basal ganglia and visceral omentum adipose tissue) identifies 378 genes overlapping the results from a TWAS of BMI, while only 162 overlapping genes are identified using gene expression in blood. Cellular perturbation analysis further supports the utility of MOTPEC for identifying trait-associated gene sets and narrowing the effect size divergence between peripheral blood and central tissues.

Interpretation

The MOTPEC framework improves the gene expression prediction accuracy for central tissues and enhances the identification of tissue-specific trait-associated genes.

Funding

This research is supported by the 10.13039/501100001809 National Natural Science Foundation of China 82204118 (D.Z.), the seed funding of the Key Laboratory of Intelligent Preventive Medicine of Zhejiang Province (2020E10004), the 10.13039/100000002 National Institutes of Health (NIH) Genomic Innovator Award R35HG010718 (E.R.G.), 10.13039/100000002 NIH /10.13039/100000051 NHGRI R01HG011138 (E.R.G.), 10.13039/100000002 NIH /10.13039/100000049 NIA R56AG068026 (E.R.G.), NIH Office of the Director U24OD035523 (E.R.G.), and NIH/NIGMS R01GM140287 (E.R.G.).

Keywords

Gene expression
TWAS
Cross-tissue prediction
Multi-omics
BMI
==== Body
pmc Research in context

Evidence before this study

Studies have investigated prediction of gene expression in central tissues using peripheral blood samples. Diverse approaches, including linear-model-based (TEEBoT), Bayesian-ridge-regression-based (B-GEX), and deep-learning-based (MTM) methods, show promise by integrating various features. However, on the one hand, improvements in prediction accuracy can be achieved by incorporating additional features. On the other hand, the practical utility needs to be validated through applications, particularly those integrating genetics-based strategies, which previous studies have not sufficiently addressed.

Added value of this study

We proposed an approach, MOTPEC, which outperforms the baseline and existing approaches. To illustrate the utility, we implemented MOTPEC-based differential gene expression (DGE) analysis and genetics-anchored transcriptome-wide association study (TWAS), identifying numerous BMI-associated genes. The results showed that MOTPEC significantly narrowed the gap between peripheral blood and central tissues.

Implications of all the available evidence

Integrating multi-modal features improves cross-tissue prediction, which can be applied in DGE analysis. A peripheral tissue-based prediction approach that utilizes multi-modal features substantially reduces the divergence between the direct use of peripheral blood and the use of central tissues for identifying trait-associated tissue-specific genes. Combining the advantages of TWAS and DGE approaches is promising for identifying genes for complex traits.

Introduction

As a fundamental tool in precision medicine, omics is widely used across diverse research areas and increasingly applied to clinical practice.1, 2, 3 The transcriptome is the most widely utilized omics layer for deciphering the molecular circuitries of complex traits.4, 5, 6 However, gene expression is highly cell type or tissue-specific.7 Empirical data has shown that gene expression may substantially differ between disease-related tissues, henceforth referred to as “central” tissues, and surrogate “peripheral” tissues such as blood. However, most central tissues are not easily accessible without the use of invasive procedures, which limit our ability to investigate causal genes in the tissues that are critical for many complex traits. High-performance prediction of tissue-specific gene expression and enhanced detection of causal gene effects may help to address this challenge.

Common methods for identifying complex trait-associated genes include differential gene expression (DGE)8 analysis, transcriptome-wide association study (TWAS),9, 10, 11 and gene similarity-based methods.12 DGE analysis is powerful and can be conducted even with small sample sizes. However, it may be error-prone in determining the direction of effect, especially within a cross-sectional design.13 In addition, results from DGE are vulnerable to confounding. The limited availability of central tissues poses a major challenge to DGE analysis. Based on genetically-regulated gene expression, TWAS addresses the concern of reverse causality but requires a large sample size. Additionally, by the nature of genetic variants, TWAS is less susceptible to certain types of confounding. However, the complexity of local linkage disequilibrium and widespread pleiotropic effects reduce the resolution of TWAS for causal gene identification.11 Combining the advantages of TWAS and DGE approaches may be promising for identifying complex trait-associated genes. To this end, obtaining reliable predicted tissue-specific gene expression for DGE analysis may increase robustness for causal gene mapping in inaccessible central tissues.

Some studies have sought to predict gene expression in central tissues using peripheral blood.14, 15, 16, 17, 18 Existing approaches include TEEBoT (Tissue Expression Estimation using Blood Transcriptome),14 B-GEX (BayR-based multi-tissue Gene EXpression),15 and MTM (Multi-tissue Transcriptome Mapping).18 The primary model in TEEBoT integrates gene expression and splicing data using LASSO regression14 while B-GEX leverages Bayesian ridge regression and considers only whole-blood gene expression.15 MTM is a unified deep learning-based multi-task learning framework, enabling tissue-specific gene expression prediction from any available tissue.18 Although these approaches have demonstrated some success in predicting central tissue gene expression for a large number of genes by using multiple features derived from blood, there is room for substantial improvement in prediction quality. We hypothesize that a multi-modal approach that integrates a more comprehensive set of regulatory features, prior biological knowledge, and connectivity information among genes in a surrogate tissue can substantially enhance prediction performance in less accessible tissues, including disease-relevant ones.

Methods

Datasets

Genotype-Tissue Expression (GTEx) data were used to generate prediction models in multiple tissues. GTEx19 includes 15201 RNA-seq samples from peripheral blood and 48 central tissues obtained from 838 donors. These donors are mostly of European American ancestry (85.3%) and aged between 50 and 70 (68.1%), with a small proportion being African. Gene expression (inverse-normal-transformed), APA events,20 splicing profiles, and genetic variants from GTEx were included as features. BMI from GTEx was utilized as phenotype data to perform DGE.21 The scripts are available at Github (https://github.com/zdangm/MOTPEC).

Co-expression modules

To characterize the co-expression modules in blood, co-expression network analysis (WGCNA) was implemented with default settings.22 Briefly, a soft threshold was chosen by considering scale independence and the mean connectivity (Supplementary Figure S8). The co-expression network was constructed to identify co-expression modules. Principal component analysis (PCA) was used to reduce the dimensionality of each co-expression module. For gene expression prediction model training, the top 20 PCs of the module that involve the target gene were included as features.

Gene expression model

We chose the LASSO method to build the prediction model, focusing on generalizability and accuracy.23 For each gene, we assumed a linear model of gene expression in each of the target tissues using molecular profiles from whole blood.(1) Model:Yg=βg0+βg∗Xg

In Equation (1), Yg is the gene expression of the target gene in a given central tissue, Xg is the blood-based molecular profile matrix for the target gene, βg0 is the intercept, and βg is the coefficient matrix. For the local features, we restricted to 1 Mb and 100 Kb on both sides of the gene of interest for genetic variants and splicing profiles, respectively. The (blood-derived) splicing profiles used for the gene expression prediction in a given central tissue consisted of the “fully processed, filtered and normalized splice phenotype” data from the GTEx portal. To reduce the dimensionality and the complexity of the high-dimensional molecular features, PCA was performed. We included the top 10 PCs for gene expression across the transcriptome, the top 20 PCs for APA events, the top 20 PCs for local genetic variants, and up to 20 PCs for splicing profiles. In addition, for relevant non-local information, we prioritized biologically informative features and potentially shared regulatory structures by including the top 10 PCs of the expression profile of transcription factors and the top 20 PCs of the co-expression module that includes the target gene.

Prediction model training

Here, we systematically evaluated a variety of prediction modeling approaches, including LASSO, elastic net, BSLMM, and random forest. LASSO, elastic net, and BSLMM outperformed random forest, while BSLMM was time consuming. Considering the principle of parsimony, we chose LASSO as the regression model, which fits our case well, both theoretically and empirically.

A penalized regression model was applied for the gene expression prediction. Equation (2) shows the loss function L(β) of the LASSO regression model24:(2) L(β)=∑i=1N(yi−yˆi)2+λ∑i=1N|βi|

To prioritize the expression level of the target gene in blood, we further included it as a feature with no penalty in the loss function. The hyper-parameter λ was selected from a default sequence, with each candidate λ evaluated under 5-fold cross-validation. Finally, λ with minimum mean cross-validated error was chosen, and the coefficient of the prediction model with this λ was used.

Prediction performance evaluation

Here, we assumed a situation in which only blood samples were collected while we were interested in identifying trait-associated genes in central tissues. Consequently, the use of blood as a proxy tissue (i.e., directly using the measured expression in blood as a proxy for the expression in the tissue of interest) was considered as a baseline. The Pearson's r between baseline measured expression (in blood) and observed expression (in the central tissue of interest) was evaluated as the prediction performance baseline. The prediction accuracy of MOTPEC was evaluated by estimating the correlation between observed expression and predicted expression in each target tissue. Furthermore, we implemented three other existing methods, TEEBoT, B-GEX and MTM, and estimated their prediction accuracy using the same settings, including the same samples on the training set (70% of the GTEx samples) and the testing set. We used the suggested parameters provided by the authors (learning rate = 0.0005, β1 = 0.5, β2 = 0.9, training epochs = 200),18 with the exception of setting the batch size to half of the train-set samples. For MTM, whole blood was used as source tissue and the other tissues as target tissues. For a fair comparison, only genes with adequate features for these four methods were used.

UMAP analysis

Uniform manifold approximation and projection (UMAP) is a nonparametric graph-based dimensionality reduction algorithm,25,26 which has been commonly used for single-cell data analysis (with known limitations).27,28 In this study, we used UMAP to reduce the dimensionality and to visualize the similarity among expression profiles across different tissues. Genes expressed in all 49 tissues were included (15,637 in total). We also utilized UMAP to show the similarity among the DGE coefficients of gene expression on the BMI phenotype in the central tissues.

Integration of DGE and TWAS to identify BMI-associated genes

DGE analysis using linear regression was implemented (3) to detect associations between genes and traits.(3) BMI=β0+β1GeneExp+∑kβ2,kCovariatesk

In equation (3), age, gender, and ancestry (converted to a binary variable) were used as covariates. The regression coefficient β1 and its significance were used as the indicator of an association between expression and BMI.

TWAS is a framework for identifying trait-associated genes based on genetically regulated expression (GReX).29 Here, prediction models were trained based on GTEx v8 samples. For each gene, the best-performing prediction model, according to the prediction accuracy estimated by five-fold cross-validation, among PrediXcan,9 UTMOST,30 and JTI,10 was applied. A summary statistics-based association test31 was applied to estimate the association between GReX and BMI. TWAS significance was defined by PFDR < 0.05, and the direction of effect was determined by z-score.

Genes that show significance in both TWAS and DGE analysis with a consistent direction of effect were defined as concordant genes.

Cellular perturbation analysis using the connectivity map

To identify potential cellular perturbations and related compounds for the concordant gene list, we queried the CMap dataset. Connectivity Map (CMap)32 is a large-scale collection of functional perturbations in cultured human cells funded by the LINCS program33 (Library of Integrated Network-Based Cellular Signatures). CMap catalogs the cellular response to chemical or genetic perturbations. Given a list of genes associated with a biological state (e.g., genes associated with obesity), we estimated the similarity between the query (the gene list) and each perturbation in the CMap database based on a nonparametric, rank-based Kolmogorov–Smirnov statistic.34 Prioritized by similarity level, perturbations with strong connections may illustrate the underlying mechanisms of the studied biological state. In this study, we generated queries using the concordant genes identified by measured, MOTPEC-based, and blood expression, respectively. Applying the clue query platform (https://clue.io/query), we interrogated the Touchstone dataset in CMap, which is described by gene expression and well-annotated.

Ethics

The empirical data used in this study is publicly available. The study design was approved by the Medical Ethics Committee of the School of Public Health, Zhejiang University (ZGL202312-7).

Statistics

The statistical approaches are included in previous sections.

Role of funders

The funders played no role in the study design, data collection, data analysis, data interpretation, or paper writing.

Results

Overview of MOTPEC

We propose an approach named MOTPEC (a Multi-mOdal-based framework to bridge the Transcriptome between PEripheral and Central tissues). MOTPEC integrates a range of blood-based gene expression regulatory features while leveraging co-regulation patterns and potential biological connections as prior knowledge, to enhance prediction accuracy. In addition, we illustrate the practical application of integrating MOTPEC-based DGE with TWAS towards enhanced identification of genes associated with body mass index (BMI).

Fig. 1 provides an overview of MOTPEC. Leveraging a diverse set of features (derived from GTEx v8), the model was trained to predict tissue-specific gene expression in 48 central tissues. The predicted expression can be used in place of observed expression to identify trait-associated genes in unavailable tissues. Local genetic variants, gene expression and splicing profile data, and alternative polyadenylation (APA) events from blood were considered as features to train the prediction model. Additionally, transcription factors and genes sharing the same co-expression module with the predicted gene (i.e., the target gene) were included as features for model training. Principal component (PC) analysis was conducted to reduce the dimensionality of the expression profile of the transcription factors. A total of 27 co-expression modules were identified in whole blood using WGCNA with default settings (Methods). The module size varied from 37 to 7434 genes. The model was trained using LASSO. The prediction performance was estimated using five-fold cross-validation.Fig. 1 The overview of the framework. We train a model to predict tissue-specific gene expression for central tissues using local genetic variants, gene expression and splicing data, and APA events in blood as features, and using transcription factor regulation, co-expression modules, and potential biological connections as prior knowledge for model training.

Prediction performance

The measured gene expression in blood was considered as the “baseline”. In other words, it serves as a surrogate when we wish to study the association between gene expression in a central tissue and a phenotype when only blood samples are available for analysis. Here, we compared the prediction performance between the baseline model and the MOTPEC prediction model. We observed a significantly higher Pearson's correlation r between the MOTPEC-predicted gene expression and the measured gene expression in the target tissue, compared to the correlation between the baseline and the measured gene expression (Fig. 2a). Moreover, we compared the overall prediction performance (again by Pearson's correlation r) among MOTPEC, TEEBoT, B-GEX, and MTM in the test set, which included 30% samples of the dataset. To ensure a fair comparison, these four tools (MOTPEC, B-GEX, TEEBoT, and MTM) were implemented in the same settings, and only genes with adequate features for these four models were used. MOTPEC outperformed TEEBoT (P = 6.7e-13), B-GEX (P < 2.2e-16), and MTM (P < 2.2e-16) in 48 non-peripheral tissues (Fig. 2b and c). The tissue-specific results were consistent with the overall comparison (Supplementary Figure S1 and Supplementary Table S1).Fig. 2 MOTPEC improves prediction of gene expression in central tissues using blood-derived features. For each panel in (a), the violin plot shows the Pearson's correlation r between baseline (whole blood, blue)/MOTPEC predicted (whole blood, brick red) expression and measured expression in each of the target tissues. The comparison of four methods (MOTPEC, TEEBoT, B-GEX, and MTM) in prediction performance. The Pearson's correlation r indicates the prediction performance in the test set in all tissues (b) and brain tissues (c).

To explore the contribution of each added class of regulatory features to enhanced prediction accuracy, we performed a leave-one-out analysis and built reduced models by removing APA events, transcriptional factors, and co-expression modules one by one, then compared the prediction accuracy of these reduced models with that of full model. Focusing on the tissues that passed nominal significance (paired t-test P < 0.05), we found that APA events, transcriptional factors, and co-expression modules enhanced prediction accuracy in 22 (45.8%), 42 (87.5%), 46 (95.8%) tissues, respectively (Supplementary Table S7), highlighting the conserved role of co-expression regulation across peripheral and non-peripheral tissues. To explore the characteristics of predictable tissue-specific genes, we performed Gene Ontology enrichment analysis for those highly predictable genes (defined as Pearson's r exceeding 0.3). As shown in Supplementary Figure S2, the highly predictable genes were enriched in fundamental cellular processes, including metabolic processes, RNA processing, translation, and transcription. Various tissues exhibited enrichment in specific processes. For instance, predictable genes in lower leg skin (sun exposed) tissue showed enrichment in ‘cellular response to external stimulus’. Predictable genes in small intestine terminal ileum showed enrichment in certain metabolic processes. Predictable genes in liver, pancreas, and subcutaneous adipose tissues exhibited enrichment in small molecule catabolic processes.

MOTPEC narrows the gap between peripheral blood and central tissues

To illustrate the tissue specificity of genes expressed in blood, we visualized the transcriptome profiles across the 49 (GTEx v8) tissues by UMAP. Blood (marked in pink), located at the bottom right corner, is the most distal from most of the remaining tissues (Fig. 3), emphasizing its specificity. We further illustrated the tissue specificity in DGE analysis by identifying BMI-associated genes as an example. We regressed the BMI on measured gene expression, with age, sex, and ancestry as covariates, to estimate the DGE coefficient on BMI for each gene in blood and each of the 48 central tissues. We reduced the dimensionality of the DGE coefficients across the tested genes in 49 tissues and visualized them in a two-dimensional space (Fig. 4). The dark yellow, dark blue, and pink circles denote brain, non-brain, and blood, respectively. Overall, the distance between blood and the remaining tissues (Fig. 3, Fig. 4) indicated that the use of blood as a proxy could be problematic.Fig. 3 Sample and tissue similarity based on gene expression profiles. Each point represents a tissue-specific sample, and each colour represents a tissue. The horizontal and vertical coordinates are the two UMAP components.

Fig. 4 MOTPEC reduces the difference in estimated DGE effect size between blood and central tissues. Tissue similarity based on the estimated effect size of gene expression on BMI from DGE analysis. The first two UMAP components were used to perform hierarchical clustering to visualize the estimated DGE coefficients (a vector across the transcriptome) in brain tissues (a) and non-brain tissues (b) from the effect of gene expression on BMI. Blue, yellow, and pink denote non-brain central tissues, brain central tissues, and peripheral blood, respectively. The dark colour denotes estimated effect size using measured gene expression while the light colour denotes estimated effect size using MOTPEC-based predicted expression. Each number represents a tissue. (c) visualizes the first two UMAP components by scatter plot. Each circle represents a tissue, marked with a different number. The Euclidean distance between the estimated DGE coefficient of blood-to-measured and MOTPEC-predicted-to-measured in brain tissues (d) and non-brain tissues (e) are shown. The tissue colour is consistent with Fig. 3.

We asked whether MOTPEC could reduce the gap between peripheral blood and central tissues. We generated MOTPEC-predicted gene expression for each of the 48 central tissues and estimated the (effect size) coefficient for MOTPEC-predicted gene expression on BMI with the same covariates adjusted for. We analyzed the coefficients estimated using MOTPEC for brain and non-brain tissues (Fig. 4), respectively. The MOTPEC-predicted gene expression and the measured gene expression clustered together in multiple brain and non-brain tissues (i.e., brain caudate basal ganglia, brain cerebellar hemisphere, brain cerebellum, brain cortex denoted by 9, 10, 11, 12 in Fig. 4a; cells cultured fibroblasts, lung, skin sun exposed lower leg denoted by 21, 32, 41 in Fig. 4b). In the majority of the tissues (i.e., adipose visceral omentum, adrenal gland, brain cerebellar hemisphere, brain cerebellum, brain substantia nigra, esophagus gastroesophageal junction, and skeletal muscle denoted by 2, 3, 10, 11, 19, 25, and 34, respectively in Fig. 4c), the MOTPEC-predicted gene expression (light circle) is located closer to the measured gene expression (dark circle) than blood (azaleine). The difference in the estimated DGE effect sizes between blood and central tissues was substantially narrowed by the use of MOTPEC-predicted expression (e.g., adipose subcutaneous, brain amygdala denoted by 1, 7). We compared the Euclidean distance in the coefficients between blood-to-measured and MOTPEC-predicted-to-measured. Notably, the Euclidean distance significantly dropped in brain tissues (paired t-test P = 0.027; Fig. 4d) and non-brain tissues (paired t-test P = 2.83×10−9; Fig. 4e). In addition, we compared MOTPEC with a statistically comparable approach TEEBoT, which is the second-best approach in the context of expression prediction accuracy, for reducing the distance for gene–trait associations. The results showed that MOTPEC (which significantly reduced the distance in 32 tissues) outperformed TEEBoT (in 26 tissues) for narrowing the gap for gene-BMI associations. Consistent results were observed for common binary traits, including T2D (25 tissues v.s. 19 tissues).

To investigate the balance between the prediction accuracy and the number of predictable genes, we systematically evaluated the relationship as the Pearson's correlation r threshold was varied from 0 to 0.7. As illustrated in Supplementary Figure S5, the number of predictable genes drops, as expected, when we increase the threshold for r. In practice, we recommend r ≥ 0.3 as the threshold for the correlation between predicted and measured gene expression, which results in approximately 5000 predictable genes across the 48 tissues. Notably, the threshold is consistent with TEEBoT.14 In Fig. 4, it is worth mentioning, as a limitation, that the two-dimensional UMAP figure may not accurately represent the original distances between the tissues in a high-dimensional space.

To explore the robustness of the MOTPEC framework, we selected six disease phenotypes (MHHTN-hypertension, MHT2D-type 2 diabetes, MHHRTATT-acute myocardial infarction, MHHRTDIS-ischemic heart disease, MHCOPD-chronic respiratory disease, MHASTHMA-asthma) in addition to BMI. Similar to BMI, we conducted DEG analysis for these six binary diseases using peripheral blood gene expression, central tissue measured gene expression, and MOTPEC-predicted gene expression, respectively. Subsequently, the gaps among the DGE results based on the three types of gene expression in each disease were estimated. We found that the gaps in the gene–trait association between using peripheral blood and using non-peripheral tissues were significantly reduced by utilizing MOTPEC-predicted expression. Taking chronic respiratory disease for an example, as shown in Supplementary Figure S3, the gaps for MOTPEC-predicted-to-measured were smaller than those for blood-to-measured in 11 (84.6%) brain tissues and 29 (82.9%) non-brain tissues. Collectively, these findings show the broad utility of MOTPEC. The visualization of the 2D projection of DGE coefficients is provided in Supplementary Figure S3 and the Euclidean distance is provided in Supplementary Figure S4.

To further investigate the utility of MOTPEC, the highly predictable genes (Pearson's r > 0.3) were used to predict disease states, as had been done for TEEBoT. We found that the predicted expression was almost as good as observed expression and outperformed blood expression in MHHTN (paired t-test P = 4.0e-10), MHT2D (paired t-test P = 5.9e-4) and MHHRTDIS (paired t-test P = 2.1e-2) (Supplementary Figure S10).

In the context of DGE analysis for BMI and six diseases, our results demonstrate that the use of gene expression in blood as a proxy tissue can show a large bias in the estimated effect size relative to the gene expression measured in the actual tissues of interest while MOTPEC can narrow this estimated effect size divergence between blood and central tissues.

Integrating MOTPEC-based DGE analysis with TWAS

TWAS is characterized by a relatively high-degree of evidence for potential causality while DGE analysis has a higher statistical power to identify associations. A method that integrates TWAS and MOTPEC-based DGE analysis may benefit from these two methods. Here, we illustrate the utility of the framework by integrating TWAS and MOTPEC-based DGE analysis to identify BMI-associated genes.

For DGE analysis, we considered three approaches: (1) DGE using blood as the baseline, (2) DGE using measured gene expression in central tissues as the ideal condition, and (3) MOTPEC-based DGE analysis. The DGE results are summarized in Supplementary Table S3.

Leveraging the large-scale GWAS for BMI (Methods), TWAS was performed to identify genes with genetically regulated expression associated with BMI. We then asked to what extent the DGE genes were enriched in genes identified by TWAS. For each of the 49 tissues, this enrichment analysis was performed (Methods). Significant enrichment was observed in 12 of the 49 tissues (PFDR < 0.05). The significance and the OR for the 12 tissues are displayed in Supplementary Figure S6. Notably, brain caudate basal ganglia and adipose visceral omentum are included in the 12 tissues.

We defined “concordant genes” as those showing significance in both TWAS and DGE with the same direction of effect. As expected, DGE analysis using measured expression resulted in the larger number of concordant genes (Fig. 5a, 3760 genes across the 48 central tissues, i.e., 3760 gene–tissue pairs). The number of concordant genes was 2612 and 2523 for MOTPEC-predicted and blood (Fig. 5a), respectively. We found only 162 overlapping concordant genes between measured expression and blood expression (Fig. 5c). Notably, MOTPEC substantially increased the number of overlapping concordant genes to 378, showing its ability to narrow the gap between the baseline condition (use of blood as a proxy) and the optimal condition (use of measured expression in available tissue) (Fig. 5b). The number of overlapping concordant genes across the 48 central tissues is shown in Fig. 5d and detailed in Supplementary Material S4.Fig. 5 TWAS and DGE concordant genes for BMI. A concordant gene is defined as a gene showing significance in both TWAS and DGE with consistent direction of effect (Methods). (a) The total number of concordant BMI-associated genes across all tissues identified by measured, predicted, and blood expression, respectively are shown. The Venn diagram plot shows the number of overlapping genes between (b) measured and MOTPEC-predicted and (c) measured and blood. (d) The number of overlapping concordant genes between MOTPEC-predicted (blue)/blood expression (red) and measured expression in each tissue.

We visualized the concordant genes in the Manhattan plot for two obesity-relevant tissues, namely visceral adipose tissue35 and basal ganglia tissue,36 which showed significant enrichment (Supplementary Figure S1). The TWAS and DGE results from the measured expression, MOTPEC-predicted expression, and the measured expression in blood, are summarized in panels a and d, panels b and e, and panels c and f, respectively, of Fig. 6. Notably, 242 and 101 concordant genes (marked as solid diamonds) were identified using the directly measured expression in visceral adipose and basal ganglia tissue, respectively (Fig. 6a and d). We found that the use of MOTPEC-predicted expression uncovered more concordant genes overlapping with the use of measured expression (Fig. 6a and d) than with the use of blood (27 v.s. 12 for visceral adipose, 23 v.s. 5 for basal ganglia, the overlapped genes are annotated). Among the identified genes in 48 tissues, some (including ADCY3, PHIP, LEPR) have already been proven to be linked to obesity-related traits.37 These results show that, compared with the use of blood as a surrogate tissue, MOTPEC-based prediction provides a better approach for identifying trait-associated genes.Fig. 6 Manhattan plots for the concordant genes identified by MOTPEC-based DGE and TWAS for BMI. For illustration, visceral omentum adipose (panels a, b, and c) and brain caudate basal ganglia (panels d, e, and f) are presented. The Manhattan plots show the concordant genes with the TWAS results for BMI. Each diamond denotes a gene whose genetically regulated expression was found to be associated with BMI (PFDR < 0.05). Among these TWAS significant genes, genes attaining significance in DGE (P < 0.05) analysis with the same direction of effect are defined as concordant genes (denoted by solid diamond). Ideally, DGE would be performed using measured expression in the tissue of interest (panels a and d). As a baseline proxy, blood-based DGE identified only a small number of concordant genes (panels c and f, overlapping concordant genes are labeled) that overlap with the concordant genes (solid diamond) in panels a and d. MOTPEC-based DGE increased the number of concordant genes (panels b and e, overlapping concordant genes are labeled) overlapping the concordant genes identified by measured expression in target tissues.

To illustrate the utility of utilizing TWAS and DGE jointly, we considered a BMI-associated gene, LEPR, which was identified by both TWAS and DGE, for example. Similar to typical TWAS signals, we observed multiple signals in the 1 Mb flanking region of LEPR in a brain tissue. Remarkably, LEPR showed a modest signal compared with the leading signal at DNAJC6 (Fig. 7a). TWAS itself had a hard time further prioritizing the genes in this region, while the MOTPEC-based DGE analysis further supported the association between LEPR and BMI (Fig. 7b). LEPR encodes leptin receptor, which is crucial for regulating body weight by maintaining energy balance while DNAJC6 encodes a protein involved in the regulation of clathrin-mediated endocytosis. Obviously, LEPR is more likely to be the causal gene in this region. In short, TWAS and DGE may complement each other to prioritize trait-associated genes.Fig. 7 The regional Manhattan plot for TWAS and DGE in the LEPR region. For the genes located within 1 Mb region of LEPR, significance of TWAS and of MOTPEC-based DGE analysis are shown in panel a and b, respectively. The red line denotes nominal significance level. Blue diamond and red circle denote genes with and without significant associations.

Simulations for illustration of risk of overfitting

Since we used GTEx for both prediction model training and gene–trait association analysis, we performed simulations to quantify the potential risk of overfitting.

As shown in Supplementary Figure S9 panel a, we initialized the measured expression for blood by assuming a normal distribution. Then we generated the measured expression in a target tissue (e.g., brain caudate basal ganglia) based on the coefficients we learned from the real data (i.e., GTEx). The phenotype (e.g., BMI) was generated based on the generated measured expression in the target tissue and the coefficients estimated from the real data (i.e., GTEx). In addition, we introduced noise drawn from a standard normal distribution ∼N (0, 1) to the effect sizes (for both blood on expression and expression on BMI). The dataset we generated, as described, is considered the “pseudo GTEx” dataset. We then generated a second dataset using the same procedure and a different seed. The second dataset is considered the “pseudo external dataset”. We note that both datasets have the same ground-truth coefficients, which were derived from empirical data.

With the two simulated datasets (“pseudo GTEx” and “pseudo external dataset”), we were able to estimate the potential risk of overfitting. We estimated the effect-size coefficients of predicted expression on the phenotype in two ways, namely the “within pseudo GTEx” estimation and the “external” estimation. For the “within pseudo GTEx” estimation, we trained the prediction model for the target tissue in the “pseudo GTEx” dataset and estimated the association between predicted expression and phenotype, i.e., consistent with what we had done in the real GTEx dataset. For the “external” estimation, we applied the model (trained in the “pseudo GTEx” dataset) into the “pseudo external dataset” and estimated the association between predicted expression and phenotype. As shown in Supplementary Figure S9 panel b, the estimated “within pseudo GTEx” coefficients did not show substantial departure (without noise: correlation test r = 0.94, P < 2.2e-16; with noise: correlation test r = 0.91, P < 2.2e-16) from the “external” estimation, indicating low overfitting and approximately unbiased estimation.

Cellular perturbation analysis further supports MOTPEC-DGE-TWAS concordant genes

Ideally, concordant genes from the integration of DGE (using measured expression in the tissue of interest) and TWAS could be used to identify trait-related cellular perturbations that show a concordant or divergent transcriptional profile for the list of genes. To investigate whether MOTPEC-based DGE could do a similar job, we queried the CMap dataset with three concordant gene lists for BMI. The three gene lists were generated by overlapping TWAS and DGE results using measured, MOTPEC-predicted, and blood expression, respectively.

Due to the minimum number of input genes, only 5 of the 15 tissues of interest (2 adipose tissues and 13 brain tissues) were applied to obtain the connectivity scores between the input concordant genes and cellular perturbations. We filtered out the perturbations without specific mechanisms of action and empirically considered the top 20 positive connectivity and the top 20 negative connectivity (Supplementary Table S6) as candidates. Two perturbations (namely PD-184352 and GR-235) were identified by both measured and MOTPEC-predicted expression integration analysis in adipose visceral omentum tissue with a positive connectivity score consistent with our TWAS and DGE findings. However, no consistent perturbations were found between measured expression and blood proxy-based DGE (Fig. 8a).Fig. 8 Integrating MOTPEC-DGE with TWAS identifies drug candidates for obesity. The Venn diagram showed the overlap among the perturbations identified using measured, MOTPEC-based, and blood expression as input queries. (a) upper: Measured v.s. MOTPEC-based; lower: Measured v.s. blood. Two perturbation-related compounds PD-184352 and GR-235 were identified by both measured and MOTPEC-predicted analysis (A upper) in adipose visceral omentum tissue. (b) The hypothetical model of TNF-α regulation of CIDEC was proposed by a previous study.38 On the one hand, elevated TNF-α activates the MEK/ERK pathway, causing phosphorylation of PPARγ and downregulation of CIDEC. On the other hand, activated MEK interacts with PPARγ, causing PPARγ export from the nucleus to cytosol, which further downregulates CIDEC. In summary, suppressing the MEK could upregulate CIDEC and decrease lipolysis.

Interestingly, both compounds have a promoting effect on obesity, which suggests a consistent biological effect for the query. PD-184352 is a MEK inhibitor. MEK inhibitor could prevent TNF-α-mediated CIDEC (Cell death-inducing DFF45-like effector C) downregulation (Fig. 8b).38 Specifically expressed in adipose tissues (Supplementary Figure S7), CIDEC promotes triglyceride accumulation.38 GR-235 is a FXR antagonist. Existing research has suggested that intestinal FXR agonism can promote adipose tissue browning and reduce obesity,39 although contradictory results have been reported.40,41 From the perspective of identifying cellular biological connections, the investigation showed that integrating MOTPEC-based DGE with TWAS outperforms just using blood as a proxy tissue.

Discussion

Mapping disease-associated genes in relevant tissues is a fundamental step to elucidating disease mechanisms. To address the challenge of understanding the etiologies of complex traits due to the limited availability of central tissues, we introduce MOTPEC, a multi-modal-based framework that improves the accuracy of gene expression prediction in central tissues by incorporating comprehensive molecular markers and biological knowledge obtained from blood. MOTPEC outperforms both existing models and baseline blood-derived models. Application of MOTPEC to the identification of BMI-associated genes illustrates its superior performance compared to blood-based models. The framework narrows the gap between the use of blood as a proxy tissue and the use of the tissue of interest for identifying trait-associated genes.

As a tissue-specific expression-predicting approach, TEEBoT integrates information at the transcriptome level using a linear model. B-GEX optimizes the model using Bayesian ridge regression, which outperforms least square regression, ridge regression, and LASSO regression. MTM enables tissue-specific gene expression prediction from any available tissue using a deep learning (nonlinear) approach. However, some potentially informative regulatory features, as well as prior knowledge of gene transcriptional regulation and the gene connectivity, are not considered in these approaches. MOTPEC improves prediction accuracy by the use of prior knowledge of both global and local regulation,42 incorporating the connections from co-expression modules and leveraging additional regulatory features such as APA events.

We illustrated the utility of the MOTPEC framework using several approaches. Taking BMI as the trait of interest, we showed that MOTPEC significantly reduced the divergence in estimated differential expression effect sizes between the use of blood and the use of central tissues. In addition, we incorporated the results from TWAS for comparison. By defining “concordant genes” as those exhibiting significance in both TWAS and DGE with the same effect direction, we observed only 162 overlapping concordant genes between measured expression and blood while MOTPEC substantially increased this overlap to 378, thus reducing the divergence between blood-based analysis and the optimal scenario of measured expression.

Besides comparing the number of overlapping concordant genes, we note that integrating TWAS and DGE is a promising way to identify potentially causal genes. TWAS is a powerful approach for the discovery of trait-associated genes. However, the widespread pleiotropic effects of genetic variants in linkage disequilibrium (LD) with the variants in the TWAS prediction models may produce false positives.10,11 For example, TWAS may implicate multiple genes in a GWAS locus, providing limited resolution of causal gene mapping. Notably, the false positives may result from the fact that the genetically regulated expression (GReX, the key element of TWAS) is purely a combination of SNPs. However, in the MOTPEC model, the genetic information is only one source of the predicted expression. Thus, the model may not suffer to the same degree from the limitations of TWAS, such as the potential for implicating multiple (false positive) genes in a locus. In other words, MOTPEC-based DGE could be informative to further improve prioritization of TWAS results given the reduced impact of local linkage disequilibrium contamination and widespread pleiotropic effects. Thus, “concordant genes” identified by both DGE and TWAS may benefit from the strengths of both TWAS and DGE. However, the “non-concordant genes” are non-negligible because it is not necessary for all causal genes to show signals for both TWAS and DGE. For example, a low-heritability causal gene is not expected to show a TWAS signal since TWAS is powered for heritable gene expression.

Our approach successfully identified BMI-associated candidate genes, including ADCY3 and PHIP, which have consistently shown associations with obesity in previous studies.37 Genome-wide association studies (GWAS) have also demonstrated a consistent association between common variants in ADCY3 as well as PHIP and BMI.43 Functional evidence has strongly supported the role of ADCY3 and PHIP as crucial mediators of energy homeostasis and promising targets for pharmacological interventions in obesity treatment.44 Specifically, animal studies have shown that selective ablation of ADCY3 in the mouse hypothalamus leads to a significant increase in body fat mass.45 Conversely, mice with a gain-of-function mutation in ADCY3 exhibited less fat accumulation and showed resistance to obesity when fed a high-fat diet.46 As indicated by cellular experiments, nuclear PHIP has been shown to enhance the transcription of pro-opiomelanocortin (POMC), a neuropeptide known to suppress appetite. This ultimately contributes to dysregulated feeding behaviour and disrupted energy metabolism.44 Notably, these two genes, ADCY3 and PHIP, were detected in our TWAS analysis of predicted expression but were missed when relying on blood measured expression. Interestingly, ADCY3 does exhibit a specific expression pattern that tends to be enriched in adipose and brain tissues (https://www.proteinatlas.org/). This finding underscores the limitation of relying solely on a proxy tissue for predicting trait-associated tissue-specific expressed genes. Furthermore, leveraging the cellular perturbation data from CMap, we identified two compounds using both MOTPEC-predicted and measured expression, underscoring the utility of MOTPEC-predicted expression in identifying trait-associated genes and related biological processes. Notably, both compounds were supported by previous studies.

In this study, we proposed a framework that significantly improved gene expression prediction in central tissues. It is worth noting that the approach can be extended to other omics data. Nevertheless, the study has clear limitations. Although GTEx has been extended to multiple omics data, including DNA methylation and protein expression, the resource is still limited by sample size, which limits further improvements in prediction accuracy.

In conclusion, we develop MOTPEC, a multi-modal gene expression modelling approach that substantially reduces the divergence between the use of peripheral blood and the use of central tissues for identifying trait-associated tissue-specific genes. The utility was demonstrated through application to the identification of BMI-associated genes and analysis of effect-size divergence for six additional (disease) phenotypes.

Contributors

D.Z. and E.R.G. designed the study. Y.X. and D.Z. wrote the manuscript. Y.Z., J.F., C.H., and C.C. provided critical comments. Y.X., C.Y., W.L., and R.M. collected and processed the data. Y.X. performed the analyses. D.Z. and E.R.G. supervised and acquired funding for the study. D.Z. and Y.Z. verified the underlying data. All authors have read and approved the final version of the manuscript.

Data sharing statement

The MOTPEC scripts are available at https://github.com/zdangm/MOTPEC.

Data from GTEx portal can be downloaded via https://www.gtexportal.org/home/and for the Giant consortium via https://portals.broadinstitute.org/collaboration/giant/index.php/GIANT_consortium.

Declaration of interests

W.L declares grants from NCI (R01CA193466 and R01CA228140). The remaining authors declare that they have no competing interests.

Appendix A Supplementary data

Supplementary Figure S1

Tissue specific comparison for prediction performance among MOTPEC, TEEBoT, BGEX, and MTM.

Supplementary Figure S2

The Gene Ontology enrichment analysis for those highly predictable genes. Each figure represents the results of enrichment analysis for a specific tissue, where the Y-axis shows the biological processes, and the X-axis represents the enrichment level.

Supplementary Figure S3

The first two UMAP components of DGE-coefficients for six diseases, namely (a) hypertension, (b) type 2 diabetes, (c) acute myocardial infarction, (d) ischemic heart disease, (e) chronic respiratory disease, and (f) asthma.

Supplementary Figure S4

The Euclidean distance between the estimated DGE coefficient of blood-to-measured and MOTPEC-predicted-to-measured in brain tissues and non-brain tissues.

Supplementary Figure S5

(a-e) Visualization of DGE-coefficients under different thresholds. (f) The median number of genes in each tissue that met the threshold under different thresholds.

Supplementary Figure S6–S10

Supplementary Tables

Acknowledgements

This research is supported by the 10.13039/501100001809 National Natural Science Foundation of China 82204118 (D.Z.) and the seed funding of the Key Laboratory of Intelligent Preventive Medicine of Zhejiang Province (2020E10004). E.R.G. wishes to acknowledge support from the National Institutes of Health (NIH) Genomic Innovator Award R35HG010718, NIH/10.13039/100000051 NHGRI R01HG011138, 10.13039/100000002 NIH /10.13039/100000049 NIA R56AG068026, 10.13039/100000052 NIH Office of the Director U24OD035523, and 10.13039/100000002 NIH /10.13039/100000057 NIGMS R01GM140287.

Appendix A Supplementary data related to this article can be found at https://doi.org/10.1016/j.ebiom.2024.105305.
==== Refs
References

1 Ibrahim R. Pasic M. Yousef G.M. Omics for personalized medicine: defining the current we swim in Expert Rev Mol Diagn 16 7 2016 719 722 26959799
2 Vassy J.L. Lautenbach D.M. McLaughinet H.M. The MedSeq Project: a randomized trial of integrating whole genome sequencing into clinical medicine Trials 15 2014 85 24645908
3 Demirel H.C. Arici M.K. Tuncbag N. Computational approaches leveraging integrated connections of multi-omic data toward clinical applications Mol Omics 18 1 2022 7 18 34734935
4 Reilly S.K. Noonan J.P. Evolution of gene regulation in humans Annu Rev Genom Hum Genet 17 2016 45 67
5 Cookson W. Liang L. Abecasis G. Moffatt M Lathrop M. Mapping complex disease traits with global gene expression Nat Rev Genet 10 3 2009 184 194 19223927
6 Lappalainen T. Sammeth M. Friedländer M.R. Transcriptome and genome sequencing uncovers functional variation in humans Nature 501 7468 2013 506 511 24037378
7 Zhang W. Voloudakis G. Rajagopal V.M. Integrative transcriptome imputation reveals tissue-specific and shared biological mechanisms mediating susceptibility to complex traits Nat Commun 10 1 2019 3834 31444360
8 Anders S. Huber W. Differential expression analysis for sequence count data Genome Biol 11 10 2010 R106 20979621
9 Gamazon E.R. Wheeler H.E. Shah K.P. A gene-based association method for mapping traits using reference transcriptome data Nat Genet 47 9 2015 1091 1098 26258848
10 Zhou D. Jiang Y. Zhong X. Cox N.J. Liu C. Gamazon E.R. A unified framework for joint-tissue transcriptome-wide association and Mendelian randomization analysis Nat Genet 52 11 2020 1239 1246 33020666
11 Wainberg M. Sinnott-Armstrong N. Mancuso N. Opportunities and challenges for transcriptome-wide association studies Nat Genet 51 4 2019 592 599 30926968
12 Weeks E.M. Jacob C.U. Nathan Y.C. Leveraging polygenic enrichments of gene features to predict genes underlying complex traits and diseases Nat Genet 55 8 2023 1267 1276 37443254
13 Porcu E. Sadler M.C. Lepik K. Differentially expressed genes reflect disease-induced rather than disease-causing changes in the transcriptome Nat Commun 12 1 2021 5647 34561431
14 Basu M. Wang K. Ruppin E. Hannenhalli S. Predicting tissue-specific gene expression from whole blood transcriptome Sci Adv 7 14 2021 7
15 Xu W.J. Liu X.S. Leng F. Li W. Blood-based multi-tissue gene expression inference with Bayesian ridge regression Bioinformatics 36 12 2020 3788 3794 32277818
16 Halloran J.W. Zhu D. Qian D.C. Prediction of the gene expression in normal lung tissue by the gene expression in blood BMC Med Genom 8 2015 77
17 Wang J. Gamazon E.R. Pierce B.L. Imputing gene expression in uncollected tissues within and beyond GTEx Am J Hum Genet 98 4 2016 697 708 27040689
18 He G. Chen M. Bian Y. Yang E. MTM: a multi-task learning framework to predict individualized tissue gene expression profiles Bioinformatics 39 6 2023 btad363
19 The GTEx Consortium atlas of genetic regulatory effects across human tissues Science 369 6509 2020 1318 1330 32913098
20 Li L. Huang K.-L. Gao Y. An atlas of alternative polyadenylation quantitative trait loci contributing to complex trait and disease heritability Nat Genet 53 7 2021 994 1005 33986536
21 Pulit S.L. Stoneman C. Morris A.P. Meta-analysis of genome-wide association studies for body fat distribution in 694 649 individuals of European ancestry Hum Mol Genet 28 1 2018 166 174
22 Langfelder P. Horvath S. WGCNA: an R package for weighted correlation network analysis BMC Bioinf 9 2008 559
23 Kim S.J. Koh K. Lustig M. Boyd S. Gorinevsky D. An interior-point method for large-scale $\ell_1$-Regularized least squares IIEEE J Sel Top Signal Process 1 4 2007 606 617
24 Tay J.K. Narasimhan B. Hastie T. Elastic net regularization paths for all generalized linear models J Stat Softw 106 2023
25 Sainburg T. McInnes L. Gentner T.Q. Parametric UMAP embeddings for representation and semisupervised learning Neural Comput 33 11 2021 2881 2907 34474477
26 McInnes L. Healy J. UMAP: uniform manifold approximation and projection for dimension reduction ArXiv 2018 10.48550/arXiv.1802.03426
27 Becht E. Mclnnes L. Healy J. Dimensionality reduction for visualizing single-cell data using UMAP Nat Biotechnol 37 1 2019 38 44
28 Narayan A. Berger B. Cho H. Assessing single-cell transcriptomic variability through density-preserving data visualization Nat Biotechnol 39 6 2021 765 774 33462509
29 Ghitza U.E. Commentary: a Gene-Based association method for mapping traits Using reference transcriptome Data Front Psychiatr 7 2016
30 Hu Y. Li M. Lu Q. A statistical framework for cross-tissue transcriptome-wide association analysis Nat Genet 51 3 2019 568 576 30804563
31 Barbeira A.N. Shah K.P. Torres J.M. MetaXcan: summary statistics based gene-level association method infers accurate PrediXcan results bioRxiv 2016 045260
32 Subramanian A. Narayan R. Corsello S.M. A next generation connectivity map: L1000 platform and the first 1,000,000 profiles Cell 171 6 2017 1437 1452.e17 29195078
33 Keenan A.B. Jenkins S.L. Jagodnik K.M. The library of integrated network-based cellular Signatures NIH program: system-level cataloging of human cells response to perturbations Cell Syst 6 1 2018 13 24 29199020
34 Subramanian A. Tamayo P. Mootha V.K. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles Proc Natl Acad Sci U S A 102 43 2005 15545 15550 16199517
35 Mathis D. Immunological goings-on in visceral adipose tissue Cell Metabol 17 6 2013 851 859
36 Tan Z. Hu Y. Ji G. Alterations in functional and structural connectivity of basal ganglia network in patients with obesity Brain Topogr 35 4 2022 453 463 35780276
37 Loos R.J.F. Yeo G.S.H. The genetics of obesity: from discovery to biology Nat Rev Genet 23 2 2022 120 133 34556834
38 Tan X.R. Cao Z.Z. Li M. Xu E.D. Wang J.J. Xiao Y.F. TNF-Α downregulates CIDEC via MEK/ERK pathway in human adipocytes Obesity 24 5 2016 1070 1080 27062372
39 Fang S. Suh J.M. Reilly S.M. Intestinal FXR agonism promotes adipose tissue browning and reduces obesity and insulin resistance Nat Med 21 2 2015 159 165 25559344
40 Li F. Jiang C. Krausz K.W. Microbiome remodelling leads to inhibition of intestinal farnesoid X receptor signalling and decreased obesity Nat Commun 4 1 2013 2384 24064762
41 Sun L. Cai J. Gonzalez F.J. The role of farnesoid X receptor in metabolic diseases, and gastrointestinal and liver cancer Nat Rev Gastroenterol Hepatol 18 5 2021 335 347 33568795
42 Aguet F. Brown A.A. Castel S.E. Genetic effects on gene expression across human tissues Nature 550 7675 2017 204 213 29022597
43 Speliotes E.K. Willer C.J. Berndt S. Association analyses of 249,796 individuals reveal 18 new loci associated with body mass index Nat Genet 42 11 2010 937 948 20935630
44 Marenne G. Hendricks A.E. Perdikari A. Exome sequencing identifies genes and gene sets contributing to severe childhood obesity, linking PHIP variants to repressed POMC transcription Cell Metab 31 6 2020 1107 1119.e12 32492392
45 Cao H. Chen X. Yang Y. Storm D.R. Disruption of type 3 adenylyl cyclase expression in the hypothalamus leads to obesity Integr Obes Diabetes 2 2 2016 225 228 27942392
46 Pitman J.L. Wheeler M.C. Lloyd D.J. Walker J.R. Glynne R.J. Gekakis N. A gain-of-function mutation in adenylate cyclase 3 protects mice from diet-induced obesity PLoS One 9 2014 e110226 10.1371/journal.pone.0110226
