
==== Front
Innovation (Camb)
Innovation (Camb)
The Innovation
2666-6758
Elsevier

S2666-6758(24)00120-6
10.1016/j.xinn.2024.100682
100682
Commentary
Interpretable foundation models as decryptors peering into the Earth system
Li Chenyu 111
Hong Danfeng 1211
Zhang Bing zb@radi.ac.cn
13∗
Liao Tianjun 4
Yokoya Naoto 5
Ghamisi Pedram 6
Chen Min 7
Wang Lizhe 8
Benediktsson Jon Atli 9
Chanussot Jocelyn 10
1 Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China
2 School of Electronic, Electrical, and Communication Engineering, University of Chinese Academy of Sciences, Beijing 100049, China
3 College of Resources and Environment, University of Chinese Academy of Sciences, Beijing 100049, China
4 Academy of Military Sciences, Beijing 100000, China
5 Graduate School of Frontier Sciences, The University of Tokyo, Chiba 277-8561, Japan
6 Helmholtz-Zentrum Dresden-Rossendorf, Helmholtz Institute Freiberg for Resource Technology, 09599 Freiberg, Germany
7 Key Laboratory of Virtual Geographic Environment (Ministry of Education of PRC), Nanjing Normal University, Nanjing 210023, China
8 School of Computer Science, China University of Geosciences, Wuhan 430078, China
9 Faculty of Electrical and Computer Engineering, University of Iceland, 102 Reykjavik, Iceland
10 University Grenoble Alpes, Inria, CNRS, Grenoble INP, LJK, 38000 Grenoble, France
∗ Corresponding author zb@radi.ac.cn
11 These authors contributed equally

05 8 2024
09 9 2024
05 8 2024
5 5 10068218 4 2024
3 8 2024
© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Published Online: August 5, 2024
==== Body
pmcThe Earth system amid the big data paradigm

The processes of the Earth system drive interactions between energy, matter, and life, and a comprehensive understanding of their full evolutionary trajectory is critical for sustainable human development. Traditional modeling primarily relies on a set of theoretical equations to simulate dynamic process such as the carbon-nitrogen cycle, solar radiation dynamics, and terrestrial ecosystem dynamics.1 Despite the extensive modeling experience of Earth scientists, the rapid advancement of Earth observation techniques has led to a significant increase in the volume of databases, with data accumulating daily or even hourly. This has exacerbated the conflict between the capacity for data collection and utilization for big Earth data. Consequently, there is an urgent need to enhance the intelligent processing and analysis of big Earth data.2

At this critical juncture, the emergence of foundational models3 has revitalized the unique advantages of maximizing information retrieval and deriving insights from big Earth data. However, the mathematical principles underpinning their success are somewhat elusive, raising concerns about trustworthiness due to the lack of a clearly defined internal chain of reasoning and decision-making processes. Therefore, interpretable foundational models are crucial. They enhance our understanding and security of geoscientific applications, break through performance limitations, and improve the controllability of their social impacts.

What is the interpretable artificial intelligence?

Interpretability refers to the extent to which humans can grasp the fundamental principles behind decision-making. It is directly proportional to our understanding of the decisions or predictions made by artificial intelligence (AI) models. Specifically, it involves addressing the issues of what, why, and how:(1) discovering what interactions drive the model’s predictions

(2) verifying why certain features are instrumental in driving model’s decision-making process

(3) assessing how the effectiveness of decisions is validated by real-world data

Interpretability venturing into foundation models

Historically, theoretical modeling and deep learning (DL) methods have been viewed as entirely independent scientific paradigms. However, there is growing evidence that these approaches are inherently complementary. Theoretical models can provide direct interpretability in practical applications, while foundation models have immense potential for making inferences that surpass human cognitive abilities.

Regarding the studied avenues, we propose a cross-driven paradigm that combines multidisciplinary theory with AI-powered technology in a synergistic collaborative system, as shown in Figure 1. First, ensuring the reliability of foundation models requires their integration with multidisciplinary theory to guide, constrain, and validate their biases. Then, utilizing the trustworthy foundation model extracts deeper, more abstract insights from big Earth data, thereby enriching the human theoretical repository.Figure 1 Interpretable foundation models adhere to theoretical principles, ensuring transparent reasoning processes, reliable decisions, and flexible correction structures

They maintain full adaptive generation in areas with sparse real-world data and extract deeper, more abstract insights from big Earth data.

Considering various application tasks and requirements, we summarize three distinct coupling strategies for interpretable foundational models.(1) Generative guidance. This paradigm typically generates complementary or missing input/output data in a serial manner with two primary objectives. The first is to enhance the generality of foundational models using theoretical constraint (such as mathematical, physical, and dynamic principles). The second is to address the issue of data sparsity in real-world scenarios by leveraging AI technology in big data mining to generate necessary data.

(2) Subproblem embedded. Also known as “plug and play,” this module primarily aims to address specific subproblems by embedding foundation models or theory modeling as proxy modules or loss functions. Its greatest advantage lies in its strong flexibility in two distinct manners. First, it constrains the inference results of DL networks by designing theoretical models as submodules or loss functions. Second, in scenarios with limited expert knowledge, it integrates aggregated DL networks as proxy functions for specific variables within optimization solvers.

(3) Principle interaction. Exploring the interpretability of AI technologies at a fundamental level reflects the latter approach, providing robust support for enhancing the human knowledge base. This involves mapping the traditional theoretical models onto DL networks at a theoretical level. In this method, the optimization process of theoretical models customizes the architecture of DL networks to ensure that each layer reflects real physical properties, thereby imparting interpretability. However, this coupling method poses significant challenges, as it requires designers to possess knowledge from both domains—traditional model-driven and data-driven approaches—that integrate physical mechanisms and principles. Despite these challenges, breakthrough results have been achieved.4,5 These studies map the optimization process of models as prior knowledge into DL networks, transforming regularization parameters, which typically require manual tuning, into learnable parameters.

The ideal way and actual gap: 3D generalizability

What are the ultimate outcomes of this synergistic collaboration system? What challenges will researchers face when pursuing this ideal? This commentary aims to reshape the existing understanding of the generalization of foundation models by crafting a 3D generalizability concept: data, discipline, and downstream.

Data assimilation: integrating and processing diverse types

Data assimilation refers to optimal combination of theoretical models and observational data to estimate the evolving state of a system over time. This involves four key aspects: (1) remote sensing (collecting data from satellite or aerial sensors), (2) observational data (gathering data from observational stations, laboratory analyses, surveys, expeditions, and field experiments; these are the observations closest to the real world, used to verify or correct errors), (3) societal sensing (deriving from interactions between human activities and the environment), and (4) modeling and reanalysis (generating data, typically due to the sparsity of observational data).

An ideal foundational model must have the capacity to assimilate big Earth data and manage the synergies between different data types, each with limited geographical applicability. Therefore, more comprehensive integration of complementary big data is crucial for achieving a more consistent representation of the Earth system across spatial, temporal, and physical processes.

Discipline crossed frontier: describing, understanding, and augmenting complex dynamic processes

Realizing the advantages of interpretable foundation models will require extensive interdisciplinary efforts, advancements in computational infrastructure, and the cultivation of an innovative and open big data culture. Processes within the Earth system are extremely complex, necessitating the integration of mathematical statistics and modeling, physical modeling, and computer programming in the design and implementation of coupling methodologies. A critical obstacle is bridging the knowledge gap between members from different specialties, particularly those specializing in Earth science or computer science.

Downstream sustainability: pre-trained intrinsic features for energy-efficient computing

Generalizability should not only be understood as adaptability to multiple downstream tasks but also as prioritizing the impact on computational resource consumption and greenhouse gas emissions. Here, we advocate for the development of general tools that align with the goals of reducing greenhouse gas emissions and promoting environmental sustainability. This can be achieved by combining knowledge and skills from multiple disciplines through efficient computational practices, the use of renewable energy, and the optimization of foundational models.

Conclusion

Recently, foundation models have shown remarkable performance in intelligent processing, effectively addressing the challenges posed by the rapid growth of big data. Unfortunately, focusing solely on input and output data without ensuring transparency in the inference process and reliability of results has led to biases and skepticism toward AI technologies, hindering their advancement. Here, we strongly advocate for the establishment of a trustworthy AI system. Data-driven methods in Earth science research should not replace theoretical modeling but, rather, complement and enrich each other. Ultimately, this approach will ensure adherence to theoretical inferences with clear and interpretable structures, thereby fostering deeper insights and understanding in Earth science research.

Acknowledgments

This work was supported by the 10.13039/501100001809 National Natural Science Foundation of China under grant 42271350 and grant 42241109 .

Declaration of interests

The authors declare no competing interests.
==== Refs
References

1 Chen M. Qian Z. Boers N. Iterative integration of deep learning in hybrid earth surface system modelling Nat. Rev. Earth Environ. 4 2023 568 581 10.1038/s43017-023-00452-7
2 Hong D. Li C. Zhang B. Multimodal artificial intelligence foundation models: Unleashing the power of remote sensing big data in earth observation The Innovation Geoscience 2 1 2024 100055 10.59717/j.xinn-geo.2024.100055
3 Hong D. Zhang B. Li X. Spectralgpt: Spectral remote sensing foundation model IEEE Trans. Pattern Anal. Mach. Intell. 46 8 2024 5227 5244 10.1109/TPAMI.2024.3362475 38568772
4 Li C. Zhang B. Hong D. Lrr-net: An interpretable deep unfolding network for hyperspectral anomaly detection IEEE Trans. Geosci. Rem. Sens. 61 2023 1 12 10.1109/TGRS.2023.3279834
5 Li C. Zhang B. Hong D. Learning disentangled priors for hyperspectral anomaly detection: A coupling model-driven and data-driven paradigm IEEE Transact. Neural Networks Learn. Syst. 2024 10.1109/TNNLS.2024.3401589
