
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

72667
10.1038/s41598-024-72667-7
Article
Self-evolving artificial intelligence framework to better decipher short-term large earthquakes
Cho In Ho icho@iastate.edu

Chapagain Ashish
https://ror.org/04rswrd78 grid.34421.30 0000 0004 1936 7312 CCEE Department, Iowa State University, Ames, IA 50011 USA
20 9 2024
20 9 2024
2024
14 2193426 7 2024
9 9 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Large earthquakes (EQs) occur at surprising loci and timing, and their descriptions remain a long-standing enigma. Finding answers by traditional approaches or recently emerging machine learning (ML)-driven approaches is formidably difficult due to data scarcity, interwoven multiple physics, and absent first principles. This paper develops a novel artificial intelligence (AI) framework that can transform raw observational EQ data into ML-friendly new features via basic physics and mathematics and that can self-evolve in a direction to better reproduce short-term large EQs. An advanced reinforcement learning (RL) architecture is placed at the highest level to achieve self-evolution. It incorporates transparent ML models to reproduce magnitude and spatial location of large EQs (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w \ge $$\end{document} 6.5) weeks before of the failure. Verifications with 40-year EQs in the western U.S. and comparisons against a popular EQ forecasting method are promising. This work will add a new dimension of AI technologies to large EQ research. The developed AI framework will help establish a new database of all EQs in terms of ML-friendly new features and continue to self-evolve in a direction of better reproducing large EQs.

Subject terms

Natural hazards
Computational science
http://dx.doi.org/10.13039/100000001 National Science Foundation CMMI-2129796 issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Predicting large earthquakes (EQs) within a practical time window and at a specific location remains a long-standing enigma. A quote from1 well reflects our limit—“Short-term deterministic earthquake prediction remains elusive and is perhaps impossible...” Despite notable advances in research communities2–16, the capability of the existing methods for predicting “individual” large EQ’s specific location and magnitude is limited17,18. Recently, geophysics communities actively utilize advanced machine learning (ML) methods19–27, advancing our knowledge about EQs and improving EQ forecasting. However, there is also a doubt whether the sophistication of ML methods can necessarily increase prediction accuracy28,29. In broad computational science, recent rule-learning ML methods30–36 show some promises, but purely EQ data-oriented and ML-driven exploration of hidden mechanisms is in its infancy. When it comes to large EQs, obtaining reliable, informative large data sets is formidably difficult, and internal data are intrinsically inaccessible due to various technological limits. We still lack precise descriptions inside the earth’s lithosphere before large EQs. Fundamental limitations regarding large EQs (i.e., lack of large quality data, multi-physics interactions, and unknown principles) hinder the direct adoption of existing ML methods as recently done in geophysics—e.g., deep learning for EQ focal mechanism37, logistic regression for EQ-induced landslide susceptibility mapping38 and extreme gradient boosting for explosion/mining-induced EQ classification39. The proposed artificial intelligence (AI) framework seeks to generate ML-friendly new data and let ML-driven prediction rules gradually evolve, for which the reinforcement learning (RL) offers the well-established architecture. The proposed AI framework will be inclusive in a sense that promising ML models (e.g.,37–39 can be incorporated as a prediction rule and their evolution can be managed by the RL architecture.

To add a new research dimension to existing attempts, this paper develops a novel AI framework with self-evolving capability that can transform raw observational EQ data into ML-friendly features and that can evolve in a direction to better reproduce short-term large EQs. Contrary to existing EQ forecasting/prediction methods, this paper focuses on deterministic reproduction of large EQs weeks before the failures. Figure 1B compares the proposed framework against the existing EQ forecasting approaches13–17. Although the ranges and scopes of EQ forecasting approaches are broad, and some overlaps may exist2, Fig. 1B intentionally draws the separation line to emphasize the notable differences of the proposed framework from the forecastings—i.e., this framework pursues (i) deterministic reproductions (not probabilistic average rate of EQs) (ii) of large events with \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w>$$\end{document} 6.5 (not wide ranges of magnitudes) (iii) in terms of specific loci and magnitudes of the peaks (not probabilistic counts in discrete spatial cells and subdivided magnitudes bins), (iv) primarily a few weeks or a month before the main EQ (not including the long-term events). Undoubtedly, EQ forecasting13–17 plays an important role in understanding and mitigating probabilistic seismic hazards based on robust statistics and models with long histories (e.g.,40–42 and international collaborations (e.g., Collaboratory for the Study of Earthquake Predictability). To some extent, Fig. 1B illustrates how the proposed framework will contribute to the existing EQ forecasting research paradigm by providing additional angles with new AI-driven technologies. This paper hypothesizes that available EQ catalogs conceal rules behind the occurrence of large EQs, which are hitherto unknown but can be learned by a fusion of data, geophysics, and ML. The author’s initial works offer proofs-of-concept that the hypothesis appears promising43,44. The proposed framework transforms decades-long EQ catalogs into ML-friendly new features via multi-layered data transformation by leveraging basic physics and mathematics44. Then, the transparent ML can distinguish and remember individual large EQs with the aid of the new features, which can be supported by the “uniqueness” of the new features of individual large EQs43. This framework’s prediction uses the threefold objective, magnitudes, three-dimensional (3D) loci, and short-term timing of large EQs, pursuing deterministic, short-term reproductions44. This framework seeks self-evolving capability by harnessing RL architecture at the highest level of the framework so that consistent improvement can be done autonomously by RL with new data (Fig. 1). This paper presents the architecture and primary cores of the self-evolving AI framework for short-term large EQs. After initial training with the past 40 years of EQ data in the western U.S. region, this paper conducted pseudo-prospective short-term predictions/reprouctions of large EQs. Comparison results support the promising performance of this framework compared to UCERF3-ETAS, one of the well-established EQ forecasting methods.

Results

Seismogenesis agent and environment

Figure 1 (A) The highest level analogy between the real-world seismogenesis and the proposed reinforcement learning (RL); (B) Conceptual difference between the EQ forecasting methods and the EQ prediction used in this paper. (C) Overall architecture of the proposed AI framework with reinforcement learning (RL) schemes for autonomous improvement in reproduction of large EQs: Core (1) transforms EQs data into ML-friendly new features (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal{I}\mathcal{I}$$\end{document}); Core (2) uses (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal{I}\mathcal{I}$$\end{document}) to generate a variety of new features in terms of pseudo physics (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {U}$$\end{document}), Gauss curvatures (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbb {K}$$\end{document}), and Fourier transform (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbb {F}$$\end{document}), giving rise to unique “state” \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S \in \mathcal {S}$$\end{document} at t; Core (3) selects “action” (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$A \in \mathcal {A}$$\end{document})—a prediction rule—according to the best-so-far “policy” \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\pi \in \mathcal {P}$$\end{document}; Core (4) calculates “reward” (R)—accumulated prediction accuracy; Core (5) improves the previous policy \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\pi $$\end{document}; Core (6) finds a better prediction rule (action), enriching the action set \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {A}$$\end{document}.

There are two analogous entities between the real-world seismogenesis and the proposed RL, i.e., seismogenesis agent and environment (Fig. 1A). In general RL, given the present state (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{(t)}$$\end{document}), an “agent” takes actions (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$A^{(t)}$$\end{document}), obtains reward (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$R^{(t+1)}$$\end{document}), and affects the environment leading to next state (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{(t+1)}$$\end{document}). These terminologies are defined in typical RL method and their analogies to the present framework are summarized in Fig. S1. The “environment” resides outside of the agent, interacts with the agent, and generates the states. In this study, a “seismogenesis agent” is introduced as the virtual entity that can take actions according to a hidden policy and determines all future EQs. The seismogenesis agent’s choice of action constantly determines the next EQ’s location, magnitude, and timing given the present state and thereby affects the surrounding environment. In this study, “environment” is assumed to include all geophysical phenomena and conditions in the lithosphere and the Earth, e.g., plate motions, strain and stress accumulation processes near/on faults, and so on. We can consider the Markov decision process (MDP) in terms of a sequence (or trajectory) and the so-called four argument probability \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$p:\mathcal {S}\times \mathcal {R}\times \mathcal {S}\times \mathcal {A}\rightarrow \mathbb {R}[0,1]$$\end{document} as1 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\begin{aligned} & S^{(0)}, A^{(0)}, R^{(1)}, S^{(1)}, A^{(1)}, R^{(2)},...,S^{(t)}, A^{(t)}, R^{(t+1)} (t\rightarrow \infty ) \end{aligned}$$\end{document}

2 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\begin{aligned} & \quad p(S', R|S, A):= \text {Prob}\left( S'=S^{(t+1)}, R=R^{(t+1)}|S=S^{(t)}, A=A^{(t)}\right) \end{aligned}$$\end{document}

According to the “Markov property”45, the present state \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{(t)}$$\end{document} is assumed to contain sufficient and complete information that can lead to the accurate prediction of the next state \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{(t+1)}$$\end{document} as would be done by using the full past history. As shall be described in detail, the reward is rooted in the prediction accuracy of future EQs in terms of loci, magnitudes, and timing, and thus the true seismogenesis agent can always receive the maximum possible reward since it always reproduces the real EQ. The “true” seismogenesis agent remains unknown and so does the policy. The primary goal of the proposed RL is to learn the hidden policy and to evolve the agent so that it can imitate the true seismogenesis agent as closely as possible. Thus, as evolution continues, the RL’s agent pursues the higher reward (i.e., the long-term return) and gradually becomes the true seismogenesis agent. There are notable benefits of the adopted RL. First, with incoming new data, improving the new features and transparent ML methods will be autonomously done by RL, without human interventions. Second, RL will search and allow many pairs of state-action (i.e., current observation-future prediction), and thus researchers may obtain with uncertainty measures many prediction rules that are “customized” to different locations, rather than a single unified prediction model. In essence, the proposed global RL framework will act as a virtual scientist who keeps improving many pre-defined control parameters of ML methods, keeps searching for a better prediction rule, and keeps customizing the best prediction rule for each location and new time.

Information fusion core

As shown in Fig. 1C, the framework’s first core is the information fusion core that can transform raw EQ catalog data into ML-friendly new features—convolved information index (II) which can quantify and integrate the spatio-temporal information of all past EQs. The term “convolved” II is used since this core mainly utilizes a sort of convolution process in the spatial and temporal domains of the observed raw EQ data in the lithosphere. By extending convolution to the time domain, this core quantifies and incorporates cumulative information from past EQs, creating spatio-temporal convolved II ( \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\overline{II}_{ST}^{(t)}$$\end{document}), which is denoted as “New Feature 1” in this framework. Detailed derivation procedure is given in Supplementary Materials and Table 1.

State generation core

The second core of the framework is to generate and determine “states” (Fig. 1C). For RL to understand, distinguish, and remember all the different EQs across the large spatio-temporal domains, the state should be based on physics and unique signatures about spatio-temporal information of past EQs. Also, the effective state should be helpful in pinpointing specific location’s past, present and future events. To meet these criteria, this core defines “point-wise state” at each spatial location \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\xi _i$$\end{document} in terms of the pseudo physics (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {U}$$\end{document}), the Gauss curvatures of the pseudo physics (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbb {K}$$\end{document}), and the Fourier transform-based features of the Gauss curvatures (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbb {F}$$\end{document}) as defined in Eq. (3). Since \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$n(\mathcal {U})=4, n(\mathbb {K})=8,$$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$n(\mathbb {F})=160$$\end{document}, a state vector \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S_j$$\end{document} has \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$n(S_j) = 172.$$\end{document} “New Feature 2” is defined as the high-dimensional, pseudo physics quantities (denoted as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {U}$$\end{document}) using the spatio-temporal convolved IIs (i.e., New Feature 1) of the information fusion core. “New Feature 3” is defined as the Gauss curvature-based signatures (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbb {K}$$\end{document}) that consider distributions of the pseudo physics quantities \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {U}$$\end{document} at each depth as surfaces. “New Feature 4” is defined as the Fourier transform-based signatures (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbb {F}$$\end{document}) that quantify the time-varying nature of the Gauss curvature-based signatures \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbb {K}$$\end{document}. Table 1 summarizes the key equations and formulas of the four-layer data transformations. With these multi-layered new features we can define “state” in the RL context as:3 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\begin{aligned} {\textbf { STATE SET }}\mathcal {S}:=\{S| S=(\mathcal {U}, \mathbb {K}, \mathbb {F}), S\in \mathbb {R}^{(n(\mathcal {U})+n(\mathbb {K})+n(\mathbb {F}))} \} \end{aligned}$$\end{document}

The primary hypothesis is that the “states” can be used as unique indicators of individual large EQs before the events, in hopes of enabling ML methods to distinguish and remember EQs. This core in essence infuses basic physics terms into new features. It should be noted that these physics-infused quantities are all “pseudo” quantities, derived from data, not from any known first principles. The author’s prior works43,44 fully describe the generations of all the New Features. Its compact summary is presented in Supplementary Materials.

Prediction core

The main objective of the Prediction Core (Fig. 1C) is to select the best “prior” action (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$A^*\in \mathcal {A}^{(t)}$$\end{document}) out of the up-to-date entire action set \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {A}^{(t)}$$\end{document} by using the policy (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\pi $$\end{document}) given the present state (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{(t)}$$\end{document}). The best action predicts the location and magnitudes of the EQ before the event (e.g., 30 days or a week ahead). This core leverages the so-called Q-value function approximation45. Q-value function (denoted \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$Q_\pi (S,A)$$\end{document}) is the collection of future returns (e.g., the accumulated prediction accuracy) of all state-action pairs following the given policy. In this paper, Q-value function stores the error-based reward, i.e., the smaller error, the higher return. One difficulty is the fact that the space of the state set \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {S}$$\end{document} is a high-dimensional continuous domain, unlike the discrete action space. The adopted tabular Q-value function contains “discrete” state-action pairs. To resolve the difficulty and to leverage the efficiency of the tabular Q-value function, it is important to determine whether the present state \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{(t)}$$\end{document} is similar to or different from all the existing states \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ S^{(\forall \tau < t)} \in \mathcal {S}$$\end{document}. For this purpose, the adopted policy uses the L2 norm (i.e., Euclidean distance) between states—e.g., \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{*}(S^{(t)}) :=\text {argmin}_{S^{(\tau )}} \Vert S^{(\tau )} - S^{(t)} \Vert _2, \forall \tau < t.$$\end{document} The L2 norm includes the state’s \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {U}$$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \mathbb {K}$$\end{document} but excludes \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbb {F}$$\end{document}, for calculation brevity. Preliminary simulations confirm that the use of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {U}$$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \mathbb {K}$$\end{document} appears to work favorably in finding the “closest” prior state \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^*(S^{(t)})$$\end{document} to the present state \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{(t)}$$\end{document}. For the policy \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\pi $$\end{document} (i.e., the probability to select an action for the given state), this core leverages the so-called “greedy policy” in which the policy chooses the action that is associated with the maximum Q-value. Details about the formal expressions of the adopted policy and specialized schemes for this policy are presented in Supplementary Materials.Figure 2 Internal searching process with the tabular Q-value functions used for the pseudo-prospective prediction by this framework 14 days before 2019/7/6 (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w 7.1$$\end{document}) Ridgecrest EQ: (A) The entire state sets stored in memory; (B) A tabular Q-value function identified to contain the closest state of the potential peaks. Vertical axis shows the reward of each state-action pair and outstanding spikes stand for the best actions of corresponding states. The last action is the dummy action that has the dummy reward as given by Eq. (14); (C) One specific Q-value function of actions and the state that is considered to be the potential peak. This process is used for all the pseudo-prospective predictions; (D) General illustration of a Jacob’s ladder for large EQ reproduction consisting of the state set with increasingly many pseudo physics terms and the action set (inspired by the Jacob’s ladder within the density functional theory46). Raw EQs data are complex multi-physics vectors in space and time.

Action update core

Action update core seeks to find new action \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$A^{*(t)}$$\end{document} (i.e., prediction rules) when there are expanded states \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{(t)} \in \varvec{S}^{(ID)}, ID \ge 18$$\end{document}. This new action will be different from existing actions (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$A^{*(t)} \ne \forall A_i\in \mathcal {A}^{(t-1)}$$\end{document}) already used in the Prediction Core. Since the state \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{(t)}$$\end{document} is assumed to be “unique,” this core seeks to find a “customized” prediction rule (action) for each new state. In general, there is no restriction of the new action (i.e., new prediction rule/model). In the present framework, the Action Update Core leverages the author’s existing work, Glass-Box Physics Rule Learner (GPRL), to find the new best action (GPRL’s full details are available in43). In essence, GPRL is built upon two pillars—flexible link functions (LFs) for exploring general rule expressions and the Bayesian evolutionary algorithm for free parameter searching. LF (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {L}$$\end{document}) can take (i) a simple two-parameter exponential form or (ii) cubic regression spline (CRS)-based flexible form. As shown in34–36, the exponential LF is useful when a physical rule of interest is likely monotonically increasing or decreasing with concave or convex shapes. In this framework, the pseudo released energy \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$E_r$$\end{document} takes the exponential LF due to the favorable performance43. Also, GPRL can utilize CRS-based LFs for higher flexibility47, which is effective when shapes of the target physics rules are highly nonlinear or complex. In this framework, the actions (prediction rules) take CRS-based LFs for generality and accuracy. The space of LFs’ parameters (denoted as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\varvec{\uptheta }$$\end{document}) is vast. Finding “best-so-far” free parameters \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\varvec{\uptheta }^*$$\end{document} is done by the Bayesian evolutionary algorithm in this framework, as successfully done by34–36. Table 2 presents the salient steps of the Bayesian evolutionary algorithm. All the new features (e.g., \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\overline{II}, \mathcal {U}, \mathbb {K}, \mathbb {F}$$\end{document}) generated by the State Generation Core are used and explored as potential candidates in the prediction rules (i.e., actions). The best-so-far prediction rule’s expression identified by this framework is presented in Supplementary Materials. As shown by44, the inclusion of the Fourier-transform-based new features \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbb {F}$$\end{document} appears to sharpen the accuracy of each prediction rule for large EQs (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w \ge 6.5$$\end{document}). It is noteworthy that the action (prediction rule) shares the same spirit as the so-called Jacob’s ladder within the density functional theory46 in a sense that EQ reproduction accuracy gradually improves with the addition of more physics and mathematical terms (see Fig. 2D). Since the developed AI framework holds the self-evolving capability, whenever new large EQs occur, the old sets (State and Action) and Q-value functions should expand autonomously. The details about autonomous expansion of state, action, and policy are presented in Supplementary Materials.Figure 3 Pseudo-Prospective Predictions by UCERF3-ETAS and Reproductions by the Present AI framework 14 days before 2019/7/6 (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w 7.1$$\end{document}) Ridgecrest EQ: (A) Real Observed EQ; (B) This AI framework’s best-so-far reproduction by the pre-trained RL policy; (C) ComCat probability spatial distribution from UCERF3-ETAS predictions of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M\ge 5$$\end{document} events with probability of 0.3% within 1 month time window; (D) ComCat magnitude-time functions plot, i.e., magnitude versus time probability function since prediction simulation starts. Time 0 means the prediction starting time. Probabilities above the minimum simulated magnitude (2.5) are shown. Favorable Evolution of New Action Learning: (E) Distribution of the entire rewards of “new” best actions given a new state set. The maximum reward \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {max}[r(S,A)]$$\end{document} of the best-so-far prediction of the peak increases to 67.13 from 47.60 in (F) which shows the distribution of the entire rewards of “old” best actions given a new state set; Predicted magnitudes by using the new best action (G) whereas (I) is using the old best action. Compared to the real EQs (H), the new best action appears to favorably evolve (E,F).

Feasibility test results of pseudo-prospective short-term predictions

After training with all EQ catalog during the past 40 years from 1990 through 2019 in the western U.S. region (i.e., longitude in (− 132.5, − 110) [deg], latitude in (30, 52.5) [deg], and depth (− 5, 20) [km]), this paper applied the trained AI framework to large EQs (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w \ge 6.5$$\end{document}). The objectives of the feasibility tests are to confirm (1) whether the initial version of AI framework can “remember and distinguish” individual large EQs and take the best actions according to the up-to-date RL policy, (2) whether the best action can accurately reproduce the location and magnitude 14 days before the failure, and (3) whether RL of the framework can self-evolve by expanding its experiences and memory autonomously.

This feasibility test is “pseudo-prospective” since this framework has been pre-trained with the past 40 years’ data, identified unique states of individual large EQs, and found the best-so-far prediction rules. Each of the best prediction rules is proven successful in reproducing large EQ’s location and magnitude 30 days before the event as shown in the author’s prior work44. It is anticipated that the framework should have “experienced” during training by which it can distinguish the individual events and can determine how to reproduce them using the best-so-far action out of the “memory,” the tabular Q-value function in this stage of framework. The objective of this feasibility test is to prove the potential before applications and expansions to the “true” prospective predictions in future research. Figure 2 schematically explains the process behind the RL-based pseudo-prospective prediction, starting from a new state, to Q-value function, and to the best-so-far action. Figure S29 presents some selected cases that confirm the promising performance of this framework. All other comparative investigations with large EQs (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w \ge 6.5$$\end{document}) during the past 40 years in the western U.S. are presented in Figs. S16 through S24. All cases confirm this framework’s promising performance compared to the EQ forecasting method. Particularly, in all pseudo-prospective predictions, RL can remember the states (i.e., the quantified information around the spatio-temporal vicinity of the large EQs during training) and also select the best actions (i.e., the customized prediction rules to the individual large EQs) from the stored policy.

As a reference comparison, a well-established EQ forecasting method UCERF3-ETAS13–16 is adopted to conduct short-term predictions 14 days before the large EQs. This may not be an apple-to-apple comparison since EQ forecasting is not mainly designed for such short-term predictions of specific large EQs before the events. Still, this comparison meaningfully provides a relative standing of the framework showing its role and difference from the existing EQ forecasting approaches. Detailed settings of UCERF3-ETAS are presented in Supplementary Materials. Figure 3 compares the pseudo-predictions by UCERF3-ETAS and this framework. UCERF3-ETAS predicts large EQs in ranges of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w \in [5,6)$$\end{document} with 0.3% probability 30 days before the onset. Despite the low probability and the magnitude error, the spatial proximity of the prediction to real EQ is noteworthy, i.e., the distance between the dashed box and the real EQ (Fig. 3C). All EQs (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w>3.5$$\end{document}) of the past year are used as self-triggering sources of UCERF3-ETAS (Fig. 3D). In contrast, Fig. 3B confirms that this AI framework can remember the unique states of the reference volumes near the hypocenter of the Ridgecrest EQ, select the best-so-far action from the up-to-date policy, and successfully reproduce the peak of the real EQ 14 days before the failure. As expected, this framework’s reproduction of large EQs appears successful in achieving the threefold objective, i.e., accuracy in magnitude, location, and short-term timing. It should be noted that this framework does not remember any information on specific dates or loci of past EQs for training and prediction. Only the point-wise states near the large EQs are remembered. Also, the best actions associated with the states are stored as policy in terms of relative selection probabilities. At the time of the feasibility test, the framework accumulated about 110 important state sets and the associated 110 best actions. Thus, the feasibility test used the up-to-date RL policy stored in 110 tabular Q-value functions.

This performance of pseudo-prospective predictions is promising since searching for a similar (or identical) state from the storage is not a trivial task given a new state. Each time step (here, 1 day), there are \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$n_s = $$\end{document}253,125 reference volumes for the present resolution of 0.1 degrees of latitude and longitude, and 5 km depth. If we use all the point-wise states of the past 10 years, it amounts to \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$10\times 365\times 253125$$\end{document}, i.e., approximately over 0.92 billion states. Figure 2 illustrates such a long RL searching sequence, starting from a new state, to Q-value function, and to the best-so-far action.

In some cases, UCERF3-ETAS appears to show promising prediction accuracy about epicenters’ loci, which can be also found in 1992/4/25 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w 7.2$$\end{document} EQ (Fig. S17C), 1992/6/28 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w 7.3$$\end{document} EQ (Fig. S17C), and 2010/4/4 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w 7.0$$\end{document} EQ (Fig. S20C). The common aspect of these cases is that they have relatively large prior EQs right before the prediction begins (time = 0), i.e., \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w \approx 6$$\end{document} about 3 months earlier (Fig. S20D), \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w \approx 6$$\end{document} about 50 days earlier (Fig. S17D), and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w \approx 5.2$$\end{document} about 40 days earlier (Fig. S16D). This may play favorably for the “self-triggering” mechanism of UCERF3-ETAS. but a general conclusion is not available since other cases do not support the consistent accuracy in loci of epicenters.

Self-evolution is the central feature of the present framework. Fig. 3E–I show how the new best action update can evolve in a positive direction. Given a new state (i.e., when a new large EQ occurs; Fig. 3H), “old” best action is used to predict the magnitudes which may not be accurate enough (Fig. 3I). Then, the Action Update Core seeks to learn “new” best action for the new state (Fig. 3G). Comparison of Fig. 3G and I clearly demonstrate the positive evolution of the new action, and Fig. 3E and F quantitatively compare the maximum reward of all state-action pairs. This example underpins that favorable evolution can take place by learning new actions.

Impact of new features on prediction rules

One of the key contributions of this paper is to generate new ML-friendly features that are based on basic physics and generic mathematics. It is informative to touch upon varying contributions of new features on the prediction rules. We conducted three separate training with 3 prediction rules that are4 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\begin{aligned} M_{pred}^{[Model A]}= & \mathcal {L}_{E}(E_r^{*(t)}) \cdot \mathcal {L}_{P}\left( S_{gm}\left( e^2 {\partial E_r^{*(t)}}/{\partial {t}} \right) \right) \end{aligned}$$\end{document}

5 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\begin{aligned} M_{pred}^{[Model B]}= & \mathcal {L}_{E}(E_r^{*(t)}) \cdot \mathcal {L}_{P}\left( S_{gm}\left( e^2 {\partial E_r^{*(t)}}/{\partial {t}} \right) \right) \cdot \mathcal {L}_{\omega }\left( S_{gm}\left( e^2\omega _{\lambda } \right) \right) \cdot \mathcal {L}_{L}\left( S_{gm}\left( 10^{-4} {\partial ^2 E_r^{*(t)}}/{\partial {\lambda ^2}} \right) \right) \end{aligned}$$\end{document}

Model A (denoted by the author inside the program) utilizes the pseudo released energy and its power whereas Model B further harnesses pseudo vorticity and Laplacian terms. Model C, the prediction rule of this paper, uses Gauss curvature and FFT-based new features, as fully explained in Eq. (10) of Supplementary Materials. The detailed definitions and explanations about the terms in above models are presented in Supplementary Information. Figure 4 shows how the new features contribute to the prediction accuracy. All predictions are made 28 days before the real EQ event (Cape Mendocino EQ, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w=7.2$$\end{document}, April 25, 1992) by using the best-so-far prediction rules. As clearly seen in Fig. 4B–D, the increasing addition of new features appears to boost the accuracy of the large EQ reproduction. Interestingly, the seemingly poor reproductions with Model A (i.e., pseudo released energy and power) and Model B (i.e., all four pseudo physics terms) still can capture the peak’s location. With the Gauss curvature and FFT-based new features (Model C, Fig. 4D), both location and magnitude are sharply reproduced. Naturally, the AI framework favors Model C over Models A and B. This comparison demonstrates that the new ML-friendly features hold significant impact on the accuracy on large EQ reproduction in magnitude, location, and short-term timing aspects. Also, it suggests that future extensions with more new features may substantially help improve the accuracy of the framework. Thus, the positive evolution of this framework appears possible, which warrants further investigation into new ML-friendly features.Figure 4 Comparison of prediction rules with different new features: (A) Real observed EQ [Cape Mendocino EQ, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w=7.2$$\end{document}, April 25, 1992]; (B–D) Reproduced EQ peaks by using the best-so-far prediction rules. (B) Model A uses the pseudo released energy and power whereas (C) Model B additionally uses vorticity and Laplacian terms (i.e., all four pseudo physics terms). (D) Model C is the rule presented by this paper in Eq. (10). All prediction rules are trained with the same hyperparameters and settings.

Conclusions and outlook

It is instructive to touch upon the optimality of the adopted approach. One of the primary goals of the RL is to find the best policy that leads to the “optimal” values of the states, i.e., \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ v^*(s) :=\underset{\pi }{\text {max}}v_{\pi }(s), \forall s\in \mathcal {S}$$\end{document}. The “Bellman optimality equation”45 formally states6 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\begin{aligned} \begin{aligned} v^*(s)&=\underset{a\in \mathcal {A}(s)}{\text {max}}q_{\pi ^*}(s,a) \\&=\underset{a}{\text {max}}\mathbb {E}[R^{(t+1)}+\gamma v^*(S^{(t+1)})|S^{(t)}=s,A^{(t)}=a] \\&=\underset{a}{\text {max}}\underset{s'}{\sum }\underset{r}{\sum }\text {Prob}(s',r|s,a)[r+\gamma v^*(s')] \end{aligned} \end{aligned}$$\end{document}

where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\pi ^*$$\end{document} is the “global” optimal policy, q(s, a) is the action-value function meaning the expected return of the state-action pair following the policy, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\gamma $$\end{document} is the discount factor of future return, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$s'$$\end{document} is the next state. In view of Eq. (6), this paper’s algorithm is partially aligned with the optimality equation, achieving the “local optimality.” To explain this, it is necessary to note what we don’t know and what we do know. On one hand, the transition probability \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {Prob}(s',r|s,a)$$\end{document} from the present state \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{(t)}=s$$\end{document} to the next state \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{(t+1)}=s'$$\end{document} is assumed to remain unknown and completely random (i.e., the hidden transition from past EQs to future EQ at a specific location). Also, the “global” optimal policy \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\pi ^*$$\end{document} is not known at the early stage of learning. On the other hand, this paper seeks to use the best action for the given state (by Eqs. 12 and 13, thereby satisfying \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${\text {max}}q_{\pi ^{(t)}}(s,a)$$\end{document} of the 1st line of Eq. 6. Therefore, the present algorithm of this paper will be able to achieve at least “local optimality” with the present policy \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\pi ^{(t)}$$\end{document}, facilitating a gradual evolution toward the globally optimal policy with gradually better policy (i.e., \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\pi ^{(t)} \rightarrow \pi ^*, t\rightarrow \infty $$\end{document}). This framework can embrace the existing geophysical approaches. For instance, the RL’s “state-transition probability” \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$p:\mathcal {S}\times \mathcal {S}\times \mathcal {A}\rightarrow \mathbb {R}[0,1]$$\end{document} is given by7 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\begin{aligned} p(S'|S, A):= \text {Prob}\left( S'=S^{(t+1)}|S=S^{(t)}, A=A^{(t)}\right) =\sum _{R\in \mathcal {R}}p(S', R|S, A) \end{aligned}$$\end{document}

Such state-transition probabilities are widely used in various forms in seismology since geophysicists derived the statistical rules such as ETAS41,48 or Gutenberg-Richter law42 based on persistent observations of the EQ transitions over a long time. In the future extension, the existing statistical laws may help the proposed framework in the form of state-transition probability. The adopted GPRL is in spirit similar to the symbolic regression49,50, which may provide more general (also complicated) forms of the prediction rules. For general forms, a future extension of Action Update Core may harness a deep neural network that takes input of all the aforementioned new features (e.g., \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\overline{II}, \mathcal {U}, \mathbb {K}, \mathbb {F}$$\end{document}) and generates output of scalar action value. Since the policy determines actions for the given state, achieving consistently improving policy is vital for the reliable evolution of prediction models. The proposed learning framework is a “continuing” process, not an “episodic” one. Also, the state-transition \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{(t)}\rightarrow S^{(t+1)}$$\end{document} (i.e., the transition from present EQ to future EQ at a location) is assumed to be stochastic and remains unknown. Therefore, future extensions may leverage general policy approximations and policy update methods51. The present state-action pair is a simple one-to-one mapping, and the present Q-value function is the simplest tabular form with clear interpretability. To accommodate complex relationships between states and actions, a future extension may find a general policy that is parameterized by \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\varvec{\uptheta }_{\pi }$$\end{document}, denoted as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\pi (A^{(t)}|S^{(t)}; \varvec{\uptheta }_{\pi })$$\end{document}. A possibility of concurrent approximation and improvement of both Q-value function \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$q(S,A;\varvec{\uptheta }_q)$$\end{document} and policy \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\pi (A|S; \varvec{\uptheta }_{\pi })$$\end{document}, i.e., update of policy (actor) and value (critic) called actor-critic method45.Figure 5 Potential intellectual benefit from this paper’s outcomes.

The proposed AI seismogenesis agent will continue to evolve with clear interpretability, and machine learning-friendly data to be generated by this framework will accumulate. Toward the short-term deterministic large EQ predictions, this approach will add a meaningful dimension to our endeavors, empowered by data and AI. This AI framework will continue to enrich the new feature database which will accelerate the data- and AI-driven research and discovery about short-term large EQs. As depicted in Fig. 5, the proposed database will help diverse ML methods to explore and discover meaningful models about short-term EQs, importantly from data. Since the database will always preserve the physical meanings of each features, the resultant AI-driven models will offer (mathematically and physically) clear interpretations to researchers. The best-so-far AI-driven large EQ model will tell us which set of physics terms play a decisive role in reproducing the individual large EQ. Relative importance of the selected physics terms will be provided in terms of weights (e.g., in deep learning models) or mathematical expressions (e.g., in rule-learning models). These will substantially complement probabilistic or statistical models available in EQ forecasting/prediction methods of the existing research communities.

Another future research direction should include the generality of the proposed AI framework. Can the learned prediction rules apply to large EQ events in the untrained spatio-temporal ranges? If so, how reliable could the prediction be? To answer the general applicability of the AI framework is of fundamental importance since it will hint at a rise of practical large EQ prediction capability. This AI framework aims to establish a foundation for such a bold capability. Currently, the present research is dedicated to training the AI framework with the past four decades, generating new ML-friendly database, and accumulating new prediction rules. Still, we can glimpse the general applicability of the AI framework. Figure 6 shows examples of reasonably promising predictions of different temporal ranges (i.e., never used for training) with the best-so-far prediction rules. Some noisy false peaks are noticeable (Fig. 6B and D), but the overall prediction appears to be meaningful. It should be noted, however, that the majority of such blind predictions to other time ranges have failed to accurately predict the location and magnitude of large EQs (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w\ge 5.5$$\end{document}). This could be attributed to the early phase of the AI framework. As of August 2024, only 6.67% of total available EQ catalog between the year 1980 and 2023 are transformed into new ML-friendly features due mainly to the expensive computation cost even with the high-performance computing. If the AI framework continues expanding its coverage to all the available EQ catalog and learn more prediction rules, thereby enriching state and action sets, it may lead to practically meaningful generality in the not-too-distant future.Figure 6 Examples of generality of the AI framework applied to different temporal regions: (A) Real observed EQ [Ferndale2 EQ, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w=6.6$$\end{document}, February 19, 1995] and (B) the predicted EQ peaks by using the best-so-far prediction rules 16 days before the EQ. (C) Real observed EQ [Cape Mendocino EQ, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w=7.2$$\end{document}, April 25, 1992] and (B) the predicted EQ peaks by using the best-so-far prediction rules 28 days before the EQ.

This paper holds transformative impacts in several aspects. This paper’s AI framework will (1) help transform decades-long EQ catalogs into ML-friendly new features in diverse forms while preserving basic physics and mathematical meanings, (2) enable transparent ML methods to distinguish and remember individual large EQs via the new features, (3) advance our capability of reproducing large EQs with sufficiently detailed magnitudes, loci, and short time ranges, (4) offer a database to which geophysics experts can facilely apply advanced ML methods and validate the practical meaning of what AI finds, and (5) serve as a virtual scientist to keep expanding the database and improving the AI framework. This paper will add a new dimension to existing EQ forecasting/prediction research.

Materials and methods

Bayesian evolutionary algorithm with threefold objective

Table 2 briefly summarizes the threefold error function that is critical for the short-term EQ deterministic prediction capability. It also summarizes the salient steps of the Bayesian evolutionary algorithm used for searching the free parameters of all the LFs.Table 1 Summary of the Bayesian evolutionary algorithm (full details available in43.

Free parameter searching	
Goal: Learn all free parameters \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\Theta $$\end{document} of all LFs. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\Theta = (\uptheta _E, \uptheta _P, \uptheta _{\omega }, \uptheta _L, \uptheta _{FT} )$$\end{document}	
Loop of solution space searching, random mutation, spawning, and Bayesian inheritance.	
s denotes an individual in the entire generation S	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet $$\end{document} Fitness: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {F}(s) = \frac{ (1+ \mathcal {J}(s))^{-1}}{ \sum _{\forall s \in S} [(1+ \mathcal {J}(s))^{-1}]}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet $$\end{document} Bayesian Fitness Score: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {F}_B(s) = \frac{1}{\kappa } \frac{\mathcal {F}(s;S^{*}) \mathcal {F}^*(s)}{\sum _{\forall s\in {S^*}} \mathcal {F}^*(s)},$$\end{document}          where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\kappa = \sum _{\forall s\in {S^*}} \frac{\mathcal {F}(s;S^{*})\mathcal {F}^{*}(s)}{\sum _{\forall s\in {S^*}} \mathcal {F}^*(s)}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet $$\end{document} Probability for selection: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {Prob}(\text {parent}_k | s) \propto \mathcal {F}_B (s),~(k=1,2)$$\end{document}	
ThreeFold Error Function (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {J}$$\end{document}) about Magnitude, Location, and False Alarms	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet \mathcal {J} = (1-a_{cnt}) \sum _{{\tilde{k}} \in Top} \omega _{MD}^{(\tilde{k})} E_{MD}^{({\tilde{k}})}/n(Top) + a_{cnt} E_{cnt}$$\end{document}	
      \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$E_{MD}^{({\tilde{k}})}$$\end{document} = Errors in magnitude & 3D spatial distance of the real and predicted hypocenters.	
      \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$E_{cnt}$$\end{document} = Error in the total count of false alarms.	
      \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$a_{cnt}$$\end{document} = Coefficient about the relative importance between \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$E_{MD}^{({\tilde{k}})}$$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$E_{cnt}.$$\end{document} 0.1 is used herein.	
      \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$Top:=$$\end{document} A set of indices of reference volumes that contains the sorted real magnitudes \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\ge 6.5$$\end{document}	
      \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\omega _{MD}^{(\tilde{k})}:=\text {exp}(M_{obs}^{(t+1)}(\varvec{\xi }_{\tilde{k}})/10.0)$$\end{document}. Magnitude-dependent scale-up factor.	

8 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\begin{aligned} & \text {Type 1: } R_{\pi }(A_i, S_j) :=\text {exp} \left( \frac{M_{real}(\varvec{\xi }_j)}{c_{Mr}} \right) \left( 1-\text {erf}\left( \frac{|M_{real}(\varvec{\xi }_j)-M_{pred}(S_j)|}{M_{real}(\varvec{\xi }_j)} \right) \right) \end{aligned}$$\end{document}

9 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\begin{aligned} & \quad \text {Type 2: } R_{\pi }(A_i, S_j) :=\omega _{dist}\text {exp} \left( \frac{{M^*}_{real}}{c_{Mr}} \right) \left( 1-\text {erf}\left( \frac{|{M^*}_{real}-M_{pred}(S_j)|}{{M^*}_{real}}\right) \right) \end{aligned}$$\end{document}

where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_{real}(\varvec{\xi }_j)$$\end{document} is the observed magnitude at \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\varvec{\xi }_j$$\end{document}; \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {erf(.)}$$\end{document} is the Gauss error function. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${M^*}_{real}$$\end{document} is the maximum observed magnitude within a distance range, denoted as “reward distance range” (RDR), which is centered at \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\varvec{\xi }_j$$\end{document}. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\varvec{\xi ^*}$$\end{document} is the spatial location vector of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${M^*}_{real}$$\end{document} (Fig. S10). Physically, RDR is a spatial range where the prediction is meaningful. Even if the peak magnitudes are nearly identical, prediction may be considered valid only when the predicted and observed peaks are close enough (Fig. S10). \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\omega _{dist}\in \mathbb {R}(0,1]$$\end{document} is a distance-dependent discount factor given as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\omega _{dist}:=\text {exp} \left( \frac{-\Vert \varvec{\xi }_j - \varvec{\xi ^*} \Vert _2}{ c_{Dr} } \right) .$$\end{document} In essence, the reward is discounted when the distance between predicted and observed peaks becomes large via \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\omega _{dist}$$\end{document} whereas the reward is scaled up when the observed real EQ is large, i.e., the larger EQ, the more reward as shown in Fig. S25. For instance, suppose that the current state-action pair \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$(S_j, A_i)$$\end{document} predicts magnitude, the closest real peak magnitude within RDR is \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M^*_{real}=7.0$$\end{document} and their distance is 100 km. Then, for the reward, the distance-dependent discount occurs by a factor of 0.606 (i.e., exp(− 100 km/200 km)) while the magnitude-dependent scaling up occurs by a factor of 106.342 (i.e., exp(7/1.5)). From preliminary investigations, it is recommended that RDR = 200 km, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$c_{Dr}$$\end{document} = 200 km, and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$c_{Mr} = 1.5$$\end{document} (see Fig. S11).

Reward calculation core

This core is to calculate “reward” (denoted as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$R_{\pi }(A_k^{(t)}, S^{(t)})$$\end{document}) that is defined by the prediction accuracy of the chosen actions \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$A_k^{(t)}$$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$(k=1,...,n(\mathcal {A}^{(t)}))$$\end{document} according to the best-so-far policy \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\pi $$\end{document} given the state \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$S^{(t)}$$\end{document} (Fig. 1C). In concept, the reward is inversely proportional to the error of individual event predictions: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$R \propto \mathcal {J}^{-1}$$\end{document} where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {J}$$\end{document} is a generic error function in terms of loci, magnitudes, and timings (explained in Table 2). It is similar to but more general than the collective measure used in the existing statistical EQ forecasting methods (e.g.,18. For comparison, the reward core calculates two types of reward—Type 1 depending on the error of peak magnitudes while Type 2 on the error of both magnitude and distance. Type 2 reward turns out to be more favorable for the present framework than Type 1. It is important to note the non-zero lower bound of the proposed reward (Type 2). According to the Type 2 reward definition (Eq. 9), the magnitude-dependent scale-up factor \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {exp}(M^*_{real}/c_{Mr})$$\end{document} of larger EQs can increase greater than 100.0 whereas the distance-dependent discount factor \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\omega _{dist}\in \mathbb R(0.5, 1.0]$$\end{document} and the magnitude error term (1-erf(.))\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\in \mathbb R(0,1]$$\end{document} as shown in Fig. S25. When a magnitude error is 100% with \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M^*_{real}=7.0$$\end{document}, the magnitude error term (1-erf(.))\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\approx 0.1573$$\end{document} whereas \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\text {exp}(M^*_{real}/c_{Mr})$$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\approx $$\end{document} 106.342. Therefore, even if there is no distance error (i.e., \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\omega _{dist}=1.0$$\end{document}), the incorrect prediction may end up with a reward of 16.727.

High-performance computing

All training and simulations are conducted on a high-performance computing (HPC) facility. One-time update of state, action, and policy with the past 10 years’ EQ catalog takes more than 40 hours on 144 cores. One-time pseudo-prospective prediction takes about 70 minutes on 36 cores. The specification of the HPC is as follows: 36 cores per node (each node has two 18-core Intel Skylake 6140 processors), 384GB memory per node, 100G infinite band interconnect, and 1.5TB local hard drive. In the future extension, higher resolutions of larger domains and faster computation can be achieved by advanced parallel computing algorithm such as53.

Simulation setting of UCERF3-ETAS for pseudo-prospective predictions

To conduct prospective predictions by using UCERF3-ETAS, we include all the fore shocks available in Comprehensive Earthquake Catalog (ComCat) of the Advanced National Seismic System. In particular, all past EQs within one-year window up to about 14 days before the main shock (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w\ge 6.5$$\end{document}) are fetched as sources for self-triggering mechanisms. In this way, it is anticipated that the UCERF3-ETAS may predict the main shock (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w\ge 6.5$$\end{document}) two weeks before the event by using prior 1-year EQs. For the data fetching, we used 100–200 km radius centered at the main shock location. We specified “--min-mag 3.5” to force the ComCat evaluation plots to use this specified magnitude-of-completeness value. We used the ShakeMap surfaces option looking for all ruptures with \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$M_w \ge 5$$\end{document} via “--finite-surf-shakemap-min-mag 5”. In total, 1000 simulations are conducted for each pseudo prospective prediction. Figure S13 presents actual UCERF3-ETAS configuration commands used to conduct the pseudo prospective predictions of the large EQs presented in this paper.Table 2 Summary of four-layer data transformations (adapted from44. Details are presented in Supplementary Materials.

[New Feature 1] Spatio-Temporal Information Index (II)	
From USGS EQ catalogs52, prepare data matrix of all past EQs, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\{\lambda , \phi , -h, M \}_{i}^{(t)}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet $$\end{document} Local II: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${II_{local}}^{(t)}({\textbf {x}}_i^{(t)}) = M_{i}^{(t)}/10 \in \mathbb {R}[0,1)$$\end{document} where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${\textbf {x}}_i^{(t)} = (x, y, z)_i^{(t)}$$\end{document}, the position vector of \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$i_{th}$$\end{document} EQ	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet $$\end{document} Spatial convolved II: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\overline{II}_S^{(t)}(\varvec{\xi }_{j}; L_k) = \int _{\text {V}} { \omega (\varvec{\xi }_j, {\textbf {x}}_{i}^{(t)}; L_k) II_{local}^{(t)}({\textbf {x}}_{i}^{(t)}) d{\textbf {x}}}$$\end{document}	
         where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\omega := \mathcal {N}({\textbf {x}}_{(i)}^{(t)}, L_k^2) = (L_k(2\pi )^{1/2})^{-N}\exp \left( -\frac{|{\textbf {x}}_{i}^{(t)} - \varvec{\xi }_j|^2}{2L_k^2}\right) $$\end{document}, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$L_k=$$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$k_{th}$$\end{document} spatial influence range	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet $$\end{document} Spatio-temporal II: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\overline{II}_{ST}^{(t)}(\varvec{\xi }_{j}; L_k, T_l) = \int { \omega (\tau ; T_l) \overline{II}_S^{(t_{past})}(\varvec{\xi }_{j}; L_k) dt_{past}}. \,\, (\tau = |t-t_{past}|, t \ge t_{past})$$\end{document}	
         where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\omega := \mathcal {N}(t, T_l^2)=(T_l(2\pi )^{1/2})^{-1}\exp \left( -\frac{\tau ^2}{2T_l^2}\right) $$\end{document}, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$T_l=$$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$l_{th}$$\end{document} temporal influence range	
[New Feature 2] Pseudo Physics Features \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal {U}=\{E_r, E_r', \varvec{\omega }, \nabla _g^2 E_r\}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet $$\end{document} Pseudo Released Energy at jth reference volume:	
         \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$E_r^{(t)}(\varvec{\xi }_{j}):= \text {max}\left[ \sum _{k=1}^{n_L=2} \sum _{l=1}^{n_T=2} \mathcal {L}^{(k,l)}(\overline{II}_{ST}^{(t)}(\varvec{\xi }_{j}; L_k, T_l); ~\varvec{\uptheta }^{(k,l)}), 0\right] $$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet $$\end{document} Pseudo Power: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$E_r'=\frac{\partial E_r^{(t)}(\varvec{\xi }_{j})}{\partial {t}}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet $$\end{document} Pseudo Vorticity: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\varvec{\omega }:= \nabla _g \times \left( \nabla _g{ \frac{\partial E_r^{(t)}(\varvec{\xi }_{j})}{\partial {t}}} \right) $$\end{document}	
         \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$= \left( \frac{\partial }{\partial \phi }\frac{\partial {E_r^{'}}}{\partial {h}} - \frac{\partial }{\partial {h}}\frac{\partial {E_r^{'}}}{\partial \phi }, \,\, \frac{\partial }{\partial {h}}\frac{\partial {E_r^{'}}}{\partial \lambda } - \frac{\partial }{\partial \lambda }\frac{\partial {E_r^{'}}}{\partial {h}}, \,\, \frac{\partial }{\partial \lambda }\frac{\partial {E_r^{'}}}{\partial \phi } - \frac{\partial }{\partial \phi }\frac{\partial {E_r^{'}}}{\partial \lambda } \right) $$\end{document}	
         where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\nabla _{g} \text {f}(\varvec{\xi }_j) = {\textbf {J}}\nabla \text {f}(\varvec{\xi }_j)$$\end{document} (Full details of the Jacobian J presented in43	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet $$\end{document} Pseudo Laplacian: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\nabla _g^2 E_r^{(t)}(\varvec{\xi }_{j}) = \frac{\partial ^2{E_r^{'}}}{\partial {\lambda ^2}}+ \frac{\partial ^2{E_r^{'}}}{\partial {\phi ^2}} + \frac{\partial ^2{E_r^{'}}}{\partial {h^2}}$$\end{document}	
[New Feature 3] Gauss Curvature-Based Features \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbb {K}=\{\kappa _1, \kappa _2\}_{\mathcal {U}}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet $$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\kappa _1 = H+C;~ \kappa _2 = H-C$$\end{document}	
         where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$H=\frac{GL-2FM+EN}{2(EG-F^2)}$$\end{document} and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$C=\sqrt{A^2 + B^2}$$\end{document} (formulae for all other coefficients presented in43	
[New Feature 4] Fourier Transform-Based Features \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbb {F}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet $$\end{document}[Step-I] Time-History of Gauss Curvatures of Pseudo Physics Quantities	
Loop: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\forall \xi _j \in \text {V}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\quad \quad \quad {\textbf {Preparation: }} (\forall t = 1,..., t_n) \text { Calculate } {K}(t;\xi _j) = ((\kappa _1, \kappa _2)_E, (\kappa _1, \kappa _2)_P, (\kappa _1, \kappa _2)_L, (\kappa _1, \kappa _2)_V)_j^{(t)} $$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\quad \quad \quad \mathbb {K}(n;\xi _j)=$$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\begin{bmatrix} ((\kappa _1, \kappa _2)_E, (\kappa _1, \kappa _2)_P, (\kappa _1, \kappa _2)_L, (\kappa _1, \kappa _2)_V)_j^{(1)} \\ \vdots \\ ((\kappa _1, \kappa _2)_E, (\kappa _1, \kappa _2)_P, (\kappa _1, \kappa _2)_L, (\kappa _1, \kappa _2)_V)_j^{(t_n)} \\ \end{bmatrix} \in \mathbb {R}^{n \times 8}$$\end{document}	
End Loop	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\bullet $$\end{document}[Step-II] Column-wise Fast Fourier Transform (FFT)	
Loop: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$(\forall \xi _j \in \text {V})$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\quad \quad \quad {\textbf {Loop: }} (i = 1,..., 8) $$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\quad \quad \quad \quad \quad ({\textbf {p}}^{(i)}_{PSD}, {\textbf {f}}^{(i)}) = \text {FFT}[{\textbf {k}}^{(i)}(n;\xi _j) = i_{th} \text { column of } \mathbb {K}(n;\xi _j) ] $$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\quad \quad \quad \quad (\overline{{\textbf {p}}}^{(i)}_{PSD}, \overline{{\textbf {f}}}^{(i)}) = \text {Sorted } ({\textbf {p}}^{(i)}_{PSD}, {\textbf {f}}^{(i)}) \text { by PSD in descending order}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\quad \quad \quad \quad (\overline{{\textbf {p}}}^{(i)}_{top}, \overline{{\textbf {f}}}^{(i)}_{top}) = \text {Top 10 largest } (\overline{{\textbf {p}}}^{(i)}_{PSD}, \overline{{\textbf {f}}}^{(i)}) \text { by PSD}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\quad \quad \quad {\textbf {End Loop}}$$\end{document}	
\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\quad \quad \quad \mathbb {F}(n;\xi _j)=$$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\begin{bmatrix} \overline{{\textbf {p}}}^{(1)}_{top}&\overline{{\textbf {f}}}^{(1)}_{top}&\dots \overline{{\textbf {p}}}^{(8)}_{top}&\overline{{\textbf {f}}}^{(8)}_{top} \end{bmatrix} \in \mathbb {R}^{10 \times 16} $$\end{document}	
End Loop	

Supplementary Information

Supplementary Information.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-024-72667-7.

Acknowledgements

The authors are grateful for the various research supports. This work was supported, in part, by the National Science Foundation (NSF) of the U.S.A. under grants CSSI-1931380 and CMMI-2129796. High-performance computing (HPC) for this study is partially supported by the HPC@ISU equipment at Iowa State University, some of which has been purchased through funding provided by the NSF CNS-2018594.

Author contributions

I.C. is responsible for all algorithms and programs presented as well as writing of the manuscript. A.C. is responsible for conducting earthquake forecasting simulations for feasibility tests.

Funding

NSF CSSI-1931380 and CMMI-2129796.

Data availability

The processed 40-years data sets consisting of the month-based epochs and the refined day-based epochs are shared on a cloud storage (available upon request). Other supplementary data and parallel programs supporting other findings of this paper will be available upon request to the corresponding author.

Declarations

Competing interests

The author declares no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Beroza GC Segou M Mousavi SM Machine learning and earthquake forecasting-next steps Nat. Commun. 2021 12 4761 10.1038/s41467-021-24952-6 34362887
Beroza, G. C., Segou, M. & Mousavi, S. M. Machine learning and earthquake forecasting-next steps. Nat. Commun. 12, 4761 (2021).34362887
2. Bayona JA Savran WH Rhoades DA Werner MJ Prospective evaluation of multiplicative hybrid earthquake forecasting models in California Geophys. J. Int. 2022 229 1736 10.1093/gji/ggac018
Bayona, J. A., Savran, W. H., Rhoades, D. A. & Werner, M. J. Prospective evaluation of multiplicative hybrid earthquake forecasting models in California. Geophys. J. Int. 229, 1736 (2022).
3. Wang L Barbot S Excitation of San Andreas tremors by thermal instabilities below the seismogenic zone Sci. Adv. 2020 6 eabb2057 10.1126/sciadv.abb2057 32917611
Wang, L. & Barbot, S. Excitation of San Andreas tremors by thermal instabilities below the seismogenic zone. Sci. Adv. 6, eabb2057 (2020).32917611
4. Bletery Q Nocquet J-M The precursory phase of large earthquakes Science 2023 381 6655 10.1126/science.adg2565
Bletery, Q. & Nocquet, J.-M. The precursory phase of large earthquakes. Science 381, 6655 (2023).
5. Ross ZE Cochran ES Trugman DT Smith JD 3D fault architecture controls the dynamism of earthquake swarms Science 2020 368 1357 1361 10.1126/science.abb0779 32554593
Ross, Z. E., Cochran, E. S., Trugman, D. T. & Smith, J. D. 3D fault architecture controls the dynamism of earthquake swarms. Science 368, 1357–1361 (2020).32554593
6. Jiang Junle Lapusta Nadia Deeper penetration of large earthquakes on seismically quiescent faults Science 2016 352 6291 1293 1297 10.1126/science.aaf1496 27284188
Jiang, Junle & Lapusta, Nadia. Deeper penetration of large earthquakes on seismically quiescent faults. Science 352(6291), 1293–1297 (2016).27284188
7. Allison KL Dunham EM Earthquake cycle simulations with rate-and-state friction and power-law viscoelasticity Tectonophysics 2018 733 9 232 256 10.1016/j.tecto.2017.10.021
Allison, K. L. & Dunham, E. M. Earthquake cycle simulations with rate-and-state friction and power-law viscoelasticity. Tectonophysics 733(9), 232–256 (2018).
8. Zhu W Allison KL Dunham EM Yang Y Fault valving and pore pressure evolution in simulations of earthquake sequences and aseismic slip Nat. Commun. 2020 11 4833 10.1038/s41467-020-18598-z 32973184
Zhu, W., Allison, K. L., Dunham, E. M. & Yang, Y. Fault valving and pore pressure evolution in simulations of earthquake sequences and aseismic slip. Nat. Commun. 11, 4833 (2020).32973184
9. Mitchell EK Fialko Y Brown KM Velocity-weakening behavior of Westerly granite at temperature up to 600 C J. Geophys. Res. Solid Earth 2016 121 9 6932 6946 10.1002/2016JB013081
Mitchell, E. K., Fialko, Y. & Brown, K. M. Velocity-weakening behavior of Westerly granite at temperature up to 600 C. J. Geophys. Res. Solid Earth 121(9), 6932–6946 (2016).
10. Xu X Sandwell DT Ward LA Milliner CW Smith-Konter BR Fang P Bock Y Surface deformation associated with fractures near the 2019 Ridgecrest earthquake sequence Science 2020 370 6516 605 608 10.1126/science.abd1690 33122385
Xu, X. et al. Surface deformation associated with fractures near the 2019 Ridgecrest earthquake sequence. Science 370(6516), 605–608. 10.1126/science.abd1690 (2020).33122385
11. Simons M Minson SE Sladen A Ortega F Jiang J Owen SE Meng L Ampuero JP Wei S Chu R Helmberger DV The 2011 magnitude 9.0 Tohoku-Oki earthquake: mosaicking the megathrust from seconds to centuries Science 2011 332 1421 1425 10.1126/science.1206731 21596953
Simons, M. et al. The 2011 magnitude 9.0 Tohoku-Oki earthquake: mosaicking the megathrust from seconds to centuries. Science 332, 1421–1425. 10.1126/science.1206731 (2011).21596953
12. Toda S Stein RS Long- and short-term stress interaction of the 2019 ridgecrest sequence and coulomb-based earthquake forecasts Bull. Seismol. Soc. Am. 2020 110 4 1765 1780 10.1785/0120200169
Toda, S. & Stein, R. S. Long- and short-term stress interaction of the 2019 ridgecrest sequence and coulomb-based earthquake forecasts. Bull. Seismol. Soc. Am. 110(4), 1765–1780 (2020).
13. Field EH Long-term time-dependent probabilities for the third uniform California earthquake rupture forecast (UCERF3) Bull. Seismol. Soc. Am. 2015 105 2A 511 543 10.1785/0120140093
Field, E. H. et al. Long-term time-dependent probabilities for the third uniform California earthquake rupture forecast (UCERF3). Bull. Seismol. Soc. Am. 105(2A), 511–543 (2015).
14. Field EH Milner KR Hardebeck JL Page MT van der Elst N Jordan TH Michael AJ Shaw BE Werner MJ A spatiotemporal clustering model for the third Uniform California Earthquake Rupture Forecast (UCERF3-ETAS): toward an operational earthquake forecast Bull. Seismol. Soc. Am. 2017 107 3 1049 1081 10.1785/0120160173
Field, E. H. et al. A spatiotemporal clustering model for the third Uniform California Earthquake Rupture Forecast (UCERF3-ETAS): toward an operational earthquake forecast. Bull. Seismol. Soc. Am. 107(3), 1049–1081 (2017).
15. Milner KR Field EH Savran WH Page MT Jordan TH Operational earthquake forecasting during the 2019 Ridgecrest, California, Earthquake sequence with the UCERF3-ETAS Model Seismol. Res. Lett. 2020 91 1567 1578 10.1785/0220190294
Milner, K. R., Field, E. H., Savran, W. H., Page, M. T. & Jordan, T. H. Operational earthquake forecasting during the 2019 Ridgecrest, California, Earthquake sequence with the UCERF3-ETAS Model. Seismol. Res. Lett. 91, 1567–1578 (2020).
16. Page Morgan T Field Edward H Milner Kevin R Powers Peter M The UCERF3 grand inversion: Solving for the long-term rate of ruptures in a fault system Bull. Seismol. Soc. Am. 2014 104 3 1184 1204 10.1785/0120130180
Page, Morgan T., Field, Edward H., Milner, Kevin R. & Powers, Peter M. The UCERF3 grand inversion: Solving for the long-term rate of ruptures in a fault system. Bull. Seismol. Soc. Am. 104(3), 1184–1204 (2014).
17. Shcherbakov, R., Zhuang, J., Zller, G. & Ogata, Y., Forecasting the magnitude of the largest expected earthquake. Nat. Commun. 10, 4051 (2019).
18. Nandan S Ram SK Ouillon G Sornette D Is seismicity operating at a critical point? Phys. Rev. Lett. 2021 126 128501 10.1103/PhysRevLett.126.128501 33834802
Nandan, S., Ram, S. K., Ouillon, G. & Sornette, D. Is seismicity operating at a critical point?. Phys. Rev. Lett. 126, 128501 (2021).33834802
19. Tan YJ Waldhauser F Ellsworth WL Zhang M Zhu W Michele M Chiaraluce L Beroza GC Segou M Machine-learning-based high-resolution earthquake catalog reveals how complex fault structures were activated during the 2016–2017 Central Italy sequence Seismic Rec. 2021 1 11 19 10.1785/0320210001
Tan, Y. J. et al. Machine-learning-based high-resolution earthquake catalog reveals how complex fault structures were activated during the 2016–2017 Central Italy sequence. Seismic Rec. 1, 11–19 (2021).
20. Bergen KJ Johnson PA de Hoop MV Beroza GC Machine learning for data-driven discovery in solid earth geoscience Science 2019 363 1299 10.1126/science.aau0323
Bergen, K. J., Johnson, P. A., de Hoop, M. V. & Beroza, G. C. Machine learning for data-driven discovery in solid earth geoscience. Science 363, 1299 (2019).
21. Mousavi SM Zhu W Ellsworth W Beroza GC Unsupervised clustering of seismic signals using deep convolutional autoencoders IEEE Geosci. Rep. Sens. Lett. 2019 16 11
Mousavi, S. M., Zhu, W., Ellsworth, W. & Beroza, G. C. Unsupervised clustering of seismic signals using deep convolutional autoencoders. IEEE Geosci. Rep. Sens. Lett. 16, 11 (2019).
22. Yang L Liu X Zhu W Zhao L Beroza GC Toward improved urban earthquake monitoring through deep-learning-based noise suppression Sci. Adv. 2022 8 eabl3564 10.1126/sciadv.abl3564 35417238
Yang, L., Liu, X., Zhu, W., Zhao, L. & Beroza, G. C. Toward improved urban earthquake monitoring through deep-learning-based noise suppression. Sci. Adv. 8, eabl3564 (2022).35417238
23. Mousavi SM Beroza GC Deep-learning seismology Science 2022 377 725 10.1126/science.abm4470
Mousavi, S. M. & Beroza, G. C. Deep-learning seismology. Science 377, 725 (2022).
24. Rouet-Leduc B Hulbert C Lubbers N Barros K Humphreys CJ Johnson PA Machine learning predicts laboratory earthquakes Geophys. Res. Lett. 2017 44 9276 9282 10.1002/2017GL074677
Rouet-Leduc, B. et al. Machine learning predicts laboratory earthquakes. Geophys. Res. Lett. 44, 9276–9282 (2017).
25. Hulbert C Rouet-Leduc B Johnson PA Ren CX Rivière J Bolton DC Marone C Similarity of fast and slow earthquakes illuminated by machine learning Nat. Geosci. 2019 12 69 74 10.1038/s41561-018-0272-8
Hulbert, C. et al. Similarity of fast and slow earthquakes illuminated by machine learning. Nat. Geosci. 12, 69–74 (2019).
26. Rouet-Leduc B Hulbert C Johnson PA Continuous chatter of the Cascadia subduction zone revealed by machine learning Nat. Geosci. 2019 12 75 79 10.1038/s41561-018-0274-6
Rouet-Leduc, B., Hulbert, C. & Johnson, P. A. Continuous chatter of the Cascadia subduction zone revealed by machine learning. Nat. Geosci. 12, 75–79 (2019).
27. DeVries PMR Viégas F Wattenberg M Meade BJ Deep learning of aftershock patterns following large earthquakes Nature 2018 560 632 634 10.1038/s41586-018-0438-y 30158606
DeVries, P. M. R., Viégas, F., Wattenberg, M. & Meade, B. J. Deep learning of aftershock patterns following large earthquakes. Nature 560, 632–634 (2018).30158606
28. Mignan A Broccardo M Neural network applications in earthquake prediction (1994–2019):Meta-analytic and statistical insights on their limitations Seismol. Res. Lett. 2020 91 4 10.1785/0220200021
Mignan, A. & Broccardo, M. Neural network applications in earthquake prediction (1994–2019):Meta-analytic and statistical insights on their limitations. Seismol. Res. Lett. 91, 4 (2020).
29. Mignan A Broccardo M One neuron versus deep learning in aftershock prediction Nature 2019 574 E1 E3 10.1038/s41586-019-1582-8 31578475
Mignan, A. & Broccardo, M. One neuron versus deep learning in aftershock prediction. Nature 574, E1–E3 (2019).31578475
30. Raissi Maziar Yazdani Alireza George Em Karniadakis, Hidden fluid mechanics: Learning velocity and pressure fields from flow visualizations Science 2020 367 1026 1030 10.1126/science.aaw4741 32001523
Raissi, Maziar & Yazdani, Alireza. George Em Karniadakis, Hidden fluid mechanics: Learning velocity and pressure fields from flow visualizations. Science 367, 1026–1030 (2020).32001523
31. Karpatne A Atluri G Faghmous JH Steinbach M Banerjee A Ganguly A Shekhar S Samatova N Kumar V Theory-guided data science: A new paradigm for scientific discovery from data IEEE Trans. Knowl. Data Eng. 2017 29 10 2318 2331 10.1109/TKDE.2017.2720168
Karpatne, A. et al. Theory-guided data science: A new paradigm for scientific discovery from data. IEEE Trans. Knowl. Data Eng. 29(10), 2318–2331 (2017).
32. Champion K Lusch B Kutz JN Brunton SL Data-driven discovery of coordinates and governing equations Proc. Natl. Acad. Sci. 2019 116 45 22445 22451 10.1073/pnas.1906995116 31636218
Champion, K., Lusch, B., Kutz, J. N. & Brunton, S. L. Data-driven discovery of coordinates and governing equations. Proc. Natl. Acad. Sci. 116(45), 22445–22451 (2019).31636218
33. Cho I Li Q Biswas R Kim J A framework for putting scientists’ eyes on glass-box physics rule learner and its application to nano-scale phenomena Nat. Commun. Phys. 2020 3 78 10.1038/s42005-020-0339-x
Cho, I., Li, Q., Biswas, R. & Kim, J. A framework for putting scientists’ eyes on glass-box physics rule learner and its application to nano-scale phenomena. Nat. Commun. Phys. 3, 78. 10.1038/s42005-020-0339-x (2020).
34. Bazroun M Yang Y Cho I Flexible and interpretable generalization of self-evolving computational materials framework Comput. Struct. 2021 260 106706 10.1016/j.compstruc.2021.106706
Bazroun, M., Yang, Y. & Cho, I. Flexible and interpretable generalization of self-evolving computational materials framework. Comput. Struct. 260, 106706. 10.1016/j.compstruc.2021.106706 (2021).
35. Cho I Yeom S Sarkar T Oh T Unraveling hidden rules behind the wet-to-dry transition of bubble array by glass-box physics rule learner Nat. Sci. Rep. 2022 12 3191
Cho, I., Yeom, S., Sarkar, T. & Oh, T. Unraveling hidden rules behind the wet-to-dry transition of bubble array by glass-box physics rule learner. Nat. Sci. Rep. 12, 3191 (2022).
36. Cho I A framework for self-evolving computational material models inspired by deep learning Int. J. Numer. Methods Eng. 2019 120 10 1202 1226 10.1002/nme.6177
Cho, I. A framework for self-evolving computational material models inspired by deep learning. Int. J. Numer. Methods Eng. 120(10), 1202–1226. 10.1002/nme.6177 (2019).
37. Kuang W Yuan C Zhang J Real-time determination of earthquake focal mechanism via deep learning Nat. Commun. 2021 12 1432 10.1038/s41467-021-21670-x 33664244
Kuang, W., Yuan, C. & Zhang, J. Real-time determination of earthquake focal mechanism via deep learning. Nat. Commun. 12, 1432 (2021).33664244
38. Pokharel B Alvioli M Lim S Assessment of earthquake-induced landslide inventories and susceptibility maps using slope unit-based logistic regression and geospatial statistics Nat. Sci. Rep. 2021 11 21333
Pokharel, B., Alvioli, M. & Lim, S. Assessment of earthquake-induced landslide inventories and susceptibility maps using slope unit-based logistic regression and geospatial statistics. Nat. Sci. Rep. 11, 21333 (2021).
39. Wang T Bian Y Zhang Y Hou X Classification of earthquakes, explosions and mining-induced earthquakes based on XGBoost algorithm Comput. Geosci. 2023 170 C 105242 10.1016/j.cageo.2022.105242
Wang, T., Bian, Y., Zhang, Y. & Hou, X. Classification of earthquakes, explosions and mining-induced earthquakes based on XGBoost algorithm. Comput. Geosci. 170(C), 105242 (2023).
40. Omori FJ On the aftershocks of earthquakes Coll. Sci. Imp. Univ. Tokyo 1984 7 111 200
Omori, F. J. On the aftershocks of earthquakes. Coll. Sci. Imp. Univ. Tokyo 7, 111–200 (1984).
41. Helmstetter A Sornette D Importance of direct and indirect triggered seismicity in the ETAS model of seismicity Geophys. Res. Lett. 2003 30 1576 10.1029/2003GL017670
Helmstetter, A. & Sornette, D. Importance of direct and indirect triggered seismicity in the ETAS model of seismicity. Geophys. Res. Lett. 30, 1576 (2003).
42. Gutenberg B Richter CF Seismicity of the Earth and Associated Phenomena 1954 Princeton Univ. Press
Gutenberg, B. & Richter, C. F. Seismicity of the Earth and Associated Phenomena (Princeton Univ. Press, 1954).
43. Cho I Gauss curvature-based unique signatures of individual large earthquakes and its implications for customized data-driven prediction Nat. Sci. Rep. 2022 12 8669 10.1038/s41598-022-12575-w
Cho, I. Gauss curvature-based unique signatures of individual large earthquakes and its implications for customized data-driven prediction. Nat. Sci. Rep. 12, 8669. 10.1038/s41598-022-12575-w (2022).
44. Cho I Sharpen data-driven prediction rules of individual large earthquakes with aid of Fourier and Gauss Nat. Sci. Rep. 2023 13 16009 10.1038/s41598-023-43181-z
Cho, I. Sharpen data-driven prediction rules of individual large earthquakes with aid of Fourier and Gauss. Nat. Sci. Rep. 13, 16009. 10.1038/s41598-023-43181-z (2023).
45. Sutton RS Barto AG Introduction to Reinforcement Learning 2017 MIT Press
Sutton, R. S. & Barto, A. G. Introduction to Reinforcement Learning (MIT Press, 2017).
46. Huang B von Rudorff GF von Lilienfeld OA The central role of density functional theory in the AI age Science 2023 381 170 175 10.1126/science.abn3445 37440654
Huang, B., von Rudorff, G. F. & von Lilienfeld, O. A. The central role of density functional theory in the AI age. Science 381, 170–175 (2023).37440654
47. Wood S Generalized Additive Models: An Introduction with R 2006 CRC Press
Wood, S. Generalized Additive Models: An Introduction with R (CRC Press, 2006).
48. Helmstetter A Kagan YY Jackson DD Comparison of short-term and time-independent earthquake forecast models for Southern California Bull. Seismol. Soc. Am. 2006 96 1 90 106 10.1785/0120050067
Helmstetter, A., Kagan, Y. Y. & Jackson, D. D. Comparison of short-term and time-independent earthquake forecast models for Southern California. Bull. Seismol. Soc. Am. 96(1), 90–106. 10.1785/0120050067 (2006).
49. Ma H Narayanaswamy A Riley P Li L Evolving symbolic density functionals Sci. Adv. 2022 8 eabq0279 10.1126/sciadv.abq0279 36083906
Ma, H., Narayanaswamy, A., Riley, P. & Li, L. Evolving symbolic density functionals. Sci. Adv. 8, eabq0279 (2022).36083906
50. Udrescu SM Tegmark M AI Feynman: A physics-inspired method for symbolic regression Sci. Adv. 2020 6 eaay2631 10.1126/sciadv.aay2631 32426452
Udrescu, S. M. & Tegmark, M. AI Feynman: A physics-inspired method for symbolic regression. Sci. Adv. 6, eaay2631 (2020).32426452
51. Ciosek K Whiteson S Expected policy gradients for reinforcement learning J. Mach. Learn. Res. 2020 21 1 51 34305477
Ciosek, K. & Whiteson, S. Expected policy gradients for reinforcement learning. J. Mach. Learn. Res. 21, 1–51 (2020).34305477
52. United States Geological Survey (USGS), Earthquake Catalog. USGS https://earthquake.usgs.gov/earthquakes/search/ (Accessed Apr 2022), (2022).
53. Cho I Porter K Multilayered grouping parallel algorithm for multiple-level multiscale analyses Int. J. Numer. Meth. Eng. 2014 100 914 932 10.1002/nme.4791
Cho, I. & Porter, K. Multilayered grouping parallel algorithm for multiple-level multiscale analyses. Int. J. Numer. Meth. Eng. 100, 914–932 (2014).
