
==== Front
Stoch Partial Differ Equ
Stoch Partial Differ Equ
Stochastic Partial Differential Equations
2194-0401
2194-041X
Springer US New York

320
10.1007/s40072-023-00320-x
Article
Importance sampling for stochastic reaction–diffusion equations in the moderate deviation regime
http://orcid.org/0000-0003-4464-5284
Gasteratos Ioannis igastera@ic.ac.uk

1
Salins Michael 2
Spiliopoulos Konstantinos 2
1 https://ror.org/041kmwe10 grid.7445.2 0000 0001 2113 8111 Department of Mathematics, Imperial College London, London, UK
2 https://ror.org/05qwgg493 grid.189504.1 0000 0004 1936 7558 Department of Mathematics and Statistics, Boston University, Boston, USA
8 12 2023
8 12 2023
2024
12 4 20812150
29 9 2022
10 6 2023
16 10 2023
© The Author(s) 2023
2023
https://creativecommons.org/licenses/by/4.0/ Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
We develop a provably efficient importance sampling scheme that estimates exit probabilities of solutions to small-noise stochastic reaction–diffusion equations from scaled neighborhoods of a stable equilibrium. The moderate deviation scaling allows for a local approximation of the nonlinear dynamics by their linearized version. In addition, we identify a finite-dimensional subspace where exits take place with high probability. Using stochastic control and variational methods we show that our scheme performs well both in the zero noise limit and pre-asymptotically. Simulation studies for stochastically perturbed bistable dynamics illustrate the theoretical results.

Keywords

Rare event simulation
Importance sampling
Monte Carlo methods
Moderate deviations
Stochastic reaction–diffusion equations
Optimal control
Mathematics Subject Classification

65C05
60G99
60F9
http://dx.doi.org/10.13039/100000121 Division of Mathematical Sciences DMS 1550918 DMS 2107856 Spiliopoulos Konstantinos http://dx.doi.org/10.13039/100000893 Simons Foundation 672441 Spiliopoulos Konstantinos issue-copyright-statement© Springer Science+Business Media, LLC, part of Springer Nature 2024
==== Body
pmcIntroduction

In this paper we are concerned with the problem of rare event simulation for the stochastic reaction–diffusion equation (SRDE)1 ∂tXϵ(t,ξ)=AXϵ(t,ξ)+f(Xϵ(t,ξ))+ϵW˙(t,ξ),(t,ξ)∈[0,∞)×(0,ℓ)Xϵ(0,ξ)=x(ξ),ξ∈(0,ℓ),NXϵ(t,ξ)=0,(t,ξ)∈[0,∞)×{0,ℓ},

where ϵ≪1, A is a uniformly elliptic second-order differential operator, f:R→R is a dissipative nonlinearity with polynomial growth and W˙ is a stochastic forcing term of intensity ϵ modeled by space-time white noise. The mixed boundary conditions are given by the linear operator N which acts on functions defined on the boundary ∂(0,ℓ) (see Sect. 2 for more details), and the initial datum x:(0,ℓ)→R is a continuous function in the kernel of N.

Systems like (1) are of interest because they exhibit metastable behavior. Assuming that the associated noiseless dynamics are non-trivial and ϵ>0, the stochastic forcing can induce transitions between neighborhoods of metastable states. As ϵ→0, transitions and exits from domains of attraction occur with very small probabilities and rigorous asymptotic analysis of exit times and places is possible within the framework of large deviations or potential theory (see e.g. [18, 28, 31, 32] and [4, 16, 27, 35, 47], as well as references within, for results in metastability theory in finite and infinite dimensions respectively).

In practice, efficient simulation of such events is challenging. On the one hand, Large Deviation Principles (LDPs) characterize the exponential decay rates of probabilities in the limit as ϵ→0 but ignore the effect of prefactors which can be significant (see [23]). On the other hand, as ϵ decreases, standard Monte-Carlo schemes require an increasingly large sample size in order to maintain a small relative error per sample. For this reason, accelerated and adaptive methods such as importance sampling or multi-level splitting become essential when it comes to rare events. For more details on the general theory and applications of such methods in a number of different models, the interested reader is referred to the book [10].

In the present work, we aim to develop a provably efficient importance sampling scheme that computes exit probabilities of Xϵ from scaled neighborhoods of a stable equilibrium point x∗. In particular, let Xxϵ denote the unique (mild) solution of (1) with initial condition x,  D⊂L2(0,ℓ) andτx∗ϵ=inf{t>0:Xx∗ϵ(t)∉D}.

For T,L>0, we focus on the estimation of probabilities P[τx∗ϵ≤T] in the case where D=Dϵ with2 Dϵ={x∈L2(0,ℓ):‖x-x∗‖L2<Lϵh(ϵ)}.

The scaling h(ϵ) is chosen so that h(ϵ)→∞ and ϵh(ϵ)→0, as ϵ→0. As ϵ→0, exit probabilities from such domains lie in an asymptotic regime that interpolates between the Central Limit Theorem (CLT) and LDP. To be precise, let Xx0 denote the (deterministic) solution of (1) with ϵ=0 and define a family of centered and re-scaled processes3 ηxϵ:=Xxϵ-Xx0ϵh(ϵ),ϵ>0.

As ϵ→0, the choices h(ϵ)=1/ϵ and h(ϵ)=1 correspond to large and Gaussian deviations of Xϵ respectively.

Exits of Xϵ from D are then equivalent to exits of ηxϵ from an L2-ball of radius L around 0 and large deviations of the family {ηxϵ}ϵ∈(0,1) are called moderate deviations of {Xxϵ}ϵ∈(0,1). Moderate Deviation Principles (MDPs) have been studied in many different contexts such as multiscale and interacting particle systems, Markov processes with jumps, small-noise stochastic dynamics, statistical estimation, option pricing and stochastic recursive algorithms see e.g. [30, 55] for SRDEs as well as [7, 11, 20, 29, 33, 34, 36, 38, 44].

Importance sampling is a variance-reduction accelerated Monte-Carlo method and its objective is to minimize the variance of the estimator by carefully chosen changes of measure. Such changes of measure "push" the dynamics towards trajectories that realize the rare event of interest. This procedure transforms tail events to more typical events, thus allowing for more efficient sampling. The simulation outcomes are then weighted by likelihood ratios so that the importance sampling estimators remain unbiased under the new probability measures. Importance sampling schemes for events in the large and moderate deviation regimes have been developed for finite-dimensional systems in [21, 23, 49, 50, 53]. In [21, 50], the authors observed that moderate-deviation based schemes provide a viable and simpler alternative to their large-deviation based counterparts, in cases where both are applicable. This is due to the fact that the MDP action functional, which characterizes exponential decay rates of probabilities, takes a much simpler form. In turn, this allows for more tractable and straightforward design of optimal changes of measure.

Importance sampling for SRDEs presents new challenges due to infinite dimensionality combined with the nonlinearity of the dynamics. Our work is close to [46] where a large deviation based scheme was developed for linear equations (i.e. when f=0). In there, the authors show that efficient changes of measure need to accomplish both variance and dimension reduction. For example, changes of measure that force infinitely many modes of the dynamics lead to estimators with very large variance when ϵ is small. A possible workaround is to show that exits from D take place in a finite-dimensional submanifold of ∂D with high probability. This was achieved in the linear case of [46] where it was proved that, under a sufficiently large spectral gap, exit from D happens in the direction of the eigenvector e1 of -A corresponding to the smallest non-zero eigenvalue. Similar results regarding the exit direction for (finite-dimensional) SDEs with a linear drift have been proved in [51] (see also Remark 5 below).

To the best of our knowledge, importance sampling for nonlinear SRDEs is rigorously studied here for the first time. The main difficulty in designing large deviation-based schemes for such equations lies in the task of identifying a finite-dimensional exit submanifold (if any). We are able to overcome this obstacle by working in the moderate deviation regime. As we show in the sequel, the latter is equivalent to linearizing the dynamics in a neighborhood of the equilibrium x∗. Consequently, the results of [46] can be applied locally at the cost of a linearization error which is, however, negligible as ϵ→0. In cases where both LDP and MDP-based schemes are available, one may think of the tradeoff between the two as follows: Moderate deviations cover the regime between central limit theorem and large deviations, so they are appropriate to characterize rare events, but not so rare that they would be in the large deviations regime. On the other hand, moderate deviations schemes are in general more tractable due to the asymptotic linearization of the dynamics that takes place. In our setting, this tradeoff is reflected in the fact that we only consider exit domains (2) in which the radius shrinks to zero as ϵ→0. Furthermore, the probability of exiting from a ball of radius ϵh(ϵ) is strictly smaller than the probability of exiting a ball of radius 1. The MDP importance sampling schemes described in this paper can provide a quantitative upper bound for the much more difficult to characterize LDP exit probabilities.

The design of an importance sampling scheme and proof of its good asymptotic and pre-asymptotic performance is the main contribution of this paper. In the course of our analysis, we prove an MDP for additive-noise SRDEs with a non-Lipschitz nonlinearity which cannot be found in the literature (see Theorem 3.1 and Remark 11). Furthermore, our theory is applied to the stochastic Allen–Cahn (also known as real Ginzburg-Landau or Chafee-Infante) equation and supplemented by simulation studies. In contrast to the linear case, there is a number of interesting cases where the aforementioned spectral gap is not satisfied. Another novel feature of this work is the construction of changes of measure that perform well asymptotically (i.e. as ϵ→0) in the absence of this condition (see Hypothesis 3c’ below).

The rest of this paper is organized as follows: In Sect. 2 we fix the notation and state our assumptions. In the first part of Sect. 3 we introduce moderate deviations and subsolution-based importance sampling and then state and prove our results on the asymptotic theory of the scheme. Section 4 is devoted to the implementation and pre-asymptotic performance analysis of our scheme. In Sect. 5 we apply the developed theory to the case where f is, up to a sign, the derivative of a double-well potential. Our examples include the stochastic Allen–Cahn equation (which features a cubic nonlinearity) with different boundary conditions as well as SRDEs with higher order polynomial nonlinearities. The results of simulation studies are then presented in Sect. 6. Finally, Appendix A collects the proofs of some useful lemmas.

Notation and assumptions

Let ℓ>0. The Hilbert space L2(0,ℓ) endowed with its usual inner product will be denoted by (H,⟨·,·⟩H). The Banach space C[0,ℓ], endowed with the supremum norm, is denoted by E. The norm of a Banach space X will be denoted by ‖·‖X and the closed ball of radius R>0 and center x0∈X, i.e. the set {x∈X:‖x-x0‖X≤R}, by BX(x0,R). We use D˚,D¯,∂D to denote interior, closure and boundary of a set D⊂X respectively. The lattice notation ∧,∨ is used to indicate minimum and maximum respectively.

For θ>0, p∈[1,∞), we denote by Wp,θ(0,ℓ) the fractional Sobolev space of x∈Lp(0,ℓ) such that[x]p,θp:=∬[0,ℓ]2|x(ξ2)-x(ξ1)|p|ξ2-ξ1|pθ+1dξ1dξ2<∞.

Wp,θ(0,ℓ), endowed with the norm ‖·‖p,θ:=‖·‖Lp(0,ℓ)+[·]p,θ, is a Banach space. W2,θ(0,ℓ) is a Hilbert space and is denoted by Hθ(0,ℓ). Moreover, for T>0 and β∈[0,1), we denote by Cβ([0,T];X) the space of β-Hölder continuous X-valued functions defined on the interval [0, T]. Cβ([0,T];X), endowed with the norm‖X‖Cβ([0,T];X):=‖X‖C([0,T];X)+[X]Cβ([0,T];X):=supt∈[0,T]‖X(t)‖X+supt≠ss,t∈[0,T]‖X(t)-X(s)‖X|t-s|β,

is a Banach space.

For any two Banach spaces X,Y we denote the space of linear bounded operators B:X→Y by L(X;Y). The latter is a Banach space when endowed with the norm ‖B‖L(X;Y):=supx∈BX(0,1)‖Bx‖Y. When the domain coincides with the co-domain, we use the simpler notation L(X). The spaces of trace-class and Hilbert-Schmidt linear operators B:H→H are denoted by L1(H) and L2(H) respectively. The former is a Banach space when endowed with the norm ‖B‖L1(H):=tr(B∗B) while the latter is a Hilbert space when endowed with the inner product ⟨B1,B2⟩L2(H):=tr(B2∗B1).

The operator A in (1) is a uniformly elliptic second-order differential operator in divergence form. In particular:4 Aϕ(ξ)=ddξ(a(ξ)dϕ(ξ)dξ),ξ∈(0,ℓ)

with a∈C1(0,ℓ) and infξ∈(0,ℓ)a(ξ)>0. The operator N acts on the boundary {0,ℓ} and can be either the identity operator (corresponding to Dirichlet boundary conditions), first-order differential operators of the typeNu(ξ)=b(ξ)u′(ξ)+c(ξ)u(ξ),ξ∈{0,ℓ}

for some b,c∈C1[0,ℓ] such that b≠0 on {0,ℓ} (corresponding to Neumann or Robin boundary conditions) orNu=(u(ℓ)-u(0),u′(ℓ)-u′(0))

for periodic boundary conditions. We denote by A the realization of the differential operator A in H, endowed with the boundary condition N. It is defined on a dense subspace Dom(A)⊂H that contains{u∈H2(0,ℓ):Nu=0}

and it generates a C0 semigroup of operators S={S(t)}t≥0⊂L(H). Moreover, the part of A in Dom(A)¯⊂E, where the closure is taken in the topology of E, generates either a C0 or an analytic semigroup for which we use the same notation (see e.g. A.27 in [17] for a definition). Regarding the spectral properties of A, we make the following assumptions:

Hypothesis 1(a)

In view of (4), the operator -A is self-adjoint. As a result, there exists a countable complete orthonormal basis {en}n∈N⊂H of eigenvectors of -A. The corresponding sequence of nonnegative eigenvalues is denoted by {an}n∈N.

Hypothesis 1(b)

The eigenvectors satisfysupn∈N‖en‖E<∞.

Remark 1

Without loss of generality, we can replace the operator A by A~=A-cI for some c>0 and the reaction term f in (1), by f~(x(ξ)):=f(x(ξ))+cx(ξ). The model is invariant under this transformation and, in light of Hypothesis 1(a), it follows that ‖S~(t)‖L(H)≤e-ct. Throughout the rest of this work we will be using A~,S~ and f~ with no further distinction in notation.

Let θ≥0. In view of Hypotheses 1(a) along with the previous remark, -A, restricted to its image, has a densely defined bounded inverse (-A)-1 which can then be uniquely extended to all of H. The fractional power (-A)-θ is defined via interpolation and is also injective. Letting (-A)θ2:=((-A)-θ2)-1 we define Hθ:=Dom((-A)θ2)=Range((-A)-θ2)⊂H. The norm ‖x‖Hθ:=‖(-A)θ2x‖H turns Hθ into a Banach space and is equivalent to the graph norm (see [41], Chapter 2.2).

Remark 2

For θ∈(0,12) the spaces Hθ(0,ℓ) and Hθ coincide via the identificationHθ(0,ℓ)=Hθ={x∈H:supt∈(0,1]t-θ/2‖S(t)x-x‖H<∞}

which holds with equivalence of norms. The latter implies that for each t≥0, the linear operator S(t)-I∈L(Hθ;H) and there exists a constant C>0 such that5 ‖S(t)-I‖L(Hθ;H)≤Ctθ/2.

The analytic semigroup S possesses the following regularizing properties (see e.g. section 4.1.1 in [13]) :

(i) For 0≤s≤r≤12 and t>0, S maps Hs(0,ℓ) to Hr(0,ℓ) and6 ‖S(t)x‖Hr≤Cr,s(t∧1)-r-s2ecr,st‖x‖Hs,x∈Hs(0,ℓ),

for some positive constants cr,s,Cr,s.

(ii) S is ultracontractive, i.e. for t>0, S(t) maps H to L∞(0,ℓ) and furthermore, for any 1≤p≤r≤∞,‖S(t)x‖Lr(0,ℓ)≤C(t∧1)-r-p2pr‖x‖Lp(0,ℓ),x∈Lp(0,ℓ).

The next set of assumptions concerns the nonlinear reaction term in (1).

Hypothesis 2(a)

f:R→R is twice continuously differentiable andf=f1+f2

where f1:R→R is globally Lipschitz continuous and f2:R→R is a non-increasing function.

Hypothesis 2(b)

There exists Cf>0 and p0≥3 such that for all x∈R and i∈{0,1,2}7 |∂x(i)f(x)|≤Cf(1+|x|p0-i).

For p≥1, f induces a superposition (or Nemytskii) operator F:E→Lp(0,ℓ) defined by F(x)(ξ):=f(x(ξ)), ξ∈(0,ℓ). In view of Hypotheses 2(a) and 2(b), F is twice Gâteaux differentiable along any direction in E and (with some abuse of notation) its Gâteaux differentials are given by DiF=∂xif, i=1,2.

The last set of assumptions concerns the stability properties of the deterministic and linearized dynamics governed by (1), after setting ϵ=0.

Hypothesis 3(a)

There exists at least one asymptotically stable equilibrium x∗∈Dom(A) of (1) solving the elliptic Sturm-Liouville problem Ax+F(x)=0.

Hypothesis 3(b)

The linear self-adjoint operator -A-DF(x∗) has a countable, non-decreasing sequence of nonnegative eigenvalues {anf}n∈N corresponding to a complete orthonormal set of eigenvectors {enf}n∈N⊂E. Therefore, the equilibrium x∗ is asymptotically stable.

Hypothesis 3(c)

The first two eigenvalues of the self-adjoint operator -A-DF(x∗) satisfy 3a1f<a2f.

This spectral gap provides a sufficient condition that allows us to identify a one-dimensional exit direction for limiting trajectories (see Lemma 3.4 below). A weaker condition under which our results continue to hold is 2a1f<a2f (see Remark 7). In fact, our asymptotic results continue to hold under the following relaxed spectral gap:

Hypothesis 3(c′)

There exists k0≥1 such that 3a1f<ak0+1f and a1f<a2f.

Note that Hypothesis 3(c) trivially implies Hypothesis 3c’ with k0=1. The latter will be used throughout Section 3 to prove asymptotic results. In Sect. 4 we restrict the pre-asymptotic analysis to schemes that work under Hypothesis 3(c).

Turning to the stochastic forcing, let (Ω,F,Ft≥0,P) be a complete filtered probability space. The space-time white noise W˙ is understood as the time-derivative of a cylindrical Wiener process W:[0,∞)×H→L2(Ω) in the sense of distributions. The latter is a Gaussian family of random variables with covariance given byE[W(t1,χ1)W(t2,χ2)]=t1∧t2⟨χ1,χ2⟩H,

for (ti,χi)∈[0,∞)×H,i=1,2. Given a separable Hilbert space (H1,⟨.,.⟩H1) such that H is a linear subspace of H1 and the inclusion map H→iH1 is Hilbert-Schmidt, W can be identified with the H1-valued Wiener processW(t)=∑n=1∞W(t,en)i(en),t≥0

with covariance operator Q=ii∗∈L1(H). This identification is assumed throughout the rest of this paper without further distinction in notation.

Having introduced the necessary notation, we can recast (1) as a stochastic evolution equation on E given by8 dXϵ(t)=[AXϵ(t)+F(Xϵ(t))]dt+ϵdW(t)Xϵ(0)=x.

A mild solution to the latter is defined as a process Xϵ satisfying for each ϵ and all t∈[0,T],9 Xϵ(t)=S(t)x+∫0tS(t-s)F(Xϵ(s))ds+ϵ∫0tS(t-s)dW(t)

with probability 1. The last term is known as a stochastic convolution and will be frequently denoted by WA. Our assumptions guarantee that the E-valued paths of WA are continuous with probability 1 and10 Esupt∈[0,T]‖WA(t)‖Ep<∞.

This can be proved by the stochastic factorization method of Da Prato-Zabczyk [17] (see also Theorem B.6 in [45]). Moreover, for each ϵ>0, (8) has a unique mild solution taking values in C([0,T];E) with probability 1 (see e.g. Theorem 2.2 in [14]).

Moderate deviations, importance sampling and asymptotic theory

General theory and main results

In this section we present some theoretical aspects of subsolution-based importance sampling in the moderate deviation regime, applied to our problem of interest. First, we recall the notion of a Moderate Deviation Principle (MDP).

Definition 3.1

Let T>0, X=H or E,x∈X and a functional Sx,T:C([0,T];X)→[0,∞] with compact sub-level sets.

(i) We say that the collection of C([0,T];X)-valued random elements {Xϵ}ϵ≪1 satisfies an MDP with action functional Sx,T if, for all continuous and bounded g:C([0,T];X)→R and all scalings h(ϵ) such that h(ϵ)→∞ and ϵh(ϵ)→0 as ϵ→011 limϵ→01h2(ϵ)logEe-h2(ϵ)g(ηxϵ)=-inf{ϕ∈C([0,T];X):ϕ(0)=0}[Sx,T(ϕ)+g(ϕ)],

where ηxϵ is defined as in (3).

(ii) A Borel set E⊂C([0,T];X) will be called an Sx,T-continuity set ifinfϕ∈E¯Sx,T(ϕ)=infϕ∈E˚Sx,T(ϕ).

As mentioned in Sect. 1 we aim to compute probabilities of the form12 P(ϵ)=P[τx∗ϵ≤T]

for ϵ≪1,T>0, where τx∗ϵ=inf{t>0:Xx∗ϵ(t)∉D} and13 D=Dϵ=B˚H(x∗,Lϵh(ϵ))

for some L>0. Passing to the moderate deviation process ηxϵ and recalling that x∗ is a (stable) equilibrium of Xx0 we see thatηx∗ϵ=Xx∗ϵ-x∗ϵh(ϵ)

and14 τx∗ϵ=inf{t>0:ηx∗ϵ(t)∉B˚H(0,L)}.

As will be shown in Sect. 3.4, η∗ϵ converges, as ϵ→0, to the solution of a linear deterministic PDE with zero initial condition. Since 0 is the unique fixed point of this PDE, the limit process is bound to stay at 0 and limϵ→0P(ϵ)=0. This is why accelerated methods that estimate P(ϵ) when ϵ is small are useful.

In this paper, we will only work with unbiased estimators. Hence, minimizing the variance of the estimator is equivalent to minimizing the second moment. As we show below, an upper bound for the exponential decay rate of the second moment of any unbiased estimator can be determined in terms of the action functional Sx,T.

Lemma 3.1

Let P(ϵ) as in (12) and P^(ϵ) be an unbiased estimator of P(ϵ) with respect to a probability measure P¯ defined on (Ω,F). For any ϕ∈C([0,T];X), let τϕ=inf{t>0:ϕ(t)∉B˚H(0,L)} and15 GT(0,0):=inf{ϕ∈C([0,T];X):ϕ(0)=0,τϕ=T}Sx∗,T(ϕ).

If {Xϵ} satisfies an MDP with action functional Sx,T and E={ϕ∈C([0,T];H):τϕ≤T} is a Sx,T-continuity set thenlim supϵ→0-1h2(ϵ)logE¯[(P^(ϵ))2]≤2GT(0,0),

where E¯ denotes expectation with respect to the measure P¯.

Proof

We haveP(ϵ)=P[τx∗ϵ≤T]=P[supt∈[0,T]‖ηx∗ϵ(t)‖H≥L]=P[η∗ϵ∈E].

Now, for any unbiased estimator P^(ϵ),E¯[P^(ϵ)2]≥E¯[P^(ϵ)]2=P(ϵ)2,

where we used Jensen’s inequality. Thuslim supϵ→0-1h2(ϵ)logE¯[P^(ϵ)2]≤2lim supϵ→0-1h2(ϵ)logP(ϵ)=-2lim infϵ→01h2(ϵ)logP(ϵ)=2inf{ϕ∈C([0,T];X):ϕ(0)=0,ϕ∈E}Sx∗,T≤2GT(0,0)

where we used the continuity property of E in the last equality. □

As in finite dimensions (see e.g. the discussion in Section 2.2 in [23]) , the previous lemma shows that 2GT(0,0) is the best possible exponential decay rate for any unbiased estimator. In turn, this motivates the following criterion for asymptotic optimality.

Definition 3.2

An unbiased estimator P^(ϵ) of P(ϵ) defined on a probability space (Ω,F,P¯) will be called asymptotically optimal iflim infϵ→0-1h2(ϵ)logE¯[(P^(ϵ))2]≥2GT(0,0).

In other words, an estimator is asymptotically optimal if its second moment achieves the best possible exponential decay rate in the limit as ϵ→0.

Importance sampling involves changes of measure chosen to guarantee that the corresponding estimators achieve optimal (or nearly optimal) asymptotic behavior. Given a measurable feedback control (or change of measure) u:[0,T]×H→H that is bounded on bounded subsets of H, we define a family of probability measures {Pϵ}ϵ>0 on (Ω,F) such that, for all ϵ, Pϵ<<P on FT anddPϵdP|FT=exp(h(ϵ)∫0T〈u(s,η∗ϵ(s)),dW(s)〉H-h2(ϵ)2∫0T‖u(s,η∗ϵ(s))‖H2ds).

Using these new measures, it is straightforward to verify thatP^(ϵ,u):=dPdPϵ1{τ∗ϵ≤T},

defined on (Ω,FT,Pϵ), is an unbiased estimator of P[τ∗ϵ≤T]. Its second moment is given by16 Qϵ(u):=Eϵ[P^(ϵ,u)2]=Eϵ[exp(-2h(ϵ)∫0τ∗ϵ〈u(s,η∗ϵ(s)),dW(s)〉H+h2(ϵ)∫0τ∗ϵ‖u(s,η∗ϵ(s))‖H2ds)1{τ∗ϵ≤T}].

As we show in the next lemma Qϵ(u) admits a variational stochastic control representation which will be useful for studying its asymptotic behavior. A similar variational formula can be found in (2.5) of [46].

Lemma 3.2

Let u:[0,T]×H→H be a measurable feedback control that is bounded on bounded subsets of H, uniformly in t∈[0,T]. Then for all ϵ>017 -1h2(ϵ)logQϵ(u)=infv∈AE[12∫0τ^∗ϵ,v‖v(s)‖H2ds-∫0τ^∗ϵ,v‖u(s,η^∗ϵ,v(s))‖H2ds],

where A is the collection of all H-valued, Ft≥0-adapted processes v defined on [0, T] such thatτ^∗ϵ,v=inf{t>0:η^∗ϵ,v(t)∉B˚H(0,L)}≤T

with probability 1,  η^∗ϵ,v solves18 dη^∗ϵ,v(t)=Aη^∗ϵ,v(t)+1ϵh(ϵ)[F(x∗+ϵh(ϵ)η^∗ϵ,v(t))-F(x∗)]dt+[v(t)-u(t,η^∗ϵ,v(t))]dt+1h(ϵ)dW(t)η^∗ϵ,v(0)=0H

andE∫0τ^∗ϵ,v‖v(s)‖H2ds<∞.

Proof

Let ϵ>0. From the Cameron-Martin-Girsanov theorem,Qϵ(u)=Eϵ[exp(-2h(ϵ)∫0τ∗ϵ〈u(s,η∗ϵ(s)),dWϵ(s)〉H-h2(ϵ)∫0τ∗ϵ‖u(s,η∗ϵ(s))‖H2ds)1{τ∗ϵ≤T}],

whereWϵ(t):=W(t)-h(ϵ)∫0tu(s,η∗ϵ(s))ds,t∈[0,T]

is a cylindrical Wiener process under Pϵ. Using yet another change of measure withdP~ϵdPϵ|FT=exp(-2h(ϵ)∫0T〈u(s,η∗ϵ(s)),dWϵ(s)〉H-2h2(ϵ)∫0T‖u(s,η∗ϵ(s))‖H2ds),

we can write19 Qϵ(u)=E~ϵ[exp(h2(ϵ)∫0τ∗ϵ‖u(s,η∗ϵ(s))‖H2ds)1{τ∗ϵ≤T}]=E[exp(h2(ϵ)∫0τ^∗ϵ‖u(s,η^∗ϵ(s))‖H2ds)1{τ^∗ϵ≤T}],

where η^∗ϵ solves{dη^∗ϵ(t)=Aη^∗ϵ(t)+1ϵh(ϵ)[F(x∗+ϵh(ϵ)η^∗ϵ(t))-F(x∗)]dt-u(t,η^∗ϵ(t))dt+1h(ϵ)dW(t),η^∗ϵ(0)=0H}

and τ^∗ϵ denotes the corresponding exit time for η^∗ϵ. This follows, once again, from the Cameron-Martin-Girsanov theorem, asW~ϵ(t):=W(t)+h(ϵ)∫0tu(s,η∗ϵ(s))ds,t∈[0,T]

is a cylindrical Wiener process under the measure P~ϵ. From (19) we see that the second moment of the estimator can be written as an exponential functional of the driving noise and, as such, it admits the variational representation (17) (see (2.5) in [46] as well as (14) in [50] for the finite-dimensional case). □

The form of the MDP action functional provides essential information for choosing changes of measure u that perform well asymptotically. In particular, if for all ϕ with Sx,T(ϕ)<∞ there exists a (local) Lagrangian Lx defined on a subset of X×H, such that20 Sx,T(ϕ)=∫0TLx(ϕ(t),ϕ˙(t))dt,

then "good" changes of measure are connected to subsolutions of the PDE21 ∂tU(t,η)+Hx(η,DηU(t,η))=0,(t,η)∈[0,T)×KU(T,η)=g¯(η),η∈K⊂H,

withg¯(η)=0,η:‖η‖H≥L∞,η:‖η‖H<L.

Here, Hx denotes the Hamiltonian corresponding to Lx via Legendre transform (up to a sign). In the problems we consider, the latter are not well-defined on the whole space but rather on a subset K×H⊂H×H, see e.g. (23) below. The notion of subsolution is meant in the sense of the following definition.

Definition 3.3

A subsolution of (21) is any U:[0,T]×K→R such that for all (t,η), U(·,η)∈C1(0,T), U(t,·)∈C1(K) in the sense of Fréchet differentiation and satisfies∂tU(t,η)+Hx(η,DηU(t,η))≥0,(t,η)∈[0,T)×KU(T,η)≤g¯(η),η∈K⊂H.

The interested reader is referred to [25] for the original development of subsolution-based importance sampling. As we will show below (Theorem 3.1 and Remark 11), when x=x∗, the MDP action functional takes the form (20) with22 Lx∗(η,v)=12‖v-[A+DF(x∗)]η‖H2,(η,v)∈Dom(A)∩E×H

and the corresponding Hamiltonian is given by23 Hx∗(η,p)=〈[A+DF(x∗)]η,p〉H-12‖p‖H2,(η,p)∈Dom(A)∩E×H.

A direct consequence of (20) is that we can construct an explicit stationary subsolution in terms of the corresponding quasipotential. The latter is given byVx∗(η)=inf{Sx∗,T(ϕ):ϕ∈C([0,T];X):ϕ(0)=0,ϕ(T)=η,T∈(0,∞)}=‖(-A)12η‖H2-〈DF(x∗)η,η〉=-〈[A+DF(x∗)]η,η〉,η∈Dom(A)

and Vx∗(η)=∞ otherwise. A physical interpretation of Vx∗(η) is that of the minimal "energy" required to push a path from 0 to the state η and its explicit form is a consequence of the fact that (8) is, in our setting, a gradient system (see e.g. [16, 17] Section 12.2.3 for SRDEs). In view of Hypotheses 1(a), 1(b), 3(a), 3(b) it follows that24 U(t,η)=a1fL2-Vx∗(η)

is a subsolution of (21) on K=Dom(A). The final condition is satisfied since a1f=infn∈Nanf.

Remark 3

In finite-dimensional systems, feedback controls (or changes of measure) defined by u(t,η)=-DηU(t,η) lead to nearly optimal asymptotic behavior (see [22] Section 2.3, [23] Theorem 2.4 for large-deviation and [50] Theorem 3.1 for moderate deviation-based schemes). A first issue that appears in infinite dimensions is that u(t,η^∗ϵ,v(t)) is not well-defined since with probability 1 and for all t,  η^∗ϵ,v(t)∉Dom(A). The latter is a consequence of the spatial irregularity of the noise.

Throughout the rest of this paper, Pnf:H→H denotes an orthogonal projection to the n-dimensional eigenspace span{ejf}j=1n and we consider the "projected" quasipotential Vx∗(Pnfη)=Vx∗(⟨η,e1f⟩He1f), the subsolution U(t,Pnfη) of (21) (with K=PnfH). The changes of measure we will use are given by25 uk0(t,η):=-DηU(t,P0fη):=2∑i=1k0aif⟨η,eif⟩Heif,

with k0 as in Hypothesis 3c’. For implementation purposes, uk0 is replaced by a sequence uk0ϵ that converges to uk0 as ϵ→0. For more details on the choice of u1ϵ see (63) and the discussion in Sect. 4 below.

We can now present our main results on the asymptotic behavior of the scheme.

Theorem 3.1

(Moderate Deviations) Let T>0,L>0 as in (13), k0 as in Hypothesis 3c’, uk0 as in (25), Qϵ as in (16) and BH(0,L)⊂H denote the closed ball of radius L centered at the origin. Moreover let uk0ϵ:[0,T]×H→H be a sequence that converges pointwise and uniformly over bounded subsets of H to uk0,26 T={y∈C([0,T];H):y(0)=0,∃τ∈(0,T]:y(τ)∈∂BH(0,L),y(t)∈BH(0,L)∀t∈[0,τ)}

andCy,x∗={v∈L2([0,T];H):y˙(t)=Ay(t)+DF(x∗)y(t)-uk0(t,y(t))+v(t)}.

Under Hypotheses 1(a)-(c), 2(a),(b), 3(a),(b),(c’) we have27 limϵ→0-1h2(ϵ)logQϵ(uk0ϵ)=infy∈Tinfv∈Cy,x∗∫0τ(12‖v(t)‖H2-‖uk0(t,y(t))‖H2)dt,

with the convention that the infimum over the empty set is ∞.

Remark 4

A few comments on (27) are in order: (1) If y∈H1((0,T);H)∩L2([0,T];Dom(A)), the set Cy,x∗ reduces to the singleton {v¯(t):=y˙(t)-Ay(t)-DF(x∗)y(t)-u(t,y(t))} and for any y∉H1((0,T);H)∩L2([0,T];Dom(A)), Cy,x∗ is empty. (2) Using the same notation, it follows that the right-hand side of (27) can be expressed asinfy∈T∫0τ(12‖v¯(t)‖H2-‖uk0(t,y(t))‖H2)dt.

(3) Since the functional on the right-hand side involves only the values of y on [0,τ] it is straightforward to see that the infimum can in fact be taken over paths y∈C([0,τ];H) that satisfy the constraints in (26).

Using the moderate deviation asymptotics of Theorem 3.1 we can then prove the following:

Theorem 3.2

(Near asymptotic optimality) Let L,T>0, k0,uk0,uk0ϵ:[0,T]×H→H as in Theorem 3.1, A as in Lemma 3.2, U as in (24) and GT as in (15). For any sequence {vϵ}⊂A such that28 -1h2(ϵ)logQϵ(uk0ϵ)≥E[12∫0τ^∗ϵ,vϵ‖vϵ(s)‖H2ds-∫0τ^∗ϵ,vϵ‖uk0ϵ(s,η^∗ϵ,vϵ(s))‖H2ds]-ϵ2

we have29 limϵ→0E〈η^∗ϵ,vϵ(τ^∗ϵ,vϵ),e1f〉H2=L2.

Moreover, we have the second moment bounds30 GT(0,0)+U(0,0)≤limϵ→0-1h2(ϵ)logQϵ(uk0ϵ)≤2GT(0,0),

where U(0,0)≤GT(0,0) and GT(0,0)⟶U(0,0) as T→∞.

The first statement above asserts that the limiting controlled trajectories exit the domain D through the boundary near the direction of the eigenvector e1f (see Hypotheses 3(c), (c’)). Finally, (30) shows that, for any finite time horizon T, our scheme is close to asymptotic optimality, according to Definition 3.2, and achieves optimal behavior in the limit ϵ→0,T→∞. Near asymptotic optimality is a common feature of importance sampling schemes for continuous-time dynamics even in finite dimensions. This is mainly a consequence of using subsolutions of (21) instead of exact solutions which are rarely given in explicit form. Our numerical studies indicate that near optimality leads to provably superior performance in comparison to standard Monte Carlo.

Remark 5

The moderate deviation regime allows us to work with the exit problem of a linear equation instead of that of the initial nonlinear SRDE (1). The "drift" of this linear equation is given by A+DF(x∗) and thus the dominant eigenpairs of this operator govern the exit time and exit place asymptotics. As mentioned in the introduction, similar statements have been proved for finite-dimensional linear equations in [51] (see e.g. Theorem 6).

On the asymptotic exit direction

In this section we study the limiting variational problem appearing on the right-hand side of (27). In particular, we will show that, under Hypothesis 3c’, changes of measure that force the dynamics in the e1f direction lead to minimal paths that exit from the ball B˚H(0,L) through the same direction. From this point on we will only use the notation Sx,T to denote the explicit action functional31 Sx,T(ϕ)=12∫0T‖ϕ˙(t)-[A+DF(Xx0(t))]ϕ(t)‖H2dt.

Moving on to the variational problem in (27), we let Ik0:T⊂C([0,T];H)→R,32 Ik0(y):=infv∈Cy,x∗∫0τ(12‖v(t)‖H2-‖uk0(t,y(t))‖H2)dt

and seek to characterize argminy∈TIk0(y). For the first part of this section we consider the case k0=1 covered by Hypothesis 3(c). The more general setting of Hypothesis 3c’ will be studied in Proposition 3.1 below. For the sake of simplicity we will drop the superscript k0 and write I≡I1 and u≡u1 unless otherwise stated.

A first observation is that I(y)<∞ if and only if y∈H1((0,T);H)∩L2([0,T];Dom(A)) and for all such y the infimum above is uniquely attained byv¯(t)=y˙(t)-Ay(t)-DF(x∗)y(t)-u(t,y(t)),t∈[0,T]

(see also Remark 4 above). Therefore, in view of (25), we can re-express I as follows:33 I(y)=∫0τ12‖v¯(t)‖H2-‖u(t,y(t))‖H2dt=∫0τ(12‖y˙(t)-Ay(t)-DF(x∗)y(t)+u(t,y(t))‖H2-‖u(t,y(t))‖H2)dt=∫0τ(12‖y˙(t)-Ay(t)-DF(x∗)y(t)‖H2-12‖u(t,y(t))‖H2)dt+∫0τ〈y˙(t),u(y(t))〉Hdt-∫0τ〈Ay(t)+DF(x∗)y(t),u(y(t))〉Hdt=Sx∗,τ(y)-∫0τ12‖u(t,y(t))‖H2dt+2a1f∫0τ⟨y(t),e1f⟩H⟨y˙(t),e1f⟩Hdt-2a1f∫0τ〈Ay(t)+DF(x∗)y(t),e1f〉H⟨y(t),e1f⟩Hdt=Sx∗,τ(y)+2a1f∫0τ⟨y(t),e1f⟩H⟨y˙(t),e1f⟩Hdt+2(a1f)2∫0τ⟨y(t),e1f⟩H2dt-∫0τ12‖u(t,y(t))‖H2dt.

The last two terms in the last display are equal due to (25). Thus,34 I(y)=Sx∗,τ(y)+a1f∫0τddt(⟨y(t),e1f⟩H2)dt=Sx∗,τ(y)+a1f(⟨y(τ),e1f⟩H2-⟨y(0),e1f⟩H2).

It is straightforward to verify that argminy∈TI(y)≠∅, i.e. the minimum value of I over the set T⊂C([0,T];H) is attained in T. Indeed, Sx∗,·:[0,T]×C([0,T];H)→[0,∞] is lower-semicontinuous and the second summand in (34) defines a continuous functional on the same set. Thus, I is itself lower-semicontinuous and furthermore T is closed in the topology of C([0,T];H) (recall that BH(0,L) in (26) is a closed ball in H).

Remark 6

We shall proceed to the characterization of minimizers in three steps. First we minimize over paths y with y(0)=0 and y(τ)=z∈∂BH(0,L). Then we minimize over the exit place z and finally over the time τ in which the path y hits the boundary ∂BH(0,L) of the closed ball BH(0,L). At this point, we emphasize that, in contrast to τ∗ϵ (14), τϕ (Lemma 3.1),τ^∗ϵ,v (Lemma 3.2) and τ^∗ϵ,vϵ (28), it is not known a priori whether the time τ is the first exit time of y from the open ball B˚H(0,L). We will show that the latter is true for minimizing paths in Lemma 3.4 and Proposition 3.1 below.

Lemma 3.3

Let y∗∈argmin{I(y):y∈C([0,τ];H),y(0)=0,y(τ)=z}. Theny∗(t)=yz,τ∗(t)=∑k=1∞sinh(akft)sinh(akfτ)⟨z,ekf⟩Hekf,t∈[0,τ].

Proof

The fact that we minimize over y∈C([0,τ];H) instead of C([0,T];H) is justified by Remark 4-3). Next notice that yz,τ∗∈C([0,τ];H) since sinh is increasing and continuous. In particular,‖yz,τ∗‖H≤∑k=1∞⟨z,ekf⟩H2=‖z‖H2.

Proceeding to the proof we have, in view of (34),I(y)=∫0τLx∗(y(t),y˙(t))dt+a1f⟨z,e1f⟩H2

with Lx∗ as in (22). Minimizers are then governed by the Euler-Lagrange equation35 ∂tDvLx∗(y(t),y˙(t))=DηLx∗(y(t),y˙(t))

which boils down to{y′′(t)=[A+DF(x∗)]2y(t),y(0)=0,y(τ)=z}.

Projecting to the eigenbasis {ekf}k∈N of A+DF(x∗) we obtaind2dt2⟨y(t),ek⟩=(akf)2⟨y(t),ek⟩,k∈N0,t∈[0,T].

Letting yk=⟨y,ekf⟩H,zk=⟨z,ekf⟩H, the general solution of the latter has the form yk(t)=c1eakft+c2e-akft and taking into account the initial and terminal conditions we obtainc1+c2=0,zk=c1eakfτ+c2e-akfτ⇒c1=zkeakτ-e-akfτ.

Thus,yk(t)=zk(eakft-e-akft)eakfτ-e-akfT=zksinh(akft)sinh(akfτ).

□

The next lemma is concerned with the exit direction when Hypothesis 3(c) holds.

Lemma 3.4

Let T>0, I as in (32) and u,T,Cy,x∗ as in Theorem 3.1. Under Hypothesis 3(c), any y∗∈argmin{I(y);y∈T}  y∗ first exits B˚H(0,L) at τ=T in the direction of the eigenvector e1f (recall Remark 6) i.e. for all k≥2,⟨y∗(τ),ekf⟩H=⟨y∗(T),ekf⟩H=0,

‖y∗(t)‖H<L for all t<T and ‖y∗(T)‖H=L.

Proof

Let ϕ∗=ϕz,τ∗ be a minimizer provided by Lemma 3.3. Notice that, since the Euler-Lagrange equations provide necessary conditions for minimality, any ϕ∗∈argmin{I(y):y∈C([0,τ];H),y(0)=0,y(τ)=z} will be of this form. After straightforward algebra we obtainI(ϕz,τ∗)=Sx∗,τ(ϕ∗)+a1f⟨z,e1f⟩H2=∑k=1∞∫0τ⟨ϕ˙∗(t)-[A+DF(x∗)]ϕ∗(t),ekf⟩H2dt+a1f⟨z,e1f⟩H2=∑k=1∞∫0τ(ϕ˙k∗(t)-akfϕk∗(t))2dt+a1f⟨z,e1f⟩H2=a1fz12+∑k=1∞akfzk21-e-2akfτ.

Now for each fixed τ, Hypothesis 3(c) guarantees that this quadratic form is minimized for z∗∈∂BH(0,L) such that zk∗=0 for all k≥2 and z1∗=±L (see e.g. Theorem 3.4 in [46]). Then,36 I(ϕz∗,τ∗)=a1fL2(1+11-e-2a1fτ)

is minimized for the largest possible τ i.e. for τ=T. Hence, since the order with which the variables are being minimized does not change the value of the minimum, we have miny∈TI(y)=I(ϕz∗,T∗) and the minimizers y∗=ϕz∗,T∗ enjoy the desired properties. Finally, note that any element y∗∈argmin{I(y);y∈T} is of the form ϕz∗,T∗. Indeed, fix the initial and terminal values y∗(0)=0,y∗(τ)=z∈∂BH(0,L) and assume that y∗ does not satisfy the Euler-Lagrange equations (35). Since the latter provide necessary conditions for minimality, it follows that y∗ is not a minimizer. Moreover, it follows from the previous calculations that if τ<T or if zk≠0 for some k≥2 then y∗ cannot be a minimizer of I. The proof is complete.□

As mentioned above, the previous lemma implies that, for any minimizing path y∗, τ is in fact the first exit time from the open ball B˚H(0,L), i.e. τ=inf{t∈[0,T]:y∗∉B˚H(0,L)} and furthermore τ=T.

Remark 7

If the sampling time T is large enough, the results of Lemma 3.4 as well as Theorems 3.1, 3.2 remain true under the weaker spectral gap assumption that 2a1f<akf for all k≥2. Since we are interested in schemes that perform well for large values of T,  this generalization comes at no cost. For more details on this relaxed condition see [46, Theorem 3.9].

Up to this point we have worked under Hypothesis 3(c) to show that minimizers of the functional I lie on the one-dimensional subspace where the change of measure u acts. In the absence of a sufficiently large spectral gap the situation is more complicated. In particular, if the sampling time T is large enough, the minimizers can be orthogonal to u. In other words, forcing the system towards its physical exit direction e1f might actually lead to controlled trajectories that exit from a subspace that is orthogonal to e1f under the change of measure. This is proved in the following lemma.

Lemma 3.5

Assume that the eigenvalues {akf}k∈N are strictly increasing, a2f≤2a1f and letT∗:=-12a2fln(1-a2f2a1f).

If T>T∗ then any minimizer y∗∈argmin{I(y);y∈T} satisfies ‖y∗(t)‖H<L for all t<T and ‖y∗(T)‖H=L. Moreover y∗ first exits B˚H(0,L) at τ=T in the direction of the eigenvector e2f (recall Remark 6) i.e. for all k≠2,⟨y∗(τ),ekf⟩H=⟨y∗(T),ekf⟩H=0.

Proof

As in the proof of Lemma 3.4 we haveI(ϕz,τ∗)=a1fz12+∑k=1∞akfzk21-e-2akfτ=:∑k=1∞λkfzk2.

We claim that, without loss of generality, we can consider τ∈(T∗,T]. Assuming the latter for now, we can compare the weights λkf to conclude that37 λ2f=a2f1-e-2a2fτ<a2f1-e-2a2fT∗=a2f1-eln(1-a2f/2a1f)=2a1f≤a1f(1+11-e-2a1fτ)=λ1f

and since x↦x/(1-e-2τx) is (strictly) increasing for all τ, it follows thatλ2f<λkf,∀k>2.

Therefore, the quadratic form is minimized for z∗∈∂BH(0,L) such that zk∗=0 for all k≠2 and z2∗=±L. Consequently38 inf(z,τ)∈∂BH(0,L)×[T∗,T]I(ϕz,τ∗)=a2fL21-e-2a2fT≥inf(z,τ)∈∂BH(0,L)×[0,T]I(ϕz,τ∗)

andinf(z,τ)∈∂BH(0,L)×[0,T]I(ϕz,τ∗)≥inf(z,τ)∈∂BH(0,L)×[0,T](infk∈Nλkf(τ))‖z‖H2=L2infτ∈[0,T](infk∈Nλkf(τ)).

Since λ2f≤λkf for all k≥2 it follows that39 inf(z,τ)∈∂BH(0,L)×[0,T]I(ϕz,τ∗)≥L2(λ1f(T)∧λ2f(T))=L2λ2f(T)=a2fL21-e-2a2fT,

which follows from (37) by setting τ=T>T∗. Since the infimum is achieved at t=T, the combination of (38) and (39) concludes the proof. □

Remark 8

Lemma 3.5 highlights the importance of sufficient spectral gaps for the design of efficient changes of measure. If Hypothesis 3(c) fails, a scheme that forces the e1f direction will be far from optimal and is expected to produce large errors for small values of ϵ. Under the assumptions of that lemma, one can repeat the arguments of the proof above to showlimϵ→0-1h2(ϵ)logQϵ(u1ϵ)=a2fL21-e-2a2fT<2GT(0,0).

If the ratio 2a1f/a2f is large, this bound translates to sub-optimal performance as ϵ→0 which does not improve as T→∞. Moreover, as we will see in Sect. 5, this ratio depends non-trivially on the interval length ℓ and is indeed large when ℓ is moderately small. This behavior is caused by the linearization of the dynamics and is completely absent when f=0. For an example that satisfies the assumptions of Lemma 3.5 see Sect. 5.1.

Before we conclude this section we consider once again the situation where the eigenvalues {akf}k∈N do not satisfy Hypothesis 3(c) but instead Hypothesis 3c’ holds. We show that the conclusions of Lemma 3.4 can be recovered by projecting to a higher dimensional eigenspace of A+DF(x∗) consisting of the first k0 eigenvalues.

Proposition 3.1

Let k0 as in Hypothesis 3c’, U as in (24) and uk0 as in (25). Under Hypothesis 3c’ any minimizer y∗∈argmin{Ik0(y);y∈T} satisfies the same properties as in Lemma 3.4.

Proof

Following the computations in (33), which carry over verbatim, we see that40 Ik0(y)=Sx∗,τ(y)+∑j=1k0[ajf(⟨y(τ),ejf⟩H2-⟨y(0),ejf⟩H2)].

Since the second term is constant for each fixed value of the exit point y(τ), the Euler-Lagrange equations and minimizers for this functional are then identical to those derived in Lemma 3.3 for I. Thus, for any minimizing path ϕz,τ∗ that hits the point z=(zk)k∈N∈∂BH(0,L) at time τ∈[0,T] we haveIk0(ϕz,τ∗)=∑j=1k0ajfzj2+∑j=1∞ajfzj21-e-2ajfτ=:∑j=1∞λk0,jfzj2.

Comparing the weights λk0,jf we see that for all 1<j≤k0λk0,1f=a1f(1+11-e-2a1fτ)<ajf(1+11-e-2ajfτ)=λk0,j,

which holds since x↦x/(1-e-2τx) is (strictly) increasing for all τ and a1f<a2f≤ajf for any j≥2. In order to show that minimizers point towards z1 it remains to compare λk0,1f with λk0,jf for j≥k0+1. Since λk0,k0+1≤λk0,k0+2≤⋯ it suffices to consider λk0,k0+1f. In view of Hypothesis 3c’ and Theorem 3.4 of [46] we conclude thatλk0,1f=a1f(1+11-e-2a1fτ)<ak0+1f1-e-2τak0+1f=λk0,k0+1f

for all τ∈[0,T]. The proof is complete.□

Tightness of η^x∗ϵ,vϵ

Let vϵ be a sequence in A satisfying the assumptions of Theorem 3.2, uk0 as in (25) and uk0ϵ:[0,T]×H→H be a sequence that converges pointwise and uniformly over bounded subsets of H to uk0. The goal of this section is to prove tightness estimates for the collection {η^x∗ϵ,vϵ:ϵ<ϵ0} of C([0,T];X)-valued random elements. Throughout the rest of this section we drop the index k0 and write u≡uk0,uϵ≡u0ϵ.

Recall that for each ϵ, η^x∗ϵ,vϵ is the unique mild solution of the controlled equation (18) with v=vϵ,u=uϵ. Existence and uniqueness is once again provided by Theorem 2.2 of [14] (see also Theorem 7.1 of [45]). The following lemma guarantees that, for ϵ small, the sequence vϵ is bounded in L2.

Lemma 3.6

There exists ϵ0>0 and a constant C>0 such thatsupϵ<ϵ0E∫0τ^∗ϵ,vϵ‖vϵ(s)‖H2ds≤C.

Proof

In view of the variational representation (17) any approximate minimizer vϵ∈A satisfiesE[12∫0τ^∗ϵ,vϵ‖vϵ(s)‖H2ds]≤-1h2(ϵ)logQϵ(u)+E∫0τ^∗ϵ,vϵ‖uϵ(s,η^∗ϵ,vϵ(s))‖H2ds+ϵ2

Now from the MDP for bounded functionals (see Definition 3.1 as well as Remark 11 below), along with Lemma 3.1, there exists a constant C>0 such that, for ϵ sufficiently small,-1h2(ϵ)logQϵ(u)≤C.

Hence, from the uniform convergence of uϵ to u and the uniform boundedness of u in bounded subsets of H and the fact that τ^∗ϵ,vϵ≤T with probability 1 the estimate follows. □

Remark 9

Without loss of generality, we can trivially extend the controls vϵ to [0, T] by letting vϵ(t)=0 for t∈[τ^∗ϵ,vϵ,T]. This convention will be in use for the rest of this section.

We shall now proceed to the proof of tightness estimates.

Lemma 3.7

Let p≥1. For all ϵ,T>0, there exist ϵ0>0,α,β>0 such that41 supϵ<ϵ0Esupt∈[0,T]‖η^∗ϵ,vϵ(t)‖Ep+supϵ<ϵ0Esupt∈[0,T]‖η^∗ϵ,vϵ(t)‖Ca+supϵ<ϵ0E‖η^∗ϵ,vϵ‖Cβ([0,T];E)<CT,ℓ,f

Proof

Using the mild formulation we have42 η^x∗ϵ,vϵ(t)=1ϵh(ϵ)∫0tS(t-s)[F(x∗+ϵh(ϵ)η^∗ϵ,vϵ(s))-F(x∗)]ds+∫0tS(t-s)[vϵ(s)-uϵ(s,η^∗ϵ,vϵ(s))]ds+1h(ϵ)∫0tS(t-s)dW(s)=:Ψϵ,vϵ(t)+Uϵ(t)+1h(ϵ)WA(t).

We now fix a version of the process Ψϵ,vϵ(t,ξ) and work path-by-path. The paths of Ψϵ,vϵ are weakly differentiable with probability 1 and∂tΨϵ,vϵ(t,ξ)=AΨϵ,vϵ(t,ξ)+1ϵh(ϵ)[F(x∗+ϵh(ϵ)η^∗ϵ,vϵ(t))-F(x∗)](ξ),

with A as in (4). Next, let t∈[0,T] and choose ξt∈[0,L] to be such that‖Ψϵ,vϵ(t)‖E=Ψϵ,vϵ(t,ξt)sign(Ψϵ,vϵ(t,ξt))

In view of Proposition A.1 in [45] (see also Proposition D.4 of [17]) we can estimate the left derivative of the supremum norm ‖Ψϵ,vϵ(t)‖E byd-dt‖Ψϵ,vϵ(t)‖E≤AΨϵ,vϵ(t,ξt)sign(Ψϵ,vϵ(t,ξt))+1ϵh(ϵ)[f(x∗(ξt)+ϵh(ϵ)η^∗ϵ,vϵ(t,ξt))-f(x∗(ξt))]sign(Ψϵ,vϵ(t,ξt)).

From the uniform ellipticity of A we have for all t∈[0,T],AΨϵ,vϵ(t,ξt)sign(Ψϵ,vϵ(t,ξt))≤0. Thus, in view of Hypothesis (2(a))d-dt‖Ψϵ,vϵ(t)‖E≤1ϵh(ϵ)[f1(x∗(ξt)+ϵh(ϵ)η^∗ϵ,vϵ(t,ξt))-f1(x∗(ξt))]sign(Ψϵ,vϵ(t,ξt))+1ϵh(ϵ)[f2(x∗(ξt)+ϵh(ϵ)η^∗ϵ,vϵ(t,ξt))-f2(x∗(ξt))]sign(ϵh(ϵ)Ψϵ,vϵ(t,ξt))≤Mf1|η^∗ϵ,vϵ(t,ξt)|+1ϵh(ϵ)[f2(x∗(ξt)+ϵh(ϵ)Ψϵ,vϵ(t,ξt)+ϵh(ϵ)Uϵ(t,ξt)+ϵWA(t,ξt))-f2(x∗(ξt))]·sign(ϵh(ϵ)Ψϵ,vϵ(t,ξt)),

where Mf1 is the Lipschitz constant of f1. To proceed, we distinguish the following two cases:

Case 1:sign(ϵh(ϵ)Ψϵ,vϵ(t,ξt))=sign(ϵh(ϵ)Ψϵ,vϵ(t,ξt)+ϵh(ϵ)Uϵ(t,ξt)+ϵWA(t,ξt)).

Since f2 is non-increasing,f2(x∗(ξt)+ϵh(ϵ)Ψϵ,vϵ(t,ξt)+ϵh(ϵ)Uϵ(t)+ϵWA(t))-f2(x∗(ξt))sign(ϵh(ϵ)Ψϵ,vϵ(t,ξt))≤0.

Hence,43 d-dt‖Ψϵ,vϵ(t)‖E≤Mf1|η^∗ϵ,vϵ(t,ξt)|≤Mf1‖η^∗ϵ,vϵ(t)‖E.

Case 2:sign(ϵh(ϵ)Ψϵ,vϵ(t,ξt))≠sign(ϵh(ϵ)Ψϵ,vϵ(t,ξt)+ϵh(ϵ)Uϵ(t,ξt)+ϵWA(t,ξt)).

In this case it is straightforward to verify thatϵh(ϵ)|Ψϵ,vϵ(t,ξt)|≤|ϵh(ϵ)Uϵ(t,ξt)+ϵWA(t,ξt)|.

The reader is referred to the proof of Theorem 6.1 of [45] for a similar argument. The latter, along with the optimality of ξt, yields44 ‖Ψϵ,vϵ(t)‖E=|Ψϵ,vϵ(t,ξt)|≤|Uϵ(t,ξt)+1h(ϵ)WA(t,ξt)|≤‖Uϵ(t)+1h(ϵ)WA(t)‖E.

Setting Ξϵ,vϵ(t):=max{‖Uϵ+WA/h(ϵ)‖C([0,T];E),‖Ψϵ,vϵ(t)‖E}, we can combine (43), (44) and the mean value inequality to obtainΞϵ,vϵ(t)-‖Uϵ+WA/h(ϵ)‖C([0,T];E)=Ξϵ,vϵ(t)-Ξϵ,vϵ(0)≤∫0td-dsΞϵ,vϵ(s)ds≤Mf1∫0t‖η^∗ϵ,vϵ(s)‖Eds≤Mf1∫0t[‖Ψϵ,vϵ(s)‖E+‖Uϵ+WA/h(ϵ)‖C([0,T];E)]ds≤2Mf1∫0tΞϵ,vϵ(s)ds.

By Grönwall’s inequality,‖Ψϵ,vϵ(t)‖E≤Ξϵ,vϵ(t)≤CT,ϕ‖Uϵ+WA/h(ϵ)‖C([0,T];E),

where CT,ϕ=e2Mf1T. Since the latter holds for all t∈[0,T] we obtain45 ‖Ψϵ,vϵ‖C([0,T];E)≤CT,ϕ‖Uϵ+WA/h(ϵ)‖C([0,T];E).

Turning to the control term,‖Uϵ(t)‖E≤C‖∫0tS(t-s)vϵ(s)ds‖Hθ(0,L)+‖∫0tS(t-s)uϵ(s,η^∗ϵ,vϵ(s))ds‖E

for any θ>1/2. This is a consequence of the embedding Wθ,p(O)↪E which holds for smooth domains O⊂Rd and all θ>d/p. From the smoothing property (6), the Cauchy-Schwarz inequality, the uniform convergence of uϵ to u and (25) we have46 ‖Uϵ(t)‖E≤CT,θ∫0t(t-s)-θ2‖vϵ(s)‖Hds+CT∫0t‖uϵ(s,η^∗ϵ,vϵ(s))‖Eds≤C(∫0t(t-s)-θds)12(∫0t‖vϵ(s)‖H2ds)12+CT∫0t(‖u(s,η^∗ϵ,vϵ(s))‖E+ρ)ds≤CθT(1-θ)/2‖vϵ‖L2([0,T];H)+ρTCT+2λ1f‖e1f‖E2∫0T‖η^∗ϵ,vϵ(s)‖Eds

which holds w.p. 1 for θ<1, ϵ sufficiently small and ρ>0. As for the stochastic convolution term we have h(ϵ)→∞ and (10) yieldsEsupt∈[0,T]‖WA(t)/h(ϵ)‖Ep≤C

for ϵ small and some C>0 independent of ϵ. The estimate is a consequence of the Sobolev embedding theorem along with heat kernel estimates and the stochastic factorization formula. Combining (42), (45), (46), Lemma 3.6 and Remark 9 we obtainEsupt∈[0,T]‖η^∗ϵ,vϵ(t)‖Ep≤CEsupt∈[0,T](‖Ψϵ,vϵ(t)‖Ep+‖Uϵ(t)‖Ep+‖WAϵ(t)/h(ϵ)‖Ep)≤Cp,T(1+1hp(ϵ)+∫0TEsups∈[0,t]‖η^∗ϵ,vϵ(s)‖Epdt)

and h(ϵ)→∞ as ϵ→0. Another application of Grönwall’s inequality leads to47 supϵ<ϵ0Esupt∈[0,T]‖η^∗ϵ,vϵ(t)‖Ep≤C

which is the first estimate in (41). Note here that C does not depend on x∗. Turning to the spatial Hölder regularity, an application of Taylor’s theorem for Gâteaux derivatives yields48 Ψϵ,vϵ(t)=1ϵh(ϵ)∫0tS(t-s)[F(x∗+ϵh(ϵ)η^∗ϵ,vϵ(s))-F(x∗)]ds=∫0tS(t-s)[DF(x∗)(η^∗ϵ,vϵ(s))+ϵh(ϵ)2D2×F(x∗+θ0ϵh(ϵ)η^∗ϵ,vϵ(s))(η^∗ϵ,vϵ(s),η^∗ϵ,vϵ(s))]ds

for some θ0∈(0,1). Let θ>1/2 and α=(2θ-1)/2. By virtue of the Sobolev embedding theorem (see e.g. Theorem 8.2 in [19]) and Hypothesis 2(b) we have‖Ψϵ,vϵ(t)‖Cα≤C‖∫0tS(t-s)[DF(x∗)(η^∗ϵ,vϵ(s))+ϵh(ϵ)D2F×(x∗+θ0ϵh(ϵ)η^∗ϵ,vϵ(s))(η^∗ϵ,vϵ(s),η^∗ϵ,vϵ(s))]ds‖Hθ≤C∫0t(t-s)-θ/2‖DF(x∗)(η^∗ϵ,vϵ(s))+ϵh(ϵ)D2F(x∗+θ0ϵh(ϵ)η^∗ϵ,vϵ(s))×(η^∗ϵ,vϵ(s),η^∗ϵ,vϵ(s))‖Hds≤Cf∫0t(t-s)-θ/2[(1+‖x∗‖Ep0-1)‖η^∗ϵ,vϵ(s)‖E+ϵh(ϵ)×(1+‖x∗+θ0ϵh(ϵ)η^∗ϵ,vϵ(s)‖Ep0-2)‖η^∗ϵ,vϵ(s)‖E2]ds≤Cf,θ,p0,x∗[1+supt∈[0,T]‖η^∗ϵ,vϵ(s)‖E0]t1-θ/2.

In view of (47),49 supϵ<ϵ0Esupt∈[0,T]‖Ψϵ,vϵ(t)‖Cα≤CT,f,θ,p0,x∗[1+supϵ<ϵ0Esupt∈[0,T]‖η^∗ϵ,vϵ(s)‖E0]<∞.

Repeating similar arguments to the ones used in (46) we see that50 supϵ<ϵ0Esupt∈[0,T]‖Uϵ(t)‖Cα≤CN,T,θ,f[1+supϵ<ϵ0Esupt∈[0,T]‖η^∗ϵ,vϵ(s)‖E]<∞.

Moreover, we have the following well-known spatial equicontinuity estimate for the stochastic convolution51 Esupt∈[0,T]‖WA(t)‖Cα≤C.

The reader is refered to [17], Theorems 5.16, 5.22 for the proof and a detailed discussion of regularity properties of stochastic convolutions. Combining the latter along with (49) and (50) we deduce that for each ϵ>0,t∈[0,T] η^∗ϵ,vϵ(t)∈Ca w.p. 1 and furthermoresupϵ<ϵ0Esupt∈[0,T]‖η^∗ϵ,vϵ(t)‖Ca<∞,

for some sufficiently small ϵ0. It remains to study the temporal equicontinuity of η^∗ϵ,vϵ. Letting s<t∈[0,T] we haveη^∗ϵ,vϵ(t)-η^∗ϵ,vϵ(s)-[S(t-s)-I]η^∗ϵ,vϵ(s)=1ϵh(ϵ)∫stS(t-r)[F(x∗+ϵh(ϵ)η^∗ϵ,vϵ(r))-F(x∗)]dr+∫stS(t-r)[vϵ(r)-uϵ(r,η^∗ϵ,vϵ(r))]dr+1h(ϵ)∫stS(t-r)dW(r)=:Ψϵ,vϵ(s,t)+Uϵ(s,t)+1h(ϵ)WA(s,t).

Hence,52 ‖η^∗ϵ,vϵ(t)-η^∗ϵ,vϵ(s)‖E≤‖Ψϵ,vϵ(s,t)‖E+‖Uϵ(s,t)‖E+‖WA(s,t)‖E+‖[S(t-s)-I]η^∗ϵ,vϵ(s)‖E.

From the estimates preceding (49) and the arguments in (46) we obtain53 supϵ<ϵ0Esups≠t∈[0,T]‖Ψϵ,vϵ(s,t)‖E|t-s|1-θ/2≤Cf,θ,p0,x∗[1+supϵ<ϵ0Esupt∈[0,T]‖η^∗ϵ,vϵ(s)‖E0]<∞

and54 supϵ<ϵ0Esups≠t∈[0,T]‖Uϵ(s,t)‖E|t-s|1-θ/2≤Cθ,N,T,f[1+supϵ<ϵ0Esupt∈[0,T]‖η^∗ϵ,vϵ(s)‖E]<∞

respectively. As for the stochastic convolution, there exists β∈(0,1) such that55 E[WA]Cβ([0,T];E)≤C

(see e.g. [17], Theorem 5.22). Finally, let θ>0,β∈(0,1/2) such that β+θ/2<1. From the Sobolev embedding theorem and (5)‖[S(t-s)-I]η^∗ϵ,vϵ(s)‖E≤C‖[S(t-s)-I]η^∗ϵ,vϵ(s)‖Hθ≤C‖[S(t-s)-I](-A)θ2η^∗ϵ,vϵ(s)‖H≤C‖S(t-s)-I‖L(Hβ;H)‖(-A)β+θ2η^∗ϵ,vϵ(s)‖H≤C(t-s)β‖η^∗ϵ,vϵ(s)‖H2β+θ.

Following the derivation of the estimates (49), (50), (51) (see also Lemma A.3 in [30]) we deduce thatEsups≠t∈[0,T]‖[S(t-s)-I]η^∗ϵ,vϵ(s)‖E|t-s|β0≤CEsups∈[0,T]‖η^∗ϵ,vϵ(s)‖H2β+θ≤C[1+1h(ϵ)+Esupt∈[0,T]‖η^∗ϵ,vϵ(s)‖E0]<∞.

From the latter and (52)-(55), there exists a sufficiently small ϵ0 and β>0 such thatsupϵ<ϵ0Esupt∈[0,T]‖η^∗ϵ,vϵ‖Cβ([0,T];E)<∞.

This proves the last estimate in (41) and completes the proof. □

From Lemma 3.7, along with an infinite-dimensional version of the Arzelà-Ascoli theorem, it follows that the family of laws of the controlled processes {η∗ϵ,vϵ}ϵ is concentrated on compact subsets of C([0,T];E), uniformly over sufficiently small values of ϵ. Thus, in view of Prokhorov’s theorem (Theorem 3.3 below), it forms a relatively compact set in the topology of weak convergence of measures in C([0,T];E). In the next section we aim to characterize the limit points as ϵ→0.

Limiting behavior of η^x∗ϵ,vϵ

Before we proceed to the main body of this section let us recall the notion of a tight family of probability measures and the classical theorem of Prokhorov.

Definition 3.4

Let Z be a Polish space and Π⊂P(Z) be a set of Borel probability measures on Z and {Pn}n∈N⊂Π. We say that (i) Pn converges weakly to a measure P∈P(Z) if for every f∈Cb(Z)limn→∞∫ZfdPn=∫ZfdP.

(ii) Π is tight if for each ϵ>0 there exists a compact set Kϵ⊂Z such that for all P∈Π,P(Z\Kϵ)<ϵ.

Prokhorov’s theorem asserts that the notions of tightness and relative weak sequential compactness are equivalent for Borel measures on Polish spaces.

Theorem 3.3

(Prokhorov) Let Z be a Polish space and Π⊂P(Z) be a tight family of Borel probability measures. Then every sequence in Π contains a weakly convergent subsequence.

Lemma 3.8

Let ϵ0 be sufficiently small, vϵ be a sequence in A satisfying the assumptions of Theorem 3.2, u as in (25) and uϵ:[0,T]×H→H be a sequence that converges pointwise and uniformly over bounded subsets of H to u. Any sequence in {(η^∗ϵ,vϵ,vϵ)}ϵ<ϵ0 has a further subsequence that converges in distribution in C([0,T];E)×L2([0,T];H) to a pair (η^∗0,v0) in the product of uniform and weak topologies. Moreover:

(i) η^∗0 is equal in law to the (unique) solution of56 {ϕ˙(t)=[A+DF(x∗)]ϕ(t)-u(t,ϕ(t))+v0(t),ϕ(0)=0},

(ii) Any sequence in {τ^∗ϵ,vϵ;ϵ<ϵ0} converges in distribution to a [0, T]-valued random variable τ^0 such thatη^∗0(τ^0)∈∂BH(0,L)

and for all t<τ^0, η^∗0(t)∈BH(0,L) with probability 1 (recall that BH(0,L) denotes a closed ball on H).

Proof

Starting from the controls vϵ, Lemma 3.6 along with Remark 9 yieldsupϵ>0E∫0T‖vϵ(t)‖H2dt<∞.

Since any bounded subset of L2([0,T];H) is relatively compact in the weak topology, we deduce from the discussion after Lemma 3.7 that the family of laws of the pairs {(η^∗ϵ,vϵ)}ϵ<ϵ0 is tight. By virtue of Prokhorov’s theorem any sequence of such elements contains a subsequence (denoted with the same notation) that converge in distribution to a pair (η^x∗,v0) of C([0,T];E)×L2([0,T];H)-valued random elements. We remark here that L2([0,T];H) with the weak topology is not globally metrizable, hence not a Polish space, and Prokhorov’s theorem is not directly applicable. However the same conclusions can be drawn by a more general version of the theorem (e.g. Theorem 8.6.7 in [8]). Invoking Skorokhod’s theorem we can now assume that this convergence happens almost surely. This theorem involves the introduction of a new probability space with respect to which the convergence takes place. This will not be reflected in our notation for the sake of convenience. We will now characterize the law of η^x∗.

(i) Recall that for all t∈[0,T]η^x∗ϵ,vϵ(t)=1ϵh(ϵ)∫0tS(t-s)[F(x∗+ϵh(ϵ)η^∗ϵ,vϵ(s))-F(x∗)]ds+∫0tS(t-s)vϵ(s)ds-∫0tS(t-s)uϵ(s,η^∗ϵ,vϵ(s))+1h(ϵ)WA(t)

with probability 1. Starting from the last term, the estimate (10) yields 1h(ϵ)WA⟶0 in Lp(Ω;C([0,T];E)) for any p≥1. Next, from Lemma 4.7 in [46] we have∫0·S(·-s)vϵ(s)ds⟶∫0·S(·-s)v0(s)ds

almost surely in C([0,T];E). As for the term involving the changes of measure uϵEsupt∈[0,T]‖∫0tS(t-s)[uϵ(s,η^∗ϵ,vϵ(s))-u(s,η^x∗(s))]ds‖E≤cE∫0T‖uϵ(s,η^∗ϵ,vϵ(s))-u(s,η^∗ϵ,vϵ(s))‖Eds+cE∫0T‖u(s,η^∗ϵ,vϵ(s))-u(s,η^x∗(s))‖Eds.

The first term on the right hand side converges to 0 by our assumptions along with (47). The almost sure convergence of η^∗ϵ,vϵ and the continuity of u (see (25)) along with the dominated convergence theorem imply the convergence of the second term to 0. Next, in view of (48), Hypothesis 2(b) and the dominated convergence theorem we haveEsupt∈[0,T]‖1ϵh(ϵ)∫0tS(t-s)[F(x∗+ϵh(ϵ)η^∗ϵ,vϵ(s))-F(x∗)-DF(x∗)(η^x∗(s))]ds‖E⟶0

as ϵ→0. Uniqueness of (56) along with a subsequence argument complete the proof.

(ii) Since [0, T] is compact in the standard topology, the family of [0, T]-valued random variables {τ^∗ϵ,vϵ}ϵ<ϵ0 is tight. Invoking Prokhorov’s and Skorokhod’s theorems once again, any sequence in this family has a subsequence that converges almost surely to a [0, T]-valued random variable τ^∗0. From the almost sure convergence of η^∗ϵ,vϵ and the definition of τ^∗ϵ,vϵ (see Lemma 3.2), η^∗ϵ,vϵ(τ^∗ϵ,vϵ)⟶η^x∗0(τ^∗0)∈∂BH(0,L) almost surely (the latter being a closed set ). Moreover, for any t<τ^∗0, there exists δ>0 and ϵ0>0 sufficiently small such that t≤τ^∗0-δ<τ^∗ϵ,vϵ for all ϵ≤ϵ0 on a set of probability 1. Thus, for ϵ sufficiently small, {η^∗ϵ,vϵ(t)}ϵ⊂BH(0,L) and the pointwise limit η^∗ϵ,v0(t)∈BH(0,L) with probability 1. □

Remark 10

A simple consequence of Lemma 3.8 is that the moderate deviation process ηxϵ (3) which results by setting u=vϵ=0 in (18), converges as ϵ→0 to the solution of the linear deterministic PDE ϕ˙(t)=[A+DF(Xx0(t))]ϕ(t) with zero initial condition, i.e. ηxϵ→0.

Proof of Theorem 3.1

Before we move on to the proof we remind the reader that the index k0 has been dropped.

Let ϵ>0. Returning to (17), choose a sequence {vϵ}⊂A of approximate minimizers such that (28) holds. Since uϵ converges uniformly to u over bounded subsets, there exists ϵ0 sufficiently small such that for any δ>0 and ϵ<ϵ057 -1h2(ϵ)logQϵ(uϵ)≥E[12∫0τ^∗ϵ,vϵ‖vϵ(s)‖H2ds-∫0τ^∗ϵ,vϵ‖u(s,η^∗ϵ,vϵ(s))‖H2ds]-ϵ-δ.

From the variational representation (17), Lemma 3.6 and the assumptions on uϵ and u there exists ϵ0 sufficiently small such thatsupϵ<ϵ0|1h2(ϵ)logQϵ(uϵ)|≤supϵ<ϵ0E12∫0τ^∗ϵ,vϵ‖vϵ(s)‖H2ds+supϵ<ϵ0E∫0T‖uϵ(s,η^∗ϵ(s))‖H2ds<∞.

Thus, there exists a sequence in ϵ over which the left hand side in (57) converges to lim infϵ→0-logQϵ(uϵ)/h2(ϵ). Since the functional J:C([0,T];E)×L2([0,T];H)×[0,T]→R,J(η,v,τ):=12∫0τ‖v(s)‖H2ds-∫0τ‖u(η(s))‖H2ds

is lower semi-continuous in the product of uniform, weak and standard topologies, we can pass to a further subsequence and apply the Portmanteau lemma along with Lemma 3.8 to obtain58 lim infϵ→0-1h2(ϵ)logQϵ(uϵ)≥lim infϵ→0E[J(η^∗ϵ,vϵ,vϵ,τ^∗ϵ,vϵ)]-δ≥E[J(η^∗0,v0,τ^∗0)]-δ=E[12∫0τ^∗0‖v0(s)‖H2ds-∫0τ^∗0‖u(η^∗0(s))‖H2ds]-δ≥infy∈Tinfv∈Cy,x∗∫0τ(12‖v(s)‖H2-‖u(y(s))‖H2)ds-δ,

with T as in (26). Since δ is arbitrary, the upper bound is complete. To obtain a lower bound we will use the conclusions of Proposition 3.1 for the limiting variational problem. To this end let y∗ satisfyinfv∈Cy∗,x∗∫0y∗(12‖v(s)‖H2-‖u(y∗(s))‖H2)ds=infy∈Tinfv∈Cy,x∗∫0τ(12‖v(s)‖H2-‖u(y(s))‖H2)ds.

As we mentioned in Sect. 3.2, the optimization problem on the left-hand side has an explicit solution attained by59 v¯(t)=y˙∗(t)-Ay∗(t)-DF(x∗)y∗(t)+u(y∗(t)),t∈[0,T]

and from Proposition 3.1, T=inf{t>0:‖y∗(t)‖H=L}=τy∗. Now consider the processes η^∗ϵ,v¯ controlled by v¯. From Lemmas 3.7, 3.8, {η^∗ϵ,v¯;ϵ>0} is tight and converges in distribution to a process η^∗v¯. From the choice of v¯ and uniqueness of solutions it follows that η^∗v¯=y∗ with probability 1. Moreover, the exit times τ^∗ϵ,v¯ converge in distribution to a random time τ^v¯ which is no less than the first exit time of y∗ from B˚H(0,L). Since the latter is equal to T it follows that τ^v¯=T with probability 1. Thus60 lim supϵ→0-1h2(ϵ)logQϵ(uϵ)≤12∫0τ^∗ϵ,v¯‖v¯(s)‖H2ds+lim supϵ→0-E∫0τ^∗ϵ,v¯‖uϵ(η^∗ϵ,v¯)‖H2ds≤12∫0T‖v¯(s)‖H2ds-lim infϵ→0E∫0τ^∗ϵ,v¯‖uϵ(η^∗ϵ,v¯)‖H2ds≤12∫0T‖v¯(s)‖H2ds-∫0T‖u(η^∗v¯)‖H2ds=infv∈Cy∗,x∗∫0y∗(12‖v(s)‖H2-‖u(y∗(s))‖H2)ds=infy∈Tinfv∈Cy,x∗∫0τ(12‖v(s)‖H2-‖u(y(s))‖H2)ds,

where the second inequality follows from lower semi-continuity. Combining (58) and (60) allows us to conclude.

Remark 11

Theorem 3.1 is essentially equivalent to an MDP for the family {Xϵ}ϵ of solutions of (8), in the space C([0,T];E). The latter is an asympotic statement for exponential functionals of g(Xϵ), where g:C([0,T];E)→R is continuous and bounded (see Definition 3.1), while the former covers exit probabilities and corresponds to the choice g=g~ withg~(η)=0,η:supt∈[0,T]‖η(t)‖H≥L∞,η:supt∈[0,T]‖η(t)‖H<L.

The case for bounded continuous test functions is in fact simpler, does not require analysis of the limiting variational problem and can be proved using very similar arguments to the ones used above. To be precise, for any continuous, bounded g:C([0,T];E)→R the variational representation (17) takes the form-1h2(ϵ)logE[e-h2(ϵ)g(ηϵ)]=infv∈AE[12∫0T‖v(t)‖H2dt+g(ηϵ,v)],

according to the classical results of [12]. The controlled process ηϵ,v solves (18) with u=0 and A is a collection of square-integrable adapted controls. The tightness and limiting statements of Lemmas 3.7, 3.8 carry over verbatim after setting u=0 and (11) then follows with the same action functional (31) by proving an upper and a lower bound as above. In particular, the upper bound is a consequence of lower-semicontinuity and the lower bound follows by considering the minimizing control v¯ in (59). In fact, this simpler MDP is used to obtain Lemma 3.6 above, which is important for the case of unbounded functionals that we consider here.

Proof of Theorem 3.2

Let {vϵ}⊂A satisfy (28). From Lemma 3.8, Theorem 3.1 and the lower semi-continuity argument in (58) we know that the triples (η^∗ϵ,vϵ,vϵ,τ^∗ϵ,vϵ) converge in distribution to a triple (η^∗0,v0,τ^∗0) andinfy∈Tinfv∈Cy,x∗∫0τ(12‖v(s)‖H2-‖u(y(s))‖H2)ds=lim supϵ→0-1h2(ϵ)logQϵ(uϵ)≥lim supϵ→0E[12∫0τ^∗ϵ,vϵ‖vϵ(s)‖H2ds-∫0τ^∗ϵ,vϵ‖uϵ(s,η^∗ϵ,vϵ(s))‖H2ds]≥lim infϵ→0E[12∫0τ^∗ϵ,vϵ‖vϵ(s)‖H2ds-∫0τ^∗ϵ,vϵ‖uϵ(s,η^∗ϵ,vϵ(s))‖H2ds]=E[12∫0τ^∗0‖v0(s)‖H2ds-∫0τ^∗0‖u(η^∗0(s))‖H2ds].

Invoking Lemma 3.8 once again we have η^∗0∈T and v0∈Cη^∗0,x∗ with probability 1. Since the left-hand side is the infimum over all such paths and controls it follows that12∫0τ^∗0‖v0(s)‖H2ds-∫0τ^∗0‖u(η^∗0(s))‖H2ds=infy∈Tinfv∈Cy,x∗∫0τ(12‖v(s)‖H2-‖u(y(s))‖H2)ds

with probability 1. Thus, from Proposition 3.1 we can conclude that τ^∗ϵ,vϵ→T in probability as ϵ→0, 〈η^∗0(T),e1f〉H2=L2 with probability 1 and (29) follows.

It remains to prove (30). We start from the upper bound which is a consequence of Lemma 3.1, provided that E={ϕ∈C([0,T];H):τϕ≤T} is a Sx∗,T-continuity set. This property can be verified from the analysis of Sect. 3.2. In particular, Lemmas 3.3, 3.4 and Proposition 3.1 remain true after setting the second summand in (34) or (40) equal to 0. Hence the infima of the action functional over {τϕ≤T},{τϕ<T} and {τϕ=T} coincide and the estimate follows. As for the lower bound, we combine Theorem 3.1, (36) and (24) to obtainGT(0,0)=a1fL21-e-2a1fT≥a1fL2=U(0,0),limT→∞GT(0,0)=U(0,0)

andlimϵ→0-1h2(ϵ)logQϵ(uϵ)=a1fL2(1+11-e-2a1fT)=U(0,0)+GT(0,0).

The latter shows that the lower bound actually holds with equality, hence the proof is complete.

Implementation and pre-asymptotic analysis of the scheme

Implementation issues and exponential mollification

In Sect. 3, we demonstrated that, under fairly general spectral gap conditions, an importance sampling scheme using the change of measure uk0 (25) achieves nearly optimal asymptotic behavior as the noise intensity ϵ→0. However, changes of measure based only on the quasipotential subsolution U (24) can lead to poor pre-asymptotic performance. This issue is present even in finite dimensions and is related to the behavior of the controlled dynamics near the origin. In [23], the authors demonstrated that, for certain choices of controls v, the second moment of the estimator degrades over time. In these situations, the system tends to spend a large amount of time near the attractor thus accumulating a large running cost which affects the variance. As a result, for fixed ϵ>0 the pre-exponential terms which are ignored by the asymptotic bounds (30) dominate and can even lead to errors that increase exponentially as T grows. For more details the reader is referred to the discussion in [23] pp.2919-2921.

In infinite dimensions, an additional challenge appears when the changes of measure act on the full space H. As we will see in Lemma 4.1 below, in order to prove that the second moment of a scheme behaves well for any fixed ϵ>0, one needs to have good control over the quantity61 Dx∗ϵ(Zx∗)(t,η):=∂tZx∗(t,η)+Hx∗(η,DηZx∗(t,η))+12h2(ϵ)tr[Dη2Zx∗(t,η)],

where Zx∗ denotes a subsolution used for the analysis of the scheme. However, any radial function Z:H→R such that Z(η)=Z¯(‖η‖H), with Z¯′′<0, satisfies tr[Dη2Z(η)]=-∞. Thus, apart from dealing with the difficulties related to unbounded operators (see Remark 3), changes of measure for SRDEs that effectively accomplish dimension reduction are necessary for provably efficient performance.

In this section we construct a scheme under Hypothesis 3(c), i.e. our changes of measure only force the e1f direction. From this point on it is understood that u≡u1 and u1ϵ≡uϵ. In order to deal with the aforementioned issues, our changes of measure uϵ will meet the following criteria: 1) The projected-quasipotential subsolution (denoted below by F1) will be used for regions of space that are sufficiently far from the origin. 2) A constant subsolution F2ϵ will dominate near zero. F2ϵ does not influence the dynamics until they enter the domain where F1 dominates. 3) To avoid issues from lack of smoothness, the combination of F1,F2 should be appropriately mollified. 4) As ϵ→0 the changes of measure uϵ converge to the asymptotically nearly optimal u. A suitable choice is provided by the exponential mollification of F1,F2ϵ.

To be precise, we define for a1f,e1f as in Hypothesis 3(c), κ∈(0,1) and δ=δ(ϵ)>0F1(η):=a1f(L2-⟨η,e1f⟩H2),F2ϵ:=a1f(L2-h(ϵ)-2κ),η∈H

and consider the exponential mollification62 Uδ(t,η):=-δlog(e-F1(η)δ+e-F2ϵδ).

We implement our scheme using the change of measure63 uϵ(t,η):=-DηUδ(t,η)=-2a1fρϵ(η)⟨e1f,η⟩He1f,

whereρϵ(η):=e-F1(η)δe-F1(η)δ+e-F2ϵδ,

δ=2/h2(ϵ) is the mollification parameter and κ is a parameter that controls the size of the neighborhood outside of which F1 dominates.

In order to derive non-asymptotic bounds for the second moment of the estimator, we will use the following min/max representation for the HamiltonianHx∗(η,p)=infvsupu[⟨p,Aη+DF(x∗)η-u+v⟩H-12‖u‖H2+14‖v‖H2]

(see e.g. [23, 24]) and for any smooth functions Ux∗,Zx∗:[0,T]×H→R we let u(t,η)=-DηUx∗(t,η) and p=DηZx∗(t,η). Thus we obtain64 infv[〈DηZx∗(t,η),Aη+DF(x∗)η-u(t,η)+v〉H-12‖u(t,η)‖H2+14‖v‖H2]=Hx∗(η,DηZx∗(t,η))-12‖DηZx∗(t,η)-DηUx∗(t,η)‖H2.

A consequence of this expression is the following pre-asymptotic bound for the second moment:

Lemma 4.1

For any smooth functions Ux∗,Zx∗:[0,T]×H→R, Dx∗ as in (61) and some θ0∈(0,1) let65 Hx∗ϵ(Zx∗)(t,η):=ϵh(ϵ)2〈DηZx∗(t,η),D2F(x∗+θ0ϵh(ϵ)η)(η,η)〉H

and66 Hx∗ϵ(Ux∗,Zx∗)(t,η):=Hx∗ϵ(Zx∗)(t,η)+Dx∗ϵ(Zx∗)(t,η)-12‖DηZx∗(t,η)-DηUx∗(t,η)‖H2.

For all ϵ>0 we have67 -1h2(ϵ)logQϵ(uϵ)≥infv∈A[2Zx∗(0,0)-2EZx∗(τ^∗ϵ,v,η^x∗ϵ,v(τ^∗ϵ,v))+2E∫0τ^∗ϵ,vHx∗ϵ(Ux∗,Zx∗)(s,η^∗ϵ,v(s))ds].

The proof makes use of Itô’s formula and is deferred to Appendix A.

Remark 12

The term Hx∗ϵ accounts for the error coming from the local approximation of the nonlinear dynamics by their linearized version around the stable equilibrium x∗. A significant part of this section is devoted to the pre-asymptotic control of this term.

The rest of this section is devoted to the pre-asymptotic analysis of Qϵ(uϵ) based on the lower bound (67) with Ux∗=Uδ(t,η),68 Z(t,η)=Zx∗(t,η)=(1-ζ)Uδ(t,η),ζ∈(0,1).

Performance analysis of the scheme

At this point we shall recall the definition of the random timesτ^∗ϵ,v=inf{t>0:η^∗ϵ,v(t)∉B˚H(0,L)}

where η^∗ϵ,v solves (18). Before we state the main result of this section, we provide the definition of exponential negligibility; a concept which will be frequently used in the sequel.

Definition 4.1

A term will be called exponentially negligible (a) in the moderate deviations range if it can be bounded from above in absolute value by C1e-c2h2(ϵ)/h2(ϵ) where C1<∞,c2>0 (b) in the large deviations range if (a) holds with 1/h2(ϵ) replaced by ϵ.

Remark 13

Since ϵh(ϵ)→0 as ϵ→0, exponential negligibility in the large deviations range implies exponential negligibility in the moderate deviations range.

The analysis of this section is summarized in the following theorem. Its proof is postponed for the end of this section and is preceded by several auxiliary estimates.

Theorem 4.1

Let T,α,ζ0,ϵ>0 and uϵ(t,η)=-DηUδ(ϵ)(t,η) with Uδ defined in (62). Assume that δ=2/h2(ϵ),κ∈(0,1-α),ζ∈(ζ0,1/2) and ϵ is sufficiently small to have h2(κ+α-1)(ϵ)≤9a1f2(ζ0-2ζ02)∧a1f2. Then, up to exponentially negligible terms in the moderate deviations range,69 -1h2(ϵ)logQϵ(uϵ)≥[(1-ζ)a1f(L2-1h(ϵ)2κ)-2log2h2(ϵ)]-CTϵh(ϵ).

‘, if h(ϵ) is such that ϵh3(ϵ)⟶0 as ϵ→0 then for ϵ sufficiently small we have70 -1h2(ϵ)logQϵ(uϵ)≥[(1-ζ)a1f(L2-1h(ϵ)2κ)-2log2h2(ϵ)].

Remark 14

Note that for a small fixed ϵ, (69) shows that, in theory, the second moment degrades as the sampling time T grows. This degradation is caused by the linearization error (65) and suggests that, in practice, good performance lies in the balance between ϵ and T. Fortunately, (70) shows that this theoretical degradation is no longer present if the scaling h(ϵ) does not grow too fast. Moreover, the simulation studies of Sect. 6 show that our scheme performs well for large T even when this growth assumption is not satisfied.

The following lemma collects a few straightforward computations that will be used below. Its proof can be found in Appendix A.

Lemma 4.2

For all (t,η)∈[0,T]×H,ζ∈(0,1), Uδ,Z as in (62), (68) and some θ0∈(0,1) we have71 Hx∗ϵ(Uδ,Z)(t,η)=2(1-ζ)(a1f)2ρϵ(η)⟨e1f,η⟩H2[1-(1-ζ)ρϵ(η)]-2ζ2(a1f)2(ρϵ(η))2⟨e1f,η⟩H2-(1-ζ)a1fρϵ(η)h2(ϵ)[1+2δρϵ(η)(1-ρϵ(η))⟨e1f,η⟩H]-ϵh(ϵ)2(1-ζ)a1fρϵ(η)⟨e1f,η⟩H×⟨e1f,D2F(x∗+θ0ϵh(ϵ)η)(η,η)⟩H.

Moving on to the main body of the analysis, let B∞(x∗,1) denote an open L∞-ball of radius 1 centered at x∗ andτ∞ϵ=inf{t>0:‖η^x∗ϵ,v(t)‖L∞≥1ϵh(ϵ)}=inf{t>0:X^x∗ϵ,v(t)∈B∞(x∗,1)c}.

Returning to (67) we have the following decomposition72 -1h2(ϵ)logQϵ(uϵ)≥infv∈A[2Zx∗(0,0)-2EZx∗(τ^∗ϵ,v,η^x∗ϵ,v(τ^∗ϵ,v))+2E∫0τ^∗ϵ,vHx∗ϵ(Ux∗,Zx∗)(s,η^∗ϵ,v(s))ds]=infv∈A[2Zx∗(0,0)-2EZx∗(τ^∗ϵ,v,η^x∗ϵ,v(τ^∗ϵ,v))+2E∫0τ^∗ϵ,v∧τ∞ϵHx∗ϵ(Ux∗,Zx∗)(s,η^∗ϵ,v(s))ds+2E∫τ^∗ϵ,v∧τ∞ϵτ^∗ϵ,vHx∗ϵ(Ux∗,Zx∗)(s,η^∗ϵ,v(s))ds].

Remark 15

This decomposition allows us to deal with the cubic power of η that appears in (65). Since we are only controlling the spatial L2-norm of the moderate deviation process, this term is problematic. In particular, estimates based in the a-priori bound (47) will introduce T-dependent constants which are not desirable for the pre-asymptotic analysis.

The last term in (72) concerns the behavior of the controlled process η^ϵ,v in the event that it exits an L∞-ball of radius 1/ϵh(ϵ) before it exits B˚H(0,L). Since the latter is a very rare event in the moderate deviations range, we expect that this term is exponentially negligible. This claim is proved in the following proposition.

Proposition 4.1

The term2E∫τ^∗ϵ,v∧τ∞ϵτ^∗ϵ,vHx∗ϵ(Uδ,Z)(s,η^∗ϵ,v(s))ds

is exponentially negligible in the moderate deviations range for ϵ sufficiently small.

Proof

Let t∈[0,T],η∈L∞∩BH(0,L), ϵ small enough to have ϵh(ϵ)<1. In view of (7),1ϵh(ϵ)|Hx∗ϵ(Z)(t,η)|≤2(1-ζ)a1fρϵ(η)|⟨e1f,×η⟩H||⟨D2F(x∗+θϵh(ϵ)η)(η,η),e1f⟩H|≤2a1f‖η‖H3‖e1f‖H‖∂x2f(x∗+θϵh(ϵ)η)e1f‖L∞≤2Cf,pa1fL3‖e1f‖L∞(1+‖x∗‖∞p0-2+‖η‖∞p0-2).

Moreover, from (71) we have|Hx∗ϵ(Uδ,Z)(t,η)-Hx∗ϵ(Z)(t,η)|≤2(1-ζ)(a1f)2ρϵ(η)⟨e1f,η⟩H2[1-(1-ζ)ρϵ(η)]+2ζ2(a1f)2(ρϵ(η))2⟨e1f,η⟩H2+(1-ζ)a1fρϵ(η)h2(ϵ)[1+2δρϵ(η)(1-ρϵ(η))⟨e1f,η⟩H]≤4(a1f)2‖e1f‖H2‖η‖H2+a1fh2(ϵ)[1+h2(ϵ)‖e1f‖H‖η‖H]≤Cℓ,f,L,

where we used that ζ,ρ∈(0,1), and h(ϵ)>1. Combining the last two estimates we deduce that for any v∈A,|∫τ^∗ϵ,v∧τ∞ϵτ^∗ϵ,vHx∗ϵ(Uδ,Z)ds|≤1{τ∞ϵ≤τ^∗ϵ,v}∫τ∞ϵτ^∗ϵ,v|Hx∗ϵ(Uδ,Z)(s,η^∗ϵ,v(s))|ds≤Cf,p0,L,ℓ1{τ∞ϵ≤τ^∗ϵ,v}∫0T(1+‖x∗‖∞p0-2+‖η^∗ϵ,v(s)‖∞p0-2)ds≤Cf,p0,L,T1{τ∞ϵ≤T}(1+‖x∗‖∞p0-2+sups∈[0,T]‖η^∗ϵ,v(s)‖∞p0-2).

An application of Hölder’s inequality along with (41) yields|E∫τ^∗ϵ,v∧τ∞ϵτ^∗ϵ,vHx∗ϵ(Uδ,Z)(s,η^∗ϵ,v(s))ds|≤Cf,p0,L,ℓ,TP[τ∞ϵ≤T]12(1+‖x∗‖∞p0-2+E[sups∈[0,T]‖η^∗ϵ,v(s)‖∞2p0-4]12)≤Cf,p,L,T,x∗P[τ∞ϵ≤T]12.

Recall now that X^x∗ϵ,v solves{dX^ϵ(t)=[AX^ϵ(t)+F(X^ϵ(t))]dt+ϵh(ϵ)[v(t)-uϵ(η^∗ϵ,v(t))]dt+ϵdW(t),X^ϵ(0)=x∗}

and, as ϵ→0, {X^x∗ϵ,v}ϵ>0 satisfies a large deviation principle in C([0,T];L∞(0,ℓ)) with action functional S~x∗,T:C([0,T];L∞(0,ℓ))→[0,∞] given byS~x∗,T(ϕ)=infu∈Pϕ12∫0T‖u(t)‖H2dt,Pϕ={u∈L2([0,T];H):∀t∈[0,T]ϕ(t)=S(t)x∗+∫0tS(t-s)[F(ϕ(s))+u(s)]ds},

where the convention inf∅=+∞ is in use (see e.g. [14], Theorems 6.2, 6.3). Passing to a convergent subsequence if necessary, we deduce thatlimϵ→0ϵlogP[τ∞ϵ≤T]≤-infϕ∈B∞(x∗,L)cS~x∗,T(ϕ),

where B∞(x∗,L):={ϕ∈C([0,T];H):supt∈[0,T]‖ϕ(t)-x∗‖L∞<1}. Hence, for ϵ sufficiently smallϵlogP[τ∞ϵ≤T]≤-infϕ∈B∞(x∗,L)cS~x∗,T(ϕ)/2

or equivalentlyP[τ∞ϵ≤T]12≤e-infϕ∈B∞(x∗,L)cS~x∗,T(ϕ)/4ϵ.

Finally, we claim that infϕ∈B∞(x∗,L)cS~x∗,T(ϕ)>0. Indeed, since the action functional is lower semi-continuous (see Lemma 5.1, [14]) and B∞(x∗,L)c⊂C([0,T];L∞(0,ℓ)) is closed, there exists a minimizer ϕ∗∈B∞(x∗,L)c. Furthermore, there exists u∗∈Pϕ∗ such that2infϕ∈B∞(x∗,L)cS~x∗,T(ϕ)=2S~x∗,T(ϕ∗)>12∫0T‖u∗(t)‖H2dt=12‖u∗‖L2([0,T];H)2>0.

The last inequality fails if and only if u∗=0 almost everywhere in [0,T]×[0,ℓ]. Since x∗ is an equilibrium of the uncontrolled system, the latter implies that ϕ∗(t)=x∗ for all t∈[0,T], hence ϕ∗∉B∞(x∗,L)c. This contradicts the initial choice of ϕ∗ and concludes the argument. Therefore, the term of interest is exponentially negligible in the large deviation range hence also in the moderate deviation range. □

Next, we turn our attention to the third term in (72). The linearization error in this term is easier to control, since the process ϵh(ϵ)η^ϵ,v is uniformly bounded by 1 in L∞-norm. This fact is used in the following lemma whose proof can be found in Appendix 1.

Lemma 4.3

For all η∈B∞(0,1/ϵh(ϵ)) there exists a constant C=Cx∗,ℓ,f>0 such that for ϵ sufficiently small we have73 Hx∗ϵ(Z)(t,η)≥-Cϵh(ϵ)2(1-ζ)a1fρϵ(η)|⟨e1f,η⟩H|.

As for Hx∗ϵ(Uδ,Z)(t,η)-Hx∗ϵ(Zx∗)(t,η), straightforward algebra along with the arguments of Lemma 4.2 of [23] yieldHx∗ϵ(Uδ,Z)(t,η)-Hx∗ϵ(Z)(t,η)≥Dx∗ϵ(Z)(t,η)-ζ22‖DηUδ(t,η)‖H2≥(1-ζ)Dx∗ϵ(Z)(t,η)-ζ-2ζ22‖DηUδ(t,η)‖H2≥1-ζ2(1-1h2(ϵ)δ)β0ϵ(η)+(1-ζ)ρϵ(η)γ1+ζ-2ζ22ρϵ(η)2‖DηF1(η)‖H2,

where the quantity β0(η):=ρϵ(η)(1-ρϵ(η))‖DηF1(η)‖H2 is nonnegative since ρϵ∈[0,1] and γ1:=Dx∗(F1)(η)=-a1f/h2(ϵ). Combining the latter with (73) and substituting δ=2/h2(ϵ) and‖DηF1(η)‖H2=4(a1f)2⟨η,e1f⟩2‖e1f‖H2=4(a1f)2⟨η,e1f⟩2,

we obtain the lower bound74 Hx∗ϵ(Uδ,Z)(t,η)≥(1-ζ)ρϵ(η)(1-ρϵ(η))(a1f)2⟨η,e1f⟩2-1h2(ϵ)(1-ζ)ρϵ(η)a1f+2(ζ-2ζ2)ρϵ(η)2(a1f)2⟨η,e1f⟩2-2Cϵh(ϵ)(1-ζ)a1fρϵ(η)|⟨e1f,η⟩H|.

At this point we partition BH(0,L)=B1f∪B2f∪B3f, where75 B1f:={η∈H:⟨η,e1f⟩H2≤h(ϵ)-2(κ+α)},B2f:={η∈H:h(ϵ)-2(κ+α)<⟨η,e1f⟩H2≤2h(ϵ)-2κ-h(ϵ)-2K},B3f:={η∈H:2h(ϵ)-2κ-h(ϵ)-2K<⟨η,e1f⟩H2≤L2}.

and the constants α,ζ,κ∈(0,1),K<0 will be chosen later. The remaining part of this section is devoted to the study of the right-hand side of (74) on each component separately.

Lemma 4.4

Let ϵ>0 small enough to have ϵh(ϵ)<1 and ζ∈(0,1/2). For all η∈B1f (75), t∈[0,T] we haveHx∗ϵ(Uδ,Z)(t,η)≥0,

up to terms that are exponentially negligible in the moderate deviations range.

We address the region B3f in the following lemma.

Lemma 4.5

Let κ∈(0,1), K=-ln3,ζ0>0, ζ∈[ζ0,1/2), ϵ>0 small enough to have ϵh(ϵ)<1 and h2(κ-1)(ϵ)≤9a1f2(ζ0-2ζ02). For all η∈B3f (75), t∈[0,T] we have either(i)Hx∗ϵ(Uδ,Z)(t,η)≥-Cϵh(ϵ),

or, if h(ϵ) is such that ϵh3(ϵ)→0, then for sufficiently small ϵ we have(ii)Hx∗ϵ(Uδ,Z)(t,η)≥0.

It remains to study the region B2f. It is the most problematic region as there is no guarantee that the weight ρϵ is exponentially negligible or of order one. The analysis is deferred to Appendix 1.

Lemma 4.6

Let α∈(0,1),κ<1-α,K=-ln3, ζ∈(ζ0,1/2), ϵ>0 small enough to have ϵh(ϵ)<1 and h2(κ+α-1)(ϵ)≤a1f2. For all η∈B2f (75), t∈[0,T] we have either(i)Hx∗ϵ(Uδ,Z)(t,η)≥-Cϵh(ϵ),

where C does not depend on ϵ or, if ϵh3(ϵ)⟶0, there exists ϵ sufficiently small such that(ii)Hx∗ϵ(Uδ,Z)(t,η)≥0.

Combining the three previous lemmas we arrive at the following regarding the third term in (72)

Lemma 4.7

There exists a constant C independent of T>0 such that for ϵ sufficiently small,(i)E∫0τ^∗ϵ,v∧τ∞ϵHx∗ϵ(Uδ,Z)(s,η^∗ϵ,v(s))ds≥-CTϵh(ϵ).

up to exponentially negligible terms in the moderate deviations range. Moreover, if h(ϵ) is such that limϵ→0ϵh3(ϵ)=0 then, for ϵ sufficiently small,(ii)E∫0τ^∗ϵ,v∧τ∞ϵHx∗ϵ(Uδ,Z)(s,η^∗ϵ,v(s))ds≥0.

up to exponentially negligible terms in the moderate deviations range.

Proof

(i) From Lemmas 4.4, 4.6(i), 4.5(i) we have∫0τ^∗ϵ,v∧τ∞ϵHx∗ϵ(Uδ,Z)(s,η^∗ϵ,v(s))ds≥-C(τ^∗ϵ,v∧τ∞ϵ)ϵh(ϵ),

with probability 1, up to exponentially negligible terms. Since τ^∗ϵ,v∧τ∞ϵ≤T with probability 1 and the constant is deterministic, the estimate follows by taking expectation.

(ii) The estimate follows from Lemmas 4.4, 4.6(ii), 4.5(ii). □

We conclude this section with the proof of Theorem 4.1.

Proof of Theorem 4.1

In view of Lemmas (4.1) and (4.7)(i), (72) yields-1h2(ϵ)logQϵ(uϵ)≥infv∈A[2Z(0,0)-2EZ(τ^∗ϵ,v,η^x∗ϵ,v(τ^∗ϵ,v))-CTϵh(ϵ)].

up to exponentially negligible terms. In view of (68), we have Z(t,η)=(1-ζ)Uδ(t,η). From Theorem 3.2 we have limϵ→0E[Z(τ^∗ϵ,v,η^x∗ϵ,v(τ^∗ϵ,v))]=0. Thus for ϵ sufficiently small we may write-1h2(ϵ)logQϵ(uϵ)≥infv∈A[Z(0,0)-CTϵh(ϵ)].

As for the first term, since Uδ is the exponential mollification of two functions, Lemma 4.1 of [23] gives thatUδ(0,0)≥F1(0)∧F2ϵ-δlog2=F2ϵ-δlog2.

Finally, the improved bound (70) follows by invoking Lemma (4.7)(ii). □

The case of a double-well potential

In this section we specialize our results to SRDEs in which the differential operator A=Δ (i.e. the second derivative operator in one spatial dimension) and the reaction term takes the form f=-Vf′, where Vf is a double-well potential as the one depicted below. This choice is possible in view of Hypotheses 2(a), 2(b) which allow arbitrary polynomial growth. Thus, we assume that Vf has two global minima and a local maximum which, for simplicity, is assumed to lie in the origin. Without loss of generality, we take f′(0)=-Vf′′(0)=1. Such SRDEs arise as scaling limits of particle systems with nearest-neighbor coupling that evolve in the inverted potential -Vf (see e.g. [4], Chapter 1) and provide one of the simplest examples of non-trivial dynamical behavior.

The deterministic reaction–diffusion equation posed on the interval (0,ℓ) has two stable equilibria x-∗,x+∗, corresponding to the global minima, and a saddle point x0∗, corresponding to the global maximum, that is identically equal to 0. The equilibria x±∗ only exist if ℓ>π for Dirichlet boundary conditions and for all ℓ>0 for Neumann and periodic conditions. Moreover, every time the interval length ℓ crosses the value kπ (for Neumann or Dirichlet b.c.) or 2kπ (for periodic b.c.), for some k∈N, two (resp. one) non-constant saddle points ±xk,ℓ∗ (resp. xp,k,ℓ∗) bifurcate from x0∗. The k-th non-constant saddle points feature k kink-antikink pairs in the periodic and Dirichlet cases and k kinks in the Neumann case. The interested reader is referred to [4], Section 2.1 and [27] for the bifurcation analysis of the problem with different boundary conditions.

For most of the sequel, we specialize the discussion to the potentialVf(x)=x44-x22,x∈R.

The corresponding bistable stochastic dynamics are governed by the (stochastic) Allen–Cahn equation76 ∂tXϵ=∂ξ2Xϵ+Xϵ-(Xϵ)3+ϵW˙.

The noiseless (ϵ=0) equation was proposed in [2] as a simple model of phase separation of two-component alloy systems. It is also known in the literature as real Ginzburg-Landau [40] (due to its connections with the physical superconductivity theory bearing the same name) or Chafe-Infante problem [15]. Transitions between the stable states x±∗ that correspond to the absolute minima ±1 are enabled by the stochastic forcing and have been studied as models of quantum tunneling phenomena [27] and thermally induced magnetization reversal of micromagnets [42]. For studies of transition times the interested reader is referred to [5, 6] and [42, 43] in the mathematical and physical literature respectively.

Stochastic Allen–Cahn with Neumann and periodic boundary conditions

The eigenvalues of the Neumann and periodic (negative) Laplacian on the interval (0,ℓ) are respectively given by77 anNe=(nπℓ)2,a±nper=(2nπℓ)2,n=0,1,2,⋯.

For ξ∈(0,ℓ),n=1,2,⋯, the corresponding eigenfunctions are78 e0Ne(ξ)=ℓ-1/2,enNe(ξ)=2ℓcos(nπξℓ),e0per(ξ)=ℓ-1/2,e±nper(ξ)=2ℓ[±sin(2nπξℓ)+cos(2nπξℓ)].

For both cases, the stable equilibria are the constant functions x±∗(ξ)=±1,ξ∈(0,ℓ). The reaction term is f(x)=x-x3 and the linearized operators Δ+DF(x±∗) acting on a function y are given by[Δ+DF(x±∗)]y(t)=y′′(t)+[(1-3x2)|x=x±∗]y(t)=y′′(t)-2y(t),t∈[0,T].

Both cases can be treated simultaneously after indexing the eigenpairs by the natural numbers. In particular, in the Neumann case, the eigenvalues {anf} from Hypothesis 3(b) are shifted eigenvalues of the (negative) Laplacian, i.e.anf=2+((n-1)πℓ)2,n=1,2,⋯

and the sequence of eigenvectors {enf} coincides with {en-1Ne}. In the periodic case we set a1f=2, a2nf=2+anper,a2n+1f=2+a-nper and e1f=e0per, e2nf=enper,e2n+1f=e-nper for n=1,2,⋯.

Turning to the spectral gap conditions, Theorems 3.1 and 3.2 hold for any value of ℓ provided that the change of measure uk0 acts on a finite-dimensional eigenspace of sufficiently high dimension. For example, consider the Neumann problem with ℓ=4π/3. For this value of ℓ, (76) has 3 saddle points (see e.g. [48], Chapter 5.3.4) and it is easy to check that Hypothesis 3(c) is violated. However, the weak spectral gap of Hypothesis 3c’ is satisfied for k0=3. Indeed, we have3a1f=6<2+9π2(4π/3)2=2+8116=7.0265=a4f.

Thus, the asymptotic results hold with the change of measureu3(t,η)=2(a1f⟨η,e1f⟩He1f+a2f⟨η,e2f⟩He2f+a3f⟨η,e3f⟩He3f).

As for the pre-asymptotic analysis of Sect. 4 and the numerical studies of the following section we work under the stronger spectral gap of Hypothesis 3(c). For the Neumann problem, this places the restriction3a1f=6<2+π2ℓ2=a2f⟺ℓ<π2

which can be weakened to ℓ<π/2 in view of Remark 7. For the periodic problem, Hypothesis 3(c) gives3a1f=6<2+4π2ℓ2=a2f⟺ℓ<π.

Finally, an example where the assumptions of Lemma 3.5 are satisfied is given by the Neumann problem with ℓ≥π/2. In this case it is straightforward to verify that a2f=2+π2/ℓ2≤4=2a1f.

Stochastic Allen–Cahn with Dirichlet boundary conditions

The eigenpairs of the Dirchlet Laplacian on the interval (0,ℓ) are explicitly given by79 anDir=(nπℓ)2,enDir(ξ)=2ℓsin(nπξℓ),n=1,2,⋯

However, exact spectral analysis and numerical simulation of the linearized operators is more involved than the periodic and Neumann cases. This is due to the fact that the stable equilibria x±∗ are non-constant functions with absolute value less than or equal to 1 that vanish at the endpoints 0,ℓ. They can be determined by solving the Sturm-Liouville problemx′′(ξ)=Vf′(x(ξ))=x3(ξ)-x(ξ),ξ∈(0,ℓ)x(0)=x(ℓ)=0.

Following [54] (see also [26]), we can parametrize x±∗ with respect to their minimum pointwise distance from the constant solutions ±1. The latter is in one-to-one correspondence with the bifurcation parameter ℓ (see (80) below).

First, note that the scaling y(ξ)=x(ℓξ) leads to the equivalent problemy′′(ξ)=ℓ2(y3(ξ)-y(ξ)),ξ∈(0,1)y(0)=y(1)=0.

For any a∈(0,1), the stable equilibria y±∗ of the latter are then given by ±y∗,y∗(ξ)≡ya∗(ξ)=asn(2K(a22-a2)ξ,a22-a2),ξ∈(0,1),

where for any m∈(0,1),K(m):=∫01dx(1-x2)(1-mx2),m∈(0,1)

is the complete elliptic integral of the first kind and sn(·,m) is the Jacobi elliptic sine function defined byx=∫0sn(x,m)dy(1-y2)(1-my2),x∈[0,K(m)].

The function sn(·,m) can be periodically extended to all of R so that K(m) is its quarter-period. We remark that there are several different parameterizations of K in the literature (e.g. in [54] K(ξ) corresponds to K(ξ) in our notation). The definition above was chosen in agreement with [1] and the corresponding built-in Matlab function.

The parameter a is the maximum value of ya∗ i.e.ya∗(inf{ξ∈(0,1):y′(ξ)=0})=a.

In order to convert to a parameterization in terms of ℓ, we first define a scaled quarter-period map M:(0,1)→R withM(a):=12∫0adxVf(a)-Vf(x)=2∫01dx1-x22-a2(1+x2)=22-a2K(a22-a2).

The correspondence of the interval length ℓ and a is then given by80 ℓ=2M(a).

As seen in the Fig. 1, M is continuous, strictly increasing and lima→1M(a)=∞. Thus M is continuously invertible. Furthermore it is straightforward to verify that lima→0M(a)=π/2. Putting the previous facts together we deduce that81 x+∗(ξ)=-x-∗(ξ)=y+∗(ξ/ℓ)=asn(2K(a22-a2)ξℓ,a22-a2)=asn(2K(a22-a2)ξ2M(a),a22-a2)=asn(ξ1-a22,a22-a2),ξ∈(0,ℓ),a=M-1(ℓ/2).

Fig. 1 The map M

Turning to the spectral properties of the linearized operators Δ+DF(x±∗), they have a countable sequence of eigenvalues-eigenvectors {(anf,enf)}n∈N, hence they satisfy Hypothesis 3(b). The first two pairs have been computed explicitly in [54] and are given by82 a1f=32a2=32M-1(ℓ/2),e1f(ξ)=e1,af(ξ)=sn(ξ1-a22,a22-a2)dn(ξ1-a22,a22-a2)

anda2f=32(2-a2)=32(2-M-1(ℓ/2)),e2f(ξ)=e2,af(ξ)=sn(ξ1-a22,a22-a2)cn(ξ1-a22,a22-a2)

where dn, cn denote the Jacobi delta amplitude and elliptic cosine functionsdn(x,m):=1-m2sn2(x,m),cn2(x,m):=1-sn2(x,m),cn(0,m)=1.

The spectral gap of Hypothesis 3(c) is then satisfied if3a1fa2f=3a22-a2<1⟺a<22⟺ℓ<2M(2/2)

where we used the monotonicity of M and 2M(2/2)≈4.0043. Plots of the equilibria ya∗ and eigenfunctions e1,af for a=0.65,0.95 are given in Figs. 2 and 3.Fig. 2 Instances of the Dirichlet stable equilibrium ya∗

Fig. 3 Instances of the Dirichlet eigenfunctions e1,af

A higher-order Ginzburg–Landau SRDE

We conclude this section with an example of an SRDE with a higher-order polynomial nonlinearity. This time we consider a potential given byVf(x)=4+μ12-12x2-μ4x4+μ+16x6,x∈R.

If μ>-1 then Vf is a double-well potential with steeper walls than the fourth-order case. Such potentials have been considered in the physical literature as higher order quantum mechanical models, see e.g. [39]. The nonlinear reaction term is given by f(x)=-Vf′(x)=x+μx3-(μ+1)x5 and f′(x)=1+3μx2-5(μ+1)x4. For μ∈(-1,0], Hypothesis 2(a) is satisfied with f1(x)=x and f2(x)=μx3-(μ+1)x5. The corresponding SRDE is given by83 ∂tXϵ=∂ξ2Xϵ+Xϵ+μ(Xϵ)3-(μ+1)(Xϵ)5+ϵW˙.

The noiseless dynamics with Neumann or periodic boundary conditions are bistable for any ℓ>0 with stable equilibria x±∗=±1 and a saddle point x0∗=0. The linearized operatorΔ+DF(±1)=Δ+(1+3μx2-5(μ+1)x4)|x=±1=Δ-2μ-4

has the same eigenfunctions as the Laplacian and eigenvalues shifted by -2μ-4.

As in the Allen–Cahn case, Theorems 3.1 and 3.2 hold for any value of ℓ provided that the change of measure uk0 acts on a finite-dimensional eigenspace of sufficiently high dimension. The pre-asymptotic analysis of Sect. 4 holds under the spectral gap of Hypothesis 3(c). In the Neumann case the spectral gap holds if3a1f=6μ+12<2μ+4+π2ℓ2=a2f⟺ℓ<π2μ+2

and in the periodic case3a1f=6μ+12<2μ+4+4π2ℓ2=a2f⟺ℓ<πμ+2.

Numerical simulations

In this section we demonstrate the theoretical results of this paper by a series of simulation studies for (8). As explained in Sect. 3, we start the process Xxϵ at a stable equilibrium x=x∗ and develop a scheme that computes exit probabilities of the formP(ϵ)=P(ϵ,T)=P[τx∗ϵ≤T]

for ϵ≪1,T>0, where τx∗ϵ=inf{t>0:Xx∗ϵ∉D} and D=B˚H(x∗,Lϵh(ϵ)). For the simulations that follow we fix L=1 and set R=R(ϵ)=ϵh(ϵ). In view of Remark 10 we haveP(ϵ)=P[supt∈[0,T]‖ηx∗ϵ(t)‖H≥1]

and the process ηx∗ϵ (3) converges in distribution to 0 as ϵ→0. Hence, for ϵ small, we are dealing with rare events. We will apply the scheme of Sect. 4 to the examples of Sect. 5 and compare its performance to the standard Monte Carlo, which corresponds to no change of measure at all. It is clear that in order to simulate the mild solutions in (9), (42) we need to discretize the equation in time and space. In the simulations below we used the exponential Euler scheme finite-dimensional Galerkin projection as it is described in [37]. In particular, with η^ϵ=(X^ϵ-x∗)/ϵh(ϵ), uϵ as in (25) and X^ϵ solving84 dX^ϵ(t)=[AX^ϵ(t)+F(X^ϵ(t))]dt+ϵh(ϵ)uϵ(η^ϵ(t))dt+ϵdW(t)X^ϵ(0)=x∗

we simulate the mild solution X^ϵ on the sampling window t∈[0,T] until it hits ∂D. Its N-th Galerkin projection is given in mild formulation byXNϵ(t)=eANtPNx∗+∫0teAN(t-s)[PNF(X^Nϵ(s))+ϵh(ϵ)PNuϵ(η^Nϵ(s))]ds+ϵWAN(t),

where PN denotes a projection to the N-dimensional subspace of H spanned by the eigenvectors e1,⋯,eN of A (not to be confused with the linearization eigenvectors enf of Hypothesis 3(b)). Turning to the time discretization, we consider a time-step h=T/Δt for some Δt∈N, discretization times tk=kh, k=0,⋯,Δt and set Θ0N:=PNx∗. The exponential Euler scheme is then given byΘk+1N=eANhΘkN+AN-1[eANh-I]-1[PNF(ΘkN)+ϵh(ϵ)PNuϵ(Θ~kN)]+ϵ∫tkk+1eAN(tk+1-s)PNdW(s),

k=0,⋯,Δt-1 where Θ~kN=(ΘkN-Θ0N)/ϵh(ϵ). Letting Θk,jN=⟨ΘkN,ej⟩H,fk,jN=⟨PNF(ΘkN),ej⟩H,uk,jN=⟨PNuϵ(Θ~kN),ej⟩H the numerical scheme for the approximation of (84) is then given by85 Θk+1,jN=e-ajhΘk,jN+1-e-ajhaj(fk,jN+ϵh(ϵ)uk,jN)+ϵ1-e-2ajh2ajwk,j

where for k=0,⋯Δt-1,j=1,⋯,N ξk,j are independent standard normal random variables.

For Neumann and periodic boundary conditions, the pairs (aj,ej) are given by (77), (78). Since the changes of measure uϵ act only in the direction of e1f and the latter coincides with e0 (i.e. a constant function) we have that uk,jN=0 when j≠0. However the eigenvalue a0 is in both cases equal to 0. Hence the exponential Euler scheme is not well-defined for j=0. For this reason we simulate Θk+1,0N via an explicit Euler scheme i.e.Θk+1,0N=Θk,0N+h[fk,0N+ϵh(ϵ)uk,0N]+ϵhwk,0,k=0,⋯,Δt-1,

where Θ0,0N:=⟨x∗,e0⟩H, fk,0N=⟨PNF(ΘkN),e0⟩H,uk,0N=⟨PNuϵ(Θ~kN),e0⟩H and wk,0 are once again independent standard normal random variables. The computation of the coefficients {fk,jN} in the Neumann (respectively periodic) case can be efficiently performed by applying a forward-backward odd (resp. periodic or Hartley-type) Fast Fourier Transform (FFT) in an iterative fashion. For more details on the discrete Fourier transform and the FFT algorithm the reader is referred to [9] (Chapters 6 and 8 respectively).

Turning to the stochastic Allen–Cahn with Dirichlet boundary conditions the simulations require an additional step. As discussed in Sect. 5, the stable equilibrium x∗ (81) is no longer a constant function and the changes of measure uϵ push towards e1f (82) which no longer coincides with a single eigenvector ekDir (79). Thus, one needs to express x∗ and e1f in terms of the eigenbasis {ekDir}k∈N of the Laplacian and then perform the exponential Euler scheme (85). If the changes of measure acted on a higher dimensional eigenspace, this step essentially reduces to a change of basis which can be computed with numerical linear-algebraic methods. Regarding the coefficients {fk,jN}, these can be computed by applying a forward-backward even Fast Fourier Transform iteratively.

Remark 16

In the examples of Sect. 5 the stable equilibrium x∗ and the spectra of the linearized operators could be found explicitly. We remark here that our scheme does not depend on explicit formulas for eigenvalues and eigenvectors as long as those can be approximated numerically and the approximated eigenvalues satisfy Hypothesis 3(c).

All the simulations below were done using a parallel MPI C code with M=5×104 Monte Carlo trajectories. The FFTs were performed with the aid of the C library FFTW. As it is standard in the related literature (see e.g. [3], Chapter VI,1), the measure of performance is relative error per sample, defined asrelative error per sample=Mst.deviation(P^ϵ)expectation[P^(ϵ)].

The smaller the relative error per sample, the more efficient the algorithm and the more accurate the estimator. However, in practice both the standard deviation and the expected value of an estimator are typically unknown, which implies that empirical relative error is often used for measurement. This means that the expected value of the estimator will be replaced by the empirical sample mean and the standard deviation of the estimator will be replaced by the empirical sample standard error. A dash line in the simulation tables indicates that no trajectory exited D before time T. Before presenting the simulation tables, let us make a few comments on the parameter values and the end conclusions of the numerical studies.

1) (Simulations for the Neumann stochastic Allen–Cahn) We estimate exit probabilities P(ϵ) for the solution Xϵ of (76) driven by additive space-time white noise on the interval (0,ℓ) and Neumann boundary conditions. For the simulations we set ℓ=1,x∗=x+∗=1, h(ϵ)=ϵ-0.1 and Galerkin projection level N=50. The numerical results can be found in Tables 1–4.Table 1 Estimated probability values P(ϵ,T) for the stochastic Allen–Cahn equation with Neumann boundary conditions using the developed importance sampling scheme with κ=0.9 and mollification parameter δ:=2/h2(ϵ)

ϵ	R=ϵh(ϵ)	T=1	T=2	T=3	T=4	T=6	T=8	
0.01	0.158489	1.93e-02	5.43e-02	8.73e-02	1.20e-01	1.80e-01	2.35e-01	
0.004	0.109856	6.65e-03	2.01e-02	3.37e-02	4.66e-02	7.19e-02	9.75e-02	
0.002	0.083255	2.60e-03	8.35e-03	1.41e-02	1.98e-02	3.13e-02	4.24e-02	
0.0008	0.057708	6.22e-04	2.09e-03	3.64e-03	5.13e-03	8.24e-03	1.12e-02	
0.0004	0.043734	1.72e-04	6.21e-04	1.08e-03	1.53e-03	2.46e-03	3.36e-03	
0.0001	0.025119	6.95e-06	2.92e-05	5.28e-05	7.62e-05	1.24e-04	1.70e-04	
0.00006	0.020477	1.77e-06	7.63e-06	1.38e-05	2.00e-05	3.26e-05	4.45e-05	
0.000008	0.009146	1.31e-09	7.45e-09	1.39e-08	2.05e-08	3.39e-08	4.72e-08	
0.000004	0.006931	4.96e-11	3.30e-10	6.21e-10	9.31e-10	1.52e-09	2.13e-09	

Table 2 Estimated relative errors per sample for the stochastic Allen–Cahn equation with Neumann boundary conditions using the developed importance sampling scheme with κ=0.9 and mollification parameter δ:=2/h2(ϵ)

ϵ	T=1	T=2	T=3	T=4	T=6	T=8	
0.01	2.1	1.2	1.0	0.9	1.0	1.3	
0.004	2.4	1.3	1.0	1.0	1.0	1.3	
0.002	2.5	1.4	1.1	1.0	1.0	1.2	
0.0008	2.7	1.5	1.1	1.0	1.0	1.2	
0.0004	2.9	1.5	1.1	1.0	0.9	1.1	
0.0001	3.3	1.6	1.2	1.0	0.9	1.1	
0.00006	3.4	1.6	1.2	1.0	0.9	1.1	
0.000008	4.2	1.7	1.2	1.0	0.9	1.0	
0.000004	4.6	1.7	1.3	1.0	0.9	1.0	

Table 3 Estimated probability values P(ϵ,T) for the stochastic Allen–Cahn equation with Neumann boundary conditions. The values reported are based on standard Monte Carlo simulation without employing some change of measure

ϵ	R=ϵh(ϵ)	T=1	T=2	T=3	T=4	T=6	T=8	
0.01	0.158489	2.07e-02	5.40e-02	8.66e-02	1.20e-01	1.80e-01	2.40e-01	
0.004	0.109856	6.80e-03	2.02e-02	3.36e-02	4.65e-02	7.15e-02	9.82e-02	
0.002	0.083255	2.60e-03	8.58e-03	1.41e-02	1.90e-02	3.05e-02	4.39e-02	
0.0008	0.057708	3.80e-04	2.18e-03	3.22e-03	4.82e-03	7.32e-03	1.13e-02	
0.0004	0.043734	1.80e-04	5.60e-04	9.40e-04	1.60e-03	2.50e-04	3.10e-03	
0.0001	0.025119	–	4.00e-05	8.00e-05	1.40e-04	8.00e-05	1.40e-04	
0.00006	0.020477	2.00e-05	2.00e-05	–	2.00e-05	6.00e-05	4.00e-05	
0.000008	0.009146	–	–	–	–	–	–	
0.000004	0.006931	–	–	–	–	–	–	

Table 4 Estimated relative errors per sample for the stochastic Allen–Cahn equation with Neumann boundary conditions. The values reported are based on standard Monte Carlo simulation without employing some change of measure. A probability of 2×10-5 means that only one out of the 5×104 trajectories exited the domain. The relative error in that case is 223.6

ϵ	T=1	T=2	T=3	T=4	T=6	T=8	
0.01	6.9	4.2	3.2	2.7	2.1	1.8	
0.004	12.1	7.0	5.4	4.5	3.6	3.0	
0.002	19.6	10.7	8.4	7.2	5.6	4.7	
0.0008	51.3	21.4	17.6	14.4	11.6	9.3	
0.0004	74.5	42.2	32.6	25.0	20.0	17.9	
0.0001	–	158.1	111.8	84.5	111.8	84.5	
0.00006	223.6	223.6	–	223.6	129.1	158.1	
0.000008	–	–	–	–	–	–	
0.000004	–	–	–	–	–	–	

Table 5 Estimated probabilities for the stochastic Allen–Cahn equation with periodic boundary conditions using the developed importance sampling scheme with κ=0.9 and mollification parameter δ:=2/h2(ϵ)

ϵ	R=ϵh(ϵ)	T=1	T=2	T=3	T=4	T=6	T=8	
0.01	0.158489	1.69e-02	4.82e-02	7.73e-02	1.07e-01	1.62e-01	2.14e-01	
0.004	0.109856	5.91e-03	1.82e-02	3.02e-02	4.23e-02	6.60e-02	8.87e-02	
0.002	0.083255	2.39e-03	7.68e-03	1.29e-02	1.82e-02	2.88e-02	3.94e-02	
0.0008	0.057708	5.62e-04	1.97e-03	3.44e-03	4.84e-03	7.71e-03	1.04e-02	
0.0004	0.043734	1.61e-04	5.92e-04	1.03e-03	1.47e-03	2.34e-03	3.24e-03	
0.0001	0.025119	6.90e-06	2.92e-05	5.27e-05	7.57e-05	1.22e-04	1.70e-04	
0.00006	0.020477	1.73e-06	7.68e-06	1.38e-05	2.00e-05	3.25e-05	4.52e-05	
0.000008	0.009146	1.34e-09	7.90e-09	1.52e-08	2.23e-08	3.66e-08	5.13e-08	
0.000004	0.006931	5.79e-11	3.65e-10	7.00e-10	1.05e-09	1.70e-09	2.37e-09	

2) (Simulations for the periodic stochastic Allen–Cahn) We estimate P(ϵ) for the solution Xϵ of (76) driven by additive space-time white noise on the interval (0,ℓ) and periodic boundary conditions. For the simulations we set x+∗=1,ℓ=1,R=R(ϵ):=ϵh(ϵ)∈(0,1),h(ϵ)=ϵ-0.1,ηϵ=(Xϵ-x+∗)/R(ϵ), Galerkin projection level N=50. The numerical results can be found in Tables 5–8.

3) (Simulations for the Dirichlet stochastic Allen–Cahn) We estimate P(ϵ) for the solution Xϵ of (76) driven by additive space-time white noise on the interval (0,ℓ) and Dirichlet boundary conditions. For the simulations we set ℓ=3.81828, x∗=x+∗=asn(ξ1-a22,a22-a2) with a=M-1(ℓ/2)=0.65, h(ϵ)=ϵ-0.1 and Galerkin projection level N=50. Note that ‖x+∗‖L2≈0.33. The numerical results can be found in Tables 9–12.

4) (Simulations for the quintic SRDE (83)) We estimate P(ϵ) for the solution Xϵ of (83) driven by additive space-time white noise on the interval (0,ℓ) and Neumann boundary conditions. For the simulations we set μ=-0.5,x+∗=1,ℓ=1,R=R(ϵ):=ϵh(ϵ)∈(0,1),h(ϵ)=ϵ-0.1,ηϵ=(Xϵ-x+∗)/R(ϵ), Galerkin projection level N=50. The numerical results can be found in Tables 13–16.Table 6 Estimated relative errors per sample for the stochastic Allen–Cahn equation with periodic boundary conditions using the developed importance sampling scheme with κ=0.9 and mollification parameter δ:=2/h2(ϵ)

ϵ	T=1	T=2	T=3	T=4	T=6	T=8	
0.01	2.1	1.2	1.0	0.9	1.0	1.3	
0.004	2.4	1.3	1.0	0.9	1.0	1.2	
0.002	2.5	1.4	1.1	0.9	1.0	1.2	
0.0008	2.7	1.4	1.1	0.9	1.0	1.1	
0.0004	2.9	1.4	1.1	1.0	0.9	1.1	
0.0001	3.2	1.5	1.1	1.0	0.9	1.1	
0.00006	3.4	1.6	1.1	1.0	0.9	1.0	
0.000008	4.2	1.7	1.2	1.0	0.9	1.0	
0.000004	4.4	1.7	1.2	1.0	0.9	1.0	

Table 7 Estimated probabilities for the stochastic Allen–Cahn equation with periodic boundary conditions. The values reported are based on standard Monte Carlo simulation without employing some change of measure

ϵ	R=ϵh(ϵ)	T=1	T=2	T=3	T=4	T=6	T=8	
0.01	0.158489	1.73e-02	4.96e-02	7.72e-02	1.07e-01	1.62e-01	2.13e-01	
0.004	0.109856	6.44e-03	1.80e-02	3.00e-02	4.27e-02	6.45e-02	8.98e-02	
0.002	0.083255	2.54e-03	8.24e-03	1.28e-02	1.81e-02	2.77e-02	3.89e-02	
0.0008	0.057708	6.20e-04	1.34e-03	3.01e-03	4.52e-03	7.94e-03	1.02e-02	
0.0004	0.043734	1.40e-04	5.20e-04	1.04e-03	1.42e-03	2.24e-03	3.30e-03	
0.0001	0.025119	2.00e-05	–	1.00e-04	1.00e-04	1.20e-04	1.40e-04	
0.00006	0.020477	–	2.00e-05	–	6.00e-05	2.00e-05	–	
0.000008	0.009146	–	–	–	–	–	–	
0.000004	0.006931	–	–	–	–	–	–	

Table 8 Estimated relative error per sample for the stochastic Allen–Cahn equation with periodic boundary conditions. The values reported are based on standard Monte Carlo simulation without employing some change of measure

ϵ	T=1	T=2	T=3	T=4	T=6	T=8	
0.01	7.5	4.4	3.5	2.9	2.3	1.9	
0.004	12.4	7.4	5.7	4.7	3.8	3.2	
0.002	19.8	11.0	8.8	7.4	5.9	5.0	
0.0008	40.1	27.3	17.9	14.8	11.2	9.9	
0.0004	84.5	43.8	30.1	26.5	21.1	17.4	
0.0001	223.6	–	100.0	100.0	91.3	84.5	
0.00006	–	223.6	–	129.1	223.6	–	
0.000008	–	–	–	–	–	–	
0.000004	–	–	–	–	–	–	

5) Standard Monte Carlo (sMC) estimation, i.e. with no change of measure does not perform well for small values of ϵ, as indicated in Tables 4,12,8. A dash line indicates that there was no successful trajectory in the simulations and thus no estimate could be provided. The relative errors per sample are getting increasingly large making most of the reported probability values of Tables 3,11,7 to be of no value.Table 9 Estimated probability values P(ϵ,T) for the stochastic Allen–Cahn equation with Dirichlet boundary conditions using the developed importance sampling scheme with κ=0.9 and mollification parameter δ:=2/h2(ϵ)

ϵ	R=ϵh(ϵ)	T=1	T=2	T=3	T=4	T=6	T=8	
0.00008	0.022974	3.96e-03	2.47e-02	5.06e-02	7.69e-01	1.26e-01	1.73e-01	
0.00001	0.010000	1.55e-04	2.19e-03	5.55e-03	9.23e-03	1.69e-02	2.43e-02	
0.000004	0.006931	2.40e-05	5.18e-04	1.47e-03	2.61e-03	5.00e-03	7.37e-03	
0.000001	0.003981	5.56e-07	3.44e-05	1.23e-04	2.32e-04	4.69e-04	7.13e-04	
0.0000008	0.003641	4.82e-07	2.07e-05	7.71e-05	1.47e-04	3.00e-04	4.52e-04	
0.0000004	0.002759	4.35e-08	3.68e-06	1.55e-05	3.06e-05	6.54e-05	1.00e-04	
0.0000002	0.002091	2.56e-09	4.82e-07	2.38e-06	5.12e-06	1.12e-05	1.75e-05	
0.0000001	0.001585	1.24e-10	5.23e-08	2.84e-07	6.34e-07	1.48e-06	2.33e-06	
0.00000008	0.001450	4.32e-11	2.27e-08	1.33e-07	3.08e-07	7.20e-07	1.14e-06	

Table 10 Estimated relative errors per sample for the stochastic Allen–Cahn equation with Dirichlet boundary conditions using the developed importance sampling scheme with κ=0.9 and mollification parameter δ:=2/h2(ϵ)

ϵ	T=1	T=2	T=3	T=4	T=6	T=8	
0.00008	5.3	2.1	1.4	1.1	0.9	0.8	
0.00001	10.1	2.8	1.7	1.3	1.0	0.9	
0.000004	14.5	3.3	1.9	1.4	1.0	0.9	
0.000001	29.1	4.0	2.1	1.5	1.0	0.9	
0.0000008	26.1	4.1	2.2	1.5	1.1	0.9	
0.0000004	36.8	4.6	2.3	1.6	1.1	0.9	
0.0000002	65.9	5.2	2.4	1.6	1.1	0.9	
0.0000001	100.4	6.0	2.6	1.7	1.1	0.9	
0.00000008	112.4	6.2	2.7	1.8	1.1	0.9	

Table 11 Estimated probability values P(ϵ,T) for the stochastic Allen–Cahn equation with Dirichlet boundary conditions. The values reported are based on standard Monte Carlo simulation without employing some change of measure

ϵ	R=ϵh(ϵ)	T=1	T=2	T=3	T=4	T=6	T=8	
0.00008	0.022974	3.64e-03	2.48e-02	5.00e-02	7.68e-02	1.27e-01	1.73e-01	
0.00001	0.010000	1.00e-04	2.22e-03	6.12e-03	9.16e-03	1.63e-02	2.33e-02	
0.000004	0.006931	2.00e-05	4.40e-04	1.84e-03	2.76e-03	4.80e-03	7.24e-03	
0.000001	0.003981	2.00e-05	4.00e-05	1.20e-04	2.20e-04	4.40e-04	6.40e-04	
0.0000008	0.003641	–	4.00e-05	4.00e-05	1.40e-04	3.20e-04	5.80e-04	
0.0000004	0.002759	–	–	2.00e-05	–	4.00e-05	8.00e-05	
0.0000002	0.002091	–	–	–	2.00e-05	–	2.00e-05	
0.0000001	0.001585	–	–	–	–	–	–	
0.00000008	0.001450	–	–	–	–	–	–	

Table 12 Estimated relative errors per sample for the stochastic Allen–Cahn equation with Dirichlet boundary conditions. The values reported are based on standard Monte Carlo simulation without employing some change of measure

ϵ	T=1	T=2	T=3	T=4	T=6	T=8	
0.00008	16.5	6.3	4.4	3.5	2.6	2.2	
0.00001	100.0	21.2	12.7	10.4	7.8	6.5	
0.000004	223.6	47.7	23.3	19.0	14.4	11.7	
0.000001	223.6	158.1	91.3	67.4	47.7	39.5	
0.0000008	–	158.1	158.1	84.5	55.9	41.5	
0.0000004	–	–	223.6	–	158.1	111.8	
0.0000002	–	–	–	223.6	–	223.6	
0.0000001	–	–	–	–	–	–	
0.00000008	–	–	–	–	–	–	

6) The importance sampling scheme for the Allen–Cahn equation outperforms sMC and performs well for all boundary conditions and probabilities ranging from 10-1 to 10-11 (see Tables 1, 9, 5). In particular, the relative errors for the former are way lower than those of sMC. As expected, the estimated probabilities resulting from sMC and importance sampling scheme agree when the relative errors are below 10.0. The relative errors per sample for the importance sampling scheme lie mostly below 1.2. The relative-error trends indicate that the accuracy improves as the sampling time grows from T=1 to T=8. The relative errors per sample as reported in Tables 2,10, 6 support the theoretical findings in that the scheme performs optimally as the theory predicts.

7) The performance of the importance sampling scheme for the quintic SRDE (83) experiences a slight degradation after T=3 (see Table 14). Nevertheless, it remains superior to that of the sMC (compare to Table 16) while the relative errors remain mostly below 2.5 and decrease with ϵ.

8) Table 17 provides a comparison between the importance sampling relative errors for the Neumann Allen–Cahn with h(ϵ)=ϵ-0.2, h(ϵ)=ϵ-0.1. The sampling time is fixed to T=2. We observe that the scaling h(ϵ)=ϵ-0.1 leads to significantly lower relative errors than h(ϵ)=ϵ-0.2. This behavior is correctly predicted by (70) since the first satisfies ϵh3(ϵ)→0 while the second does not. We remark that, despite the higher relative errors, the importance sampling scheme with h(ϵ)=ϵ-0.2 still outperforms sMC. Complete simulation tables for h(ϵ)=ϵ-0.2 are available upon request.Table 13 Estimated probabilities for the quintic SRDE (83) with Neumann boundary conditions and μ=-0.5 using the developed importance sampling scheme with κ=0.999 and mollification parameter δ:=2/h2(ϵ). The other parameters are h(ϵ)=ϵ-0.1,ℓ=1,x∗=1

ϵ	R=ϵh(ϵ)	T=1	T=2	T=3	T=4	T=6	T=8	
0.008	0.144956	3.87e-03	1.08e-02	1.77e-03	2.24e-02	3.84e-02	5.27e-02	
0.003	0.097915	6.22e-04	1.81e-03	2.98e-03	4.18e-03	6.43e-03	9.04e-03	
0.002	0.083255	2.62e-04	7.66e-04	1.30e-03	1.79e-03	2.86e-03	3.85e-03	
0.0008	0.057708	2.96e-05	8.99e-05	1.47e-04	2.09e-04	3.30e-04	4.47e-04	
0.0006	0.051435	1.38e-05	4.16e-05	6.97e-05	9.90e-05	1.55e-04	2.09e-04	
0.0002	0.033145	4.68e-07	1.50e-06	2.56e-06	3.57e-06	5.71e-06	7.62e-06	
0.00006	0.020477	4.71e-09	1.60e-08	2.73e-08	3.89e-08	6.09e-08	8.41e-08	
0.00002	0.013195	2.42e-11	8.84e-11	1.52e-10	2.16e-10	3.41e-10	4.82e-10	
0.000008	0.009146	1.16e-13	4.47e-13	7.72e-13	1.10e-12	1.80e-12	2.45e-12	

Table 14 Estimated relative errors per sample for the quintic SRDE (83) with Neumann boundary conditions and μ=-0.5 using the developed importance sampling scheme with κ=0.999 and mollification parameter δ:=2/h2(ϵ). The rest of the parameters are h(ϵ)=ϵ-0.1,ℓ=1,x∗=1

ϵ	T=1	T=2	T=3	T=4	T=6	T=8	
0.008	2.1	1.6	1.6	2.1	2.5	4.4	
0.003	2.3	1.7	1.6	1.7	2.4	5.1	
0.002	2.3	1.6	1.5	1.6	2.3	3.4	
0.0008	2.4	1.6	1.4	1.5	2.1	3.0	
0.0006	2.4	1.6	1.4	1.5	2.3	2.9	
0.0002	2.4	1.5	1.3	1.4	1.9	2.7	
0.00006	2.4	1.4	1.2	1.3	1.8	2.5	
0.00002	2.4	1.3	1.2	1.2	1.7	2.7	
0.000008	2.4	1.3	1.1	1.2	1.6	2.5	

Table 15 Estimated probability values for the quintic SRDE (83) with Neumann boundary conditions and μ=-0.5. The values reported are based on standard Monte Carlo simulation without employing some change of measure

ϵ	R=ϵh(ϵ)	T=1	T=2	T=3	T=4	T=6	T=8	
0.008	0.144956	3.60e-03	1.13e-02	1.83e-02	2.44e-02	3.83e-02	5.11e-02	
0.003	0.097915	5.80e-04	1.72e-03	2.38e-03	3.76e-03	6.76e-03	8.08e-03	
0.002	0.083255	2.40e-04	7.60e-04	1.26e-03	1.90e-03	3.02e-03	3.56e-03	
0.0008	0.057708	8.00e-05	1.20e-04	1.60e-04	2.80e-04	1.80e-04	5.60e-04	
0.0006	0.051435	2.00e-05	8.00e-05	4.00e-05	1.60e-04	1.40e-04	1.60e-04	
0.0002	0.033145	–	–	–	2.00e-05	–	–	
0.00006	0.020477	–	–	–	–	–	–	
0.00002	0.013195	–	–	–	–	–	–	
0.000008	0.009146	–	–	–	–	–	–	

Table 16 Estimated relative errors per sample for the quintic SRDE (83) with Neumann boundary conditions and μ=-0.5. The values reported are based on standard Monte Carlo simulation without employing some change of measure

ϵ	T=1	T=2	T=3	T=4	T=6	T=8	
0.008	16.6	9.4	7.3	6.3	5.0	4.3	
0.003	41.5	24.1	20.5	16.3	12.1	11.1	
0.002	64.5	36.3	28.2	22.9	18.2	16.7	
0.0008	111.8	91.3	79.1	59.8	74.5	42.2	
0.0006	223.6	111.8	158.1	79.1	84.5	79.1	
0.0002	–	–	–	223.6	–	–	
0.00006	–	–	–	–	–	–	
0.00002	–	–	–	–	–	–	
0.000008	–	–	–	–	–	–	

Table 17 Comparison of relative errors and probabilities produced by the importance sampling scheme for the Neumann Allen–Cahn equation with different moderate deviation scalings h(ϵ). The rest of the parameters are x∗=1,ℓ=1,κ=0.9,T=2

ϵ/T=2,h(ϵ)=ϵ-0.1	Prob.	Rel. error/
sample	ϵ/T=2,h(ϵ)=ϵ-0.2	Prob.	Rel. error/
sample	
0.01	5.43e-02	1.2	0.08	8.11e-02	1.7	
0.004	2.01e-02	1.3	0.05	3.35e-02	2.1	
0.002	8.35e-03	1.4	0.03	9.85e-03	2.5	
0.0008	2.09e-03	1.5	0.01	2.01e-04	4.3	
0.0004	6.21e-04	1.5	0.008	6.84e-05	3.8	
0.0001	2.92e-05	1.6	0.006	1.44e-05	5.1	
0.00006	7.63e-06	1.6	0.004	1.23e-06	9.2	
0.000008	7.45e-09	1.7	0.002	4.64e-09	5.0	
0.000004	3.30e-10	1.7	0.001	2.97e-12	5.4	

Table 18 Comparison of the relative errors produced by the importance sampling scheme for the Neumann Allen–Cahn equation with different Galerkin projection levels N. The rest of the parameters are x∗=1,ℓ=1,κ=0.9,h(ϵ)=ϵ-0.1,T=3

ϵ/T=3	N=50	N=100	N=150	
0.01	1.0	1.0	1.0	
0.004	1.0	1.0	1.0	
0.002	1.1	1.1	1.1	
0.0008	1.1	1.1	1.1	
0.0004	1.1	1.1	1.1	
0.0001	1.2	1.1	1.2	
0.00006	1.2	1.2	1.2	
0.000008	1.2	1.2	1.2	
0.000004	1.3	1.2	1.2	

9) In Table 18, we work with the Neumann Allen–Cahn and compare relative errors for different levels N of the Galerkin approximation with N=50,100,150. The sampling time is fixed to T=3. We notice that the relative errors are practically of the same order. This indicates that the first mode really dominates the rare event. Another observation we made is that the total simulation time increased significantly as we increased N. These considerations led us to conclude that N=50 is an efficient and sufficiently good lower dimensional approximation to the corresponding SPDE.

Numerical results for stochastic Allen–Cahn with Neumann boundary conditions

In this section, we provide numerical simulation results validating our theory for the stochastic Allen–Cahn equation with Neumann boundary conditions studied in Subsection 5.1.

Numerical results for stochastic Allen–Cahn with periodic boundary conditions

In this section, we provide numerical simulation results validating our theory for the stochastic Allen–Cahn equation with periodic boundary conditions studied in Subsection 5.1.

Numerical results for stochastic Allen–Cahn with Dirichlet boundary conditions

In this section, we provide numerical simulation results validating our theory for the stochastic Allen–Cahn equation with Dirichlet boundary conditions studied in Subsection 5.2.

Numerical results for the quintic SRDE (83) with Neumann boundary conditions

In this section, we provide numerical simulation results validating our theory for the for the quintic SRDE (83) with Neumann boundary conditions studied in Subsection 5.3.

Numerical comparisons of relative errors and probabilities for different parameter values

In this section, we provide numerical simulation results validating our theory for the stochastic Allen–Cahn equation with Neumann boundary conditions studied in Subsection 5.1. In particular, we now explore the effect of different moderate deviation scalings h(ϵ) and of different Galerkin projection levels N.

Conclusions and future work

In this paper we studied the problem of rare event simulation for small-noise SRDEs via moderate deviation-based importance sampling. Taking advantage of the linearized limiting dynamics of the process η^ϵ,v (42), we constructed changes of measure that behave optimally in the limit as ϵ→0 under the fairly general spectral gap condition of Hypothesis 3c’. Working under the more restrictive Hypothesis 3(c) we designed an importance sampling scheme with changes of measure that act on a one-dimensional eigenspace of the operator A+DF(x∗). We were then able to show that this scheme performs well pre-asymptotically and supplemented the theoretical results with numerical simulations for gradient-type SRDEs corresponding to a double-well potential. Such systems have wide applicability and provided good examples to illustrate our theory. Nevertheless, there are other types of nonlinearities which satisfy our assumptions, e.g. f=-Vf′, where the potential Vf(x)=sinx has more than two global minima.

The design and pre-asymptotic analysis of a scheme under the weaker spectral gap of Hypothesis 3c’ provides an interesting direction for future work. This would allow for the simulation of rare events for SRDEs under bifurcation (e.g. when ℓ>π in the Neumann Allen–Cahn case). The asymptotic optimality of such a scheme is guaranteed by Theorem 3.2. Even though the presence of non-constant saddle points with one unstable direction facilitates exits from D (13), the pre-asymptotic analysis of Sect. 4 is expected to be more complicated in this setting. This is due to the fact that the changes of measure uk0 (25) act on k0- dimensional subspaces of H. One then has to show that the linearization error is negligible by considering the behavior of the system on carefully chosen partitions of a k0-dimensional section of D.

Throughout this work we have considered SRDEs in one spatial dimension. In higher dimensions, equations like (76) are singular and a-priori ill-posed. Thus, one has to consider SRDEs with a spatially colored stochastic forcing or employ renormalization techniques. Metastability results for the renormalized two-dimensional Allen–Cahn can be found e.g. in [5, 52] and references therein. Importance sampling for linear equations (i.e. f=0) with colored noise has been considered in [46]. In the latter, the spatial covariance operator Q is assumed to be trace-class and diagonalizable with respect to the eigenbasis of the differential operator A. Carrying the analysis of this paper over to higher spatial dimensions is challenging since A and the linearized operator A+DF(x∗) do not necessarily have the same eigenbasis (e.g. in the case of the Allen–Cahn with Dirichlet boundary conditions in spatial dimension 2). In particular, the analysis of the exit direction in Sect. 4 would have to be generalized and take into account the non-commutativity of Q and A+DF(x∗).

Finally, we expect that the results of this paper can be used to design importance sampling schemes for simulating rare events in slow-fast systems of SRDEs. Similar work for multiscale diffusions in finite dimensions has been done in [50] and an MDP for multiscale SRDEs was recently proved in [30].

Appendix A

A.1 Proof of Lemma 4.1

A formal application of Itô’s formula to Zx∗(t,η^x∗ϵ,v(t)) and Taylor’s theorem to F yieldZx∗(τ^∗ϵ,v,η^x∗ϵ,v(τ^∗ϵ,v))-Zx∗(0,0)=∫0τ^∗ϵ,v∂tZx∗(s,η^x∗ϵ,v(s))ds+∫0τ^∗ϵ,v〈DηZx∗(s,η^x∗ϵ,v(s)),[A+DF(x∗)]η^x∗ϵ,v(s)〉Hds+ϵh(ϵ)2∫0τ^∗ϵ,v〈DηZx∗(s,η^x∗ϵ,v(s)),D2F(x∗+θ0ϵh(ϵ)η^∗ϵ,v(s))(η^∗ϵ,v(s),η^∗ϵ,v(s))〉Hds+∫0τ^∗ϵ,v〈DηZx∗(s,η^x∗ϵ,v(s)),v(s)-u(s,η^∗ϵ,v(s))〉Hds+12h2(ϵ)∫0τ^∗ϵ,vtr[Dη2Zx∗(s,η^x∗ϵ,v(s))]ds+1h(ϵ)∫0τ^∗ϵ,v〈DηZx∗(s,η^x∗ϵ,v(s)),dW(s)〉H≥∫0τ^∗ϵ,v[12‖u(s,η^∗ϵ,v(s))‖H2-14‖v(s)‖H2]ds+1h(ϵ)∫0τ^∗ϵ,v〈DηZx∗(s,η^x∗ϵ,v(s)),dW(s)〉H+ϵh(ϵ)2∫0τ^∗ϵ,v〈DηZx∗(s,η^x∗ϵ,v(s)),D2F(x∗+θ0ϵh(ϵ)η^∗ϵ,v(s))(η^∗ϵ,v(s),η^∗ϵ,v(s))〉Hds+∫0τ^∗ϵ,v{∂tZx∗(s,η^x∗ϵ,v(s))+Hx∗(η^x∗ϵ,v(s),DηZx∗(s,η^x∗ϵ,v(s)))+12h2(ϵ)tr[Dη2Zx∗(s,η^x∗ϵ,v(s))]}ds-12∫0τ^∗ϵ,v‖DηZx∗(s,η^x∗ϵ,v(s))-DηUx∗(s,η^x∗ϵ,v(s))‖H2ds,

where θ0∈(0,1) and we used (64) to obtain the inequality above. Taking expectation then yieldsE∫0τ^∗ϵ,v[12‖v(s)‖H2-‖u(s,η^∗ϵ,v(s))‖H2]ds≥2Zx∗(0,0)-2EZx∗(τ^∗ϵ,v,η^x∗ϵ,v(τ^∗ϵ,v))+2E∫0τ^∗ϵ,vHx∗ϵ(Ux∗,Zx∗)(s,η^∗ϵ,v(s))ds.

In view of (17) we conclude that-1h2(ϵ)logQϵ(uϵ)≥infv∈A[2Zx∗(0,0)-2EZx∗(τ^∗ϵ,v,η^x∗ϵ,v(τ^∗ϵ,v))+2E∫0τ^∗ϵ,vHx∗ϵ(Ux∗,Zx∗)(s,η^∗ϵ,v(s))ds].

The proof is complete.

A.2 Proof of Lemma 4.2

From the chain rule we computeDηZ(t,η)=(1-ζ)DηUδ(t,η)=(1-ζ)e-F1(η)δe-F1(η)δ+e-F2ϵδDηF1(η)=-a1f(1-ζ)e-F1(η)δe-F1(η)δ+e-F2ϵδe1f=-(1-ζ)a1fρϵ(η)2e1f⟨e1f,η⟩H.

Moreover,Dη2Z(t,η)=-2(1-ζ)a1f[ρϵ(η)e1f⊗e1f+1δ(ρϵ(η))2(1-ρϵ(η))2⟨e1f,η⟩He1f⊗e1f]=-2(1-ζ)a1fρϵ(η)[1+2δρϵ(η)(1-ρϵ(η))⟨e1f,η⟩H]e1f⊗e1f

and12h2(ϵ)tr[Dη2Z(t,η)]=-(1-ζ)a1fρϵ(η)h2(ϵ)[1+2δρϵ(η)(1-ρϵ(η))⟨e1f,η⟩H]∑k∈N0〈(e1f⊗e1f)ekf,ekf〉H=-(1-ζ)a1fρϵ(η)h2(ϵ)[1+2δρϵ(η)(1-ρϵ(η))⟨e1f,η⟩H].

In view of (23) we haveHx∗(η,DηZ(t,η))=-a1f〈η,DηZ(t,η)〉H-12‖DηZ(t,η)‖H2=2(1-ζ)(a1f)2ρϵ(η)⟨e1f,η⟩H2-4(1-ζ)2(a1f)2(ρϵ(η))22×⟨e1f,η⟩H2‖e1f‖H2=2(1-ζ)(a1f)2ρϵ(η)⟨e1f,η⟩H2[1-(1-ζ)ρϵ(η)].

Finally,-12‖DηZ(t,η)-DηUδ(t,η)‖H2=-ζ22‖DηUδ(t,η)‖H2=-2ζ2(a1f)2(ρϵ(η))2⟨e1f,η⟩H2.

In view of these computations we obtainHx∗ϵ(Uδ,Z)(t,η)=Hx∗ϵ(Z)(t,η)+Dx∗ϵ(Z)(t,η)-12‖DηZ(t,η)-DηUδ(t,η)‖H2=ϵh(ϵ)2〈DηZ(t,η),D2F(x∗+θ0ϵh(ϵ)η)(η,η)〉H+∂tZ(t,η)+Hx∗(η,DηZ(t,η))+12h2(ϵ)tr[Dη2Z(t,η)]-12‖DηZ(t,η)-DηU(t,η)‖H2=-ϵh(ϵ)(1-ζ)a1fρϵ(η)⟨e1f,η⟩H〈e1f,D2F(x∗+θ0ϵh(ϵ)η)(η,η)〉H+0+2(1-ζ)(a1f)2ρϵ(η)⟨e1f,η⟩H2[1-(1-ζ)ρϵ(η)]-(1-ζ)a1fρϵ(η)h2(ϵ)[1+2δρϵ(η)(1-ρϵ(η))⟨e1f,η⟩H]-2ζ2(a1f)2(ρϵ(η))2⟨e1f,η⟩H2.

The proof is complete.

A.3 Proof of Lemma 4.3

In view of the last term in (71) we have〈DηZ(t,η),D2F(x∗+θ0ϵh(ϵ)η)(η,η)〉H≥-ϵh(ϵ)2(1-ζ)a1fρϵ(η)|⟨e1f,η⟩H|⟨e1f,D2F(x∗+θ0ϵh(ϵ)η)(η,η)⟩H|≥-ϵh(ϵ)2(1-ζ)a1fρϵ(η)|⟨e1f,η⟩H|‖e1f‖H‖η2‖L∞‖∂x2f(x∗+θ0ϵh(ϵ)η)‖H≥-Cfϵh(ϵ)2(1-ζ)a1fρϵ(η)|⟨e1f,η⟩H|(1+‖(x∗+θ0ϵh(ϵ)η)p0-2‖H))≥-Cf,ℓϵh(ϵ)2(1-ζ)a1fρϵ(η)|⟨e1f,η⟩H|×(1+‖x∗‖∞p0-2+(θ0)p0-2‖ϵh(ϵ)η‖∞p0-2),

where we used the growth assumption (7) in the third inequality above. The estimate follows since ‖η‖L∞≤1 and θ0ϵh(ϵ)<1.

A.4 Proof of Lemma 4.4

In the region B1f we haveρϵ(η)=e-F1(η)-F2ϵδe-F1(η)-F2ϵδ+1

and e-F1(η)-F2ϵδ≤e-a1fh2(ϵ)h-2κ(ϵ)(1-h-2α(ϵ))/2. For any reasonable choice of h(ϵ), (e.g. continuous, positive, strictly decreasing) the latter implies that the second term in (74) is exponentially negligible. Ignoring the nonnegative terms, note that the linearization error (last term) is bounded from above in absolute value byCfϵh(ϵ)‖η‖Hρϵ(η)≤Cf,Le-ch2(ϵ)≤Ch-2(ϵ)e-ch2(ϵ)

which holds by assuming that ϵh(ϵ)<1 and applying the inequality e-cx≤(β/e)βγ-βx-β, which holds for all β,γ,x>0, with γ=c,x=h2(ϵ) and β=1. The proof is complete.

A.5 Proof of Lemma 4.5

(i) Our choice of K implies that e-K/(e-K+1)=3/4. Hence in this region we have ρϵ(η)≥3/4. From (74) and the fact that ρϵ>(ρϵ)2 we haveHx∗ϵ(Uδ,Z)(t,η)≥ρϵ(η)2(a1f)2⟨η,e1f⟩2((1-ζ)(1-ρϵ(η))+2(ζ-2ζ2))-1h2(ϵ)(1-ζ)ρϵ(η)a1f-2Cϵh(ϵ)(1-ζ)a1fρϵ(η)|⟨e1f,η⟩H|≥2(ζ0-2ζ02)ρϵ(η)2(a1f)2⟨η,e1f⟩2-1h2(ϵ)(1-ζ)ρϵ(η)a1f-2CLϵh(ϵ)(1-ζ)a1fρϵ(η)≥2(ζ0-2ζ02)916(a1f)2⟨η,e1f⟩2-1h2(ϵ)(1-ζ)a1f-2CLϵh(ϵ)(1-ζ)a1f≥1816(ζ0-2ζ02)(a1f)2(2h(ϵ)-2κ-h(ϵ)-2K)-12h2(ϵ)a1f-CLϵh(ϵ)a1f≥-1816(ζ0-2ζ02)(a1f)2h(ϵ)-2K-CLϵh(ϵ)a1f≥-CLϵh(ϵ)a1f,

where we used that ζ≥ζ0 and the Cauchy-Schwarz inequality to obtain the second inequality, ρϵ∈[3/4,1) for the third, ζ<1/2 and 2h(ϵ)-2κ-h(ϵ)-2K<⟨η,e1f⟩H2 in the fourth, and h2(κ-1)(ϵ)<9a1f(ζ0-2ζ02)/2, K<0 in the last line.

(ii) If h(ϵ) is such that limϵ→0ϵh3(ϵ)=0, the estimate follows from the first inequality in the last line of the previous display. In particular, we can writeHx∗ϵ(Uδ,Z)(t,η)≥ϵh(ϵ)(-1816(ζ0-2ζ02)(a1f)2Kϵh3(ϵ)-CLa1f)

and the estimate follows by letting ϵ be small enough to have ϵh3(ϵ)<-K1816(ζ0-2ζ02)(a1f)/CL (recall that -K>0.)

A.6 Proof of Lemma 4.6

(i) Case 1: Assume that η is such that ρϵ(η)≥1/2. Then the arguments and estimates of Lemma 4.5 carry over verbatim. The only difference lies in the range of ϵ.

Case 2: Assume that η is such that ρϵ(η)<1/2. Focusing on the first two terms in (74) we see thatHx∗ϵ(Uδ,Z)(t,η)≥(1-ζ)ρϵ(η)2(a1f)2⟨η,e1f⟩2-1h2(ϵ)(1-ζ)ρϵ(η)a1f+2(ζ-2ζ2)(ρϵ(η))2(a1f)2⟨η,e1f⟩2-2Cϵh(ϵ)(1-ζ)a1fρϵ(η)|⟨e1f,η⟩H|≥ρϵ(η)(1-ζ)a1f(a1f2⟨η,e1f⟩2-1h2(ϵ))+2(ζ-2ζ2)ρϵ(η)2(a1f)2⟨η,e1f⟩2-2CLϵh(ϵ)(1-ζ)a1fρϵ(η)2≥ρϵ(η)(1-ζ)a1f(a1f2h(ϵ)-2(κ+α)-1h2(ϵ))+2(ζ-2ζ2)ρϵ(η)2(a1f)2⟨η,e1f⟩2-2CLϵh(ϵ)(1-ζ)a1fρϵ(η)2≥2(ζ-2ζ2)ρϵ(η)2(a1f)2⟨η,e1f⟩2-2CLϵh(ϵ)(1-ζ)a1fρϵ(η)2

where we used that h(ϵ)2(κ-1)≤h(ϵ)2(κ+α-1)≤a1f/2 to obtain the last line. The estimate then follows since the first term on the last line is nonnegative.

(ii) Continuing from the last inequality we haveHx∗ϵ(Uδ,Z)(t,η)≥2ρϵ(η)2a1f((ζ0-2ζ02)a1fh(ϵ)-2(κ+α)-CLϵh(ϵ))

and the estimate follows by letting ϵ small enough to have ϵh(ϵ)1+2(κ+α)≤ϵh3(ϵ)≤(ζ0-2ζ02)a1f/CL.

This work was partially supported by the National Science Foundation (DMS 1550918, DMS 2107856, DMS 2311500) and Simons Foundation Award 672441.

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Abramowitz M Stegun IA Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables 1964 Washington, D.C. US Government Printing Office
Abramowitz, M., Stegun, I.A.: Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables, vol. 55. US Government Printing Office, Washington, D.C. (1964)
2. Allen SM Cahn JW A microscopic theory for antiphase boundary motion and its application to antiphase domain coarsening Acta Metall. 1979 27 6 1085 1095
Allen, S.M., Cahn, J.W.: A microscopic theory for antiphase boundary motion and its application to antiphase domain coarsening. Acta Metall. 27(6), 1085–1095 (1979)
3. Asmussen S Glynn PW Stochastic Simulation: Algorithms and Analysis 2007 Berlin Springer
Asmussen, S., Glynn, P.W.: Stochastic Simulation: Algorithms and Analysis, vol. 57. Springer, Berlin (2007)
4. Berglund, N.: An introduction to singular stochastic PDEs: Allen–Cahn equations, metastability and regularity structures. arXiv:1901.07420 (2019)
5. Berglund N Di Gesù G Weber H An Eyring–Kramers law for the stochastic Allen–Cahn equation in dimension two Electron. J. Probab. 2017 22 1 27
Berglund, N., Di Gesù, G., Weber, H.: An Eyring–Kramers law for the stochastic Allen–Cahn equation in dimension two. Electron. J. Probab. 22, 1–27 (2017)
6. Berglund N Gentz B Sharp estimates for metastable lifetimes in parabolic SPDEs: Kramers’ law and beyond Electron. J. Probab. 2013 18 1 58
Berglund, N., Gentz, B.: Sharp estimates for metastable lifetimes in parabolic SPDEs: Kramers’ law and beyond. Electron. J. Probab. 18, 1–58 (2013)
7. Bezemek, Z. Spiliopoulos, K.: Moderate deviations for fully coupled multiscale weakly interacting particle systems. arXiv:2202.08403 (2022)
8. Bogachev VI Measure Theory 2007 Berlin Springer
Bogachev, V.I.: Measure Theory, vol. 1. Springer, Berlin (2007)
9. Brigham EO The Fast Fourier Transform and Its Applications 1988 Hoboken Prentice-Hall, Inc.
Brigham, E.O.: The Fast Fourier Transform and Its Applications. Prentice-Hall, Inc., Hoboken (1988)
10. Budhiraja, A, Dupuis, P: Analysis and approximation of rare events. In: Representations and Weak Convergence Methods. Series Prob. Theory and Stoch. Modelling, vol. 94 (2019)
11. Budhiraja A Dupuis P Ganguly A Moderate deviation principles for stochastic differential equations with jumps Ann. Probab. 2016 44 3 1723 1775
Budhiraja, A., Dupuis, P., Ganguly, A.: Moderate deviation principles for stochastic differential equations with jumps. Ann. Probab. 44(3), 1723–1775 (2016)
12. Budhiraja A Dupuis P Maroulas V Large deviations for infinite dimensional stochastic dynamical systems Ann. Probab. 2008 36 1390 1420
Budhiraja, A., Dupuis, P., Maroulas, V.: Large deviations for infinite dimensional stochastic dynamical systems. Ann. Probab. 36, 1390–1420 (2008)
13. Cerrai S Second Order PDE’s in Finite and Infinite Dimension: A Probabilistic Approach 2001 Berlin Springer
Cerrai, S.: Second Order PDE’s in Finite and Infinite Dimension: A Probabilistic Approach. Springer, Berlin (2001)
14. Cerrai S Röckner M Large deviations for stochastic reaction–diffusion systems with multiplicative noise and non-Lipshitz reaction term Ann. Probab. 2004 32 1B 1100 1139
Cerrai, S., Röckner, M.: Large deviations for stochastic reaction–diffusion systems with multiplicative noise and non-Lipshitz reaction term. Ann. Probab. 32(1B), 1100–1139 (2004)
15. Chafee N Infante EF A bifurcation problem for a nonlinear partial differential equation of parabolic type Appl. Anal. 1974 4 1 17 37
Chafee, N., Infante, E.F.: A bifurcation problem for a nonlinear partial differential equation of parabolic type. Appl. Anal. 4(1), 17–37 (1974)
16. Da Prato G Pritchard AJ Zabczyk J On minimum energy problems SIAM J. Control Optim. 1991 29 1 209 221
Da Prato, G., Pritchard, A.J., Zabczyk, J.: On minimum energy problems. SIAM J. Control Optim. 29(1), 209–221 (1991)
17. Da Prato G Zabczyk J Stochastic Equations in Infinite Dimensions 2014 Cambridge Cambridge University Press
Da Prato, G., Zabczyk, J.: Stochastic Equations in Infinite Dimensions. Cambridge University Press, Cambridge (2014)
18. Day MV Large deviations results for the exit problem with characteristic boundary J. Math. Anal. Appl. 1990 147 1 134 153
Day, M.V.: Large deviations results for the exit problem with characteristic boundary. J. Math. Anal. Appl. 147(1), 134–153 (1990)
19. Di Nezza E Palatucci G Valdinoci E Hitchhiker’s guide to the fractional Sobolev spaces Bulletin des sciences mathématiques 2012 136 5 521 573
Di Nezza, E., Palatucci, G., Valdinoci, E.: Hitchhiker’s guide to the fractional Sobolev spaces. Bulletin des sciences mathématiques 136(5), 521–573 (2012)
20. Dupuis P Johnson D Moderate deviations for recursive stochastic algorithms Stoch. Syst. 2015 5 1 87 119
Dupuis, P., Johnson, D.: Moderate deviations for recursive stochastic algorithms. Stoch. Syst. 5(1), 87–119 (2015)
21. Dupuis P Johnson D Moderate deviations-based importance sampling for stochastic recursive equations Adv. Appl. Probab. 2017 49 4 981 1010
Dupuis, P., Johnson, D.: Moderate deviations-based importance sampling for stochastic recursive equations. Adv. Appl. Probab. 49(4), 981–1010 (2017)
22. Dupuis P Spiliopoulos K Wang H Importance sampling for multiscale diffusions Multiscale Model. Simul. 2012 10 1 1 27
Dupuis, P., Spiliopoulos, K., Wang, H.: Importance sampling for multiscale diffusions. Multiscale Model. Simul. 10(1), 1–27 (2012)
23. Dupuis P Spiliopoulos K Zhou X Escaping from an attractor: Importance sampling and rest points I Ann. Appl. Probab. 2015 25 5 2909 2958
Dupuis, P., Spiliopoulos, K., Zhou, X.: Escaping from an attractor: Importance sampling and rest points I. Ann. Appl. Probab. 25(5), 2909–2958 (2015)
24. Dupuis P Wang H Importance sampling, large deviations, and differential games Stoch. Int. J. Probab. Stoch. Process. 2004 76 6 481 508
Dupuis, P., Wang, H.: Importance sampling, large deviations, and differential games. Stoch. Int. J. Probab. Stoch. Process. 76(6), 481–508 (2004)
25. Dupuis P Wang H Subsolutions of an Isaacs equation and efficient schemes for importance sampling Math. Oper. Res. 2007 32 3 723 757
Dupuis, P., Wang, H.: Subsolutions of an Isaacs equation and efficient schemes for importance sampling. Math. Oper. Res. 32(3), 723–757 (2007)
26. Espichán Carrillo JA Maia A Jr Mostepanenko VM Jacobi elliptic solutions of λϕ4 theory in a finite domain Int. J. Modern Phys. A 2000 15 17 2645 2659
Espichán Carrillo, J.A., Maia, A., Jr., Mostepanenko, V.M.: Jacobi elliptic solutions of theory in a finite domain. Int. J. Modern Phys. A 15(17), 2645–2659 (2000)
27. Faris WG Jona-Lasinio G Large fluctuations for a nonlinear heat equation with noise J. Phys. A Math. General 1982 15 10 3025
Faris, W.G., Jona-Lasinio, G.: Large fluctuations for a nonlinear heat equation with noise. J. Phys. A Math. General 15(10), 3025 (1982)
28. Freidlin MI Wentzell AD Random Perturbations of Dynamical Systems 1998 Berlin Springer
Freidlin, M.I., Wentzell, A.D.: Random Perturbations of Dynamical Systems. Springer, Berlin (1998)
29. Gao F Zhao X Delta method in large deviations and moderate deviations for estimators Ann. Stat. 2011 39 2 1211 1240
Gao, F., Zhao, X.: Delta method in large deviations and moderate deviations for estimators. Ann. Stat. 39(2), 1211–1240 (2011)
30. Gasteratos I Salins M Spiliopoulos K Moderate deviations for systems of slow–fast stochastic reaction–diffusion equations Stoch. Partial Differ. Equ. Anal. Comput. 2022 11 2 503 598
Gasteratos, I., Salins, M., Spiliopoulos, K.: Moderate deviations for systems of slow–fast stochastic reaction–diffusion equations. Stoch. Partial Differ. Equ. Anal. Comput. 11(2), 503–598 (2022)
31. Gayrard V Bovier A Eckhoff M Klein M Metastability in reversible diffusion processes I: sharp asymptotics for capacities and exit times J. Eur. Math. Soc. 2004 6 4 399 424
Gayrard, V., Bovier, A., Eckhoff, M., Klein, M.: Metastability in reversible diffusion processes I: sharp asymptotics for capacities and exit times. J. Eur. Math. Soc. 6(4), 399–424 (2004)
32. Gayrard V Bovier A Klein M Metastability in reversible diffusion processes II: Precise asymptotics for small eigenvalues J. Eur. Math. Soc. 2005 7 1 69 99
Gayrard, V., Bovier, A., Klein, M.: Metastability in reversible diffusion processes II: Precise asymptotics for small eigenvalues. J. Eur. Math. Soc. 7(1), 69–99 (2005)
33. Grama IG On moderate deviations for martingales Ann. Probab. 1997 25 1 152 183
Grama, I.G.: On moderate deviations for martingales. Ann. Probab. 25(1), 152–183 (1997)
34. Guillin A Averaging principle of SDE with small diffusion: moderate deviations Ann. Probab. 2003 31 1 413 443
Guillin, A.: Averaging principle of SDE with small diffusion: moderate deviations. Ann. Probab. 31(1), 413–443 (2003)
35. Högele, M.: Metastability of the Chafee-Infante equation with small heavy-tailed Lévy Noise (2011)
36. Jacquier A Spiliopoulos K Pathwise moderate deviations for option pricing Math. Finance 2020 30 2 426 463
Jacquier, A., Spiliopoulos, K.: Pathwise moderate deviations for option pricing. Math. Finance 30(2), 426–463 (2020)
37. Jentzen A Kloeden PE Overcoming the order barrier in the numerical approximation of stochastic partial differential equations with additive space-time noise Proc. R. Soc. A Math. Phys. Eng. Sci. 2009 465 2102 649 667
Jentzen, A., Kloeden, P.E.: Overcoming the order barrier in the numerical approximation of stochastic partial differential equations with additive space-time noise. Proc. R. Soc. A Math. Phys. Eng. Sci. 465(2102), 649–667 (2009)
38. Kallenberg WCM On moderate deviation theory in estimation Ann. Stat. 1983 11 498 504
Kallenberg, W.C.M.: On moderate deviation theory in estimation. Ann. Stat. 11, 498–504 (1983)
39. Kuznetsov AN Tinyakov PG Periodic instanton bifurcations and thermal transition rate Phys. Lett. B 1997 406 1–2 76 82
Kuznetsov, A.N., Tinyakov, P.G.: Periodic instanton bifurcations and thermal transition rate. Phys. Lett. B 406(1–2), 76–82 (1997)
40. Landau LD Ginzburg VL On the theory of superconductivity Zh. Eksp. Teor. Fiz. 1950 20 1064
Landau, L.D., Ginzburg, V.L.: On the theory of superconductivity. Zh. Eksp. Teor. Fiz. 20, 1064 (1950)
41. Lunardi A Analytic Semigroups and Optimal Regularity in Parabolic Problems 2012 Berlin Springer
Lunardi, A.: Analytic Semigroups and Optimal Regularity in Parabolic Problems. Springer, Berlin (2012)
42. Maier RS Stein DL Droplet nucleation and domain wall motion in a bounded interval Phys. Rev. Lett. 2001 87 270601 11800865
Maier, R.S., Stein, D.L.: Droplet nucleation and domain wall motion in a bounded interval. Phys. Rev. Lett. 87, 270601 (2001)11800865
43. Maier, RS., Stein, D.L.: Effects of weak spatiotemporal noise on a bistable one-dimensional system. In: Noise in Complex Systems and Stochastic Dynamics, Vol. 5114, pp. 67–78. International Society for Optics and Photonics (2003)
44. Morse MR Spiliopoulos, Konstantinos: moderate deviations for systems of slow-fast diffusions Asymptot. Anal. 2017 105 3–4 97 135
Morse, M.R.: Spiliopoulos, Konstantinos: moderate deviations for systems of slow-fast diffusions. Asymptot. Anal. 105(3–4), 97–135 (2017)
45. Salins, M.: Systems of small-noise stochastic reaction–diffusion equations satisfy a large deviations principle that is uniform over all initial data. arXiv:2008.01140 (2020)
46. Salins M Spiliopoulos K Rare event simulation via importance sampling for linear SPDE’s Stoch. Partial Differ. Equ. Anal. Comput. 2017 5 4 652 690
Salins, M., Spiliopoulos, K.: Rare event simulation via importance sampling for linear SPDE’s. Stoch. Partial Differ. Equ. Anal. Comput. 5(4), 652–690 (2017)
47. Salins M Spiliopoulos K Metastability and exit problems for systems of stochastic reaction–diffusion equations Ann. Probab. 2021 49 5 2317 2370
Salins, M., Spiliopoulos, K.: Metastability and exit problems for systems of stochastic reaction–diffusion equations. Ann. Probab. 49(5), 2317–2370 (2021)
48. Schneider G Uecker H Nonlinear PDEs 2017 Providence American Mathematical Society
Schneider, G., Uecker, H.: Nonlinear PDEs, vol. 182. American Mathematical Society, Providence (2017)
49. Spiliopoulos, K.: Importance sampling for metastable and multiscale dynamical systems. In: Stochastic Processes, Multiscale Modeling, and Numerical Methods for Computational Cellular Biology, pp. 29–53. Springer (2017)
50. Spiliopoulos K Morse MR Importance sampling for slow–fast diffusions based on moderate deviations Multiscale Model. Simul. 2020 18 1 315 350
Spiliopoulos, K., Morse, M.R.: Importance sampling for slow–fast diffusions based on moderate deviations. Multiscale Model. Simul. 18(1), 315–350 (2020)
51. Spiliopoulos K Morse MR Importance sampling for slow-fast diffusions based on moderate deviations Multiscale Model. Simul. 2020 18 1 315 350
Spiliopoulos, K., Morse, M.R.: Importance sampling for slow-fast diffusions based on moderate deviations. Multiscale Model. Simul. 18(1), 315–350 (2020)
52. Tsatsoulis P Weber H Exponential loss of memory for the 2-dimensional Allen-Cahn equation with small noise Probab. Theory Related Fields 2020 177 1 257 322
Tsatsoulis, P., Weber, H.: Exponential loss of memory for the 2-dimensional Allen-Cahn equation with small noise. Probab. Theory Related Fields 177(1), 257–322 (2020)
53. Vanden-Eijnden E Weare J Rare event simulation of small noise diffusions Commun. Pure Appl. Math. 2012 65 12 1770 1803
Vanden-Eijnden, E., Weare, J.: Rare event simulation of small noise diffusions. Commun. Pure Appl. Math. 65(12), 1770–1803 (2012)
54. Wakasa T Exact eigenvalues and eigenfunctions associated with linearization for Chafee-Infante problem Funkcialaj Ekvacioj 2006 49 2 321 336
Wakasa, T.: Exact eigenvalues and eigenfunctions associated with linearization for Chafee-Infante problem. Funkcialaj Ekvacioj 49(2), 321–336 (2006)
55. Wang, R., Zhang, T.: Moderate deviations for stochastic reaction–diffusion equations with multiplicative noise. Potential Anal. 42(1), 99–113 (2015)
