
==== Front
J Optim Theory Appl
J Optim Theory Appl
Journal of Optimization Theory and Applications
0022-3239
1573-2878
Springer US New York

2500
10.1007/s10957-024-02500-8
Article
Second Order Dynamics Featuring Tikhonov Regularization and Time Scaling
Csetnek Ernö Robert
Karapetyants Mikhail A. mikhail.karapetyants@univie.ac.at

https://ror.org/03prydq77 grid.10420.37 0000 0001 2286 1424 Faculty of Mathematics, University of Vienna, Oskar-Morgenstern-Platz 1, 1090 Vienna, Austria
Communicated by Russell Luke.

21 8 2024
21 8 2024
2024
202 3 13851420
11 9 2023
11 7 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by/4.0/ Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
In a Hilbert setting we aim to study a second order in time differential equation, combining viscous and Hessian-driven damping, containing a time scaling parameter function and a Tikhonov regularization term. The dynamical system is related to the problem of minimization of a nonsmooth convex function. In the formulation of the problem as well as in our analysis we use the Moreau envelope of the objective function and its gradient and heavily rely on their properties. We show that there is a setting where the newly introduced system preserves and even improves the well-known fast convergence properties of the function and Moreau envelope along the trajectories and also of the gradient of Moreau envelope due to the presence of time scaling. Moreover, in a different setting we prove strong convergence of the trajectories to the element of minimal norm from the set of all minimizers of the objective. The manuscript concludes with various numerical results.

Keywords

Nonsmooth convex optimization
Damped inertial dynamics
Hessian-driven damping
Time scaling
Moreau envelope
Proximal operator
Tikhonov regularization
Mathematics Subject Classification

37N40
46N10
49M99
65K05
65K10
90C25
http://dx.doi.org/10.13039/501100002428 Austrian Science Fund project W 1260 Karapetyants Mikhail A. http://dx.doi.org/10.13039/100018987 Ministerul Cercetării, Inovării şi Digitalizarii project number PN-III-P1-1.1-TE-2021-0138, within PNCDI III Csetnek Ernö Robert issue-copyright-statement© Springer Science+Business Media, LLC, part of Springer Nature 2024
==== Body
pmcIntroduction

In the Hilbert setting H, where ⟨·,·⟩ denotes the inner product and the norm is defined as usual ‖·‖=⟨·,·⟩, we will study the convergence properties of the following second order in time differential equation1 x¨(t)+αtx˙(t)+βddt∇Φλ(t)(x(t))+b(t)∇Φλ(t)(x(t))+ε(t)x(t)=0fort≥t0,

with initial conditions x(t0)=x0∈H, x˙(t0)=x˙0∈H, where α,βandt0>0, λ:[t0,+∞)↦R+ and b:[t0,+∞)↦R+ are non-negative, non-decreasing and differentiable, Φ:H↦R¯=R∪{±∞} is a proper, convex and lower semicontinuous function and Φλ is its Moreau envelope of the index λ>0 and the function ε:[t0,+∞)↦R+ is continuously differentiable and non-increasing with the property limt→+∞ε(t)=0. In addition, we assume that argminΦ, which is the set of global minimizers of Φ, is not empty and denote by Φ∗ the optimal objective value of Φ. The system (1) has a connection to the minimization problemminx∈HΦ(x)

of a proper, convex and lower semicontinuous function Φ. Studying such systems provides better understanding of their discrete counterpart—optimization algorithms, since there is a strong connection between them, and the question of transitioning from one to another attracts a lot of attention in the modern literature.

One of the main goals of this research is to improve (compared to [23]) the fast rates of convergence for the Moreau envelope of the objective function and the objective function itself to Φ∗, as well as for the gradient of the Moreau envelope of the objective function in terms of the Moreau parameter function λ and the time scaling function b. Moreover, we also deduce the strong convergence of the trajectory of the dynamics to the minimal norm element of argminΦ. We introduce two settings with different assumptions for each result. To conclude we provide multiple numerical results in order to illustrate our theoretical discoveries.

Nonsmooth Optimization with Time Scaling

In the smooth setting the pioneering research in studying second order dynamical systems was conducted by Su–Boyd–Candes [30] for the sake of obtaining faster asymptotic convergence for convex functions. They managed to deduce the rates of convergence of the function values being of the order 1t2. Later Attouch–Peypouquet–Redont [20] also established the weak (and in some particular cases the strong) convergence of the trajectories to a minimizer of the objective function. In [19] the same authors continued the development in this direction by adding Hessian-driven damping term in order to obtain the rates for the gradient of the objective function and to eliminate any possible oscillations in the dynamical behaviour of the trajectories.

Concerning the nonsmooth setting we must point out that the Moreau envelope of a proper, convex and lower semicontinuous function Φ:H→R¯ proved to be of a significant importance in designing continuous-time approaches and numerical algorithms for the minimization of nonsmooth functions. The rigorous definition of this construction isΦλ:H→R,Φλ(x)=infy∈HΦ(y)+12λ‖x-y‖2,

where λ>0 is the parameter of the Moreau envelope (see, for instance, [21]). One of the most important properties of Moreau approximation is that for every λ>0, the functions Φ and Φλ share the same optimal objective value and also the same set of minimizers. Moreover, Φλ is convex and continuously differentiable with2 ∇Φλ(x)=1λ(x-proxλΦ(x))∀x∈H,

and ∇Φλ is 1λ-Lipschitz continuous, whereproxλΦ:H→H,proxλΦ(x)=argminy∈HΦ(y)+12λ‖x-y‖2,

denotes the proximal operator of Φ of parameter λ. The last fact we would like to mention is that for every x∈H, the function λ∈(0,+∞)→Φλ(x) is nonincreasing and differentiable (see [14], Lemma A1), namely,3 ddλΦλ(x)=-12‖∇Φλ(x)‖2∀λ>0.

Our research is a logical continuation of the one conducted in [24], where authors applied the time rescaling technique to a nonsmooth optimization problem (for more information on time scaling see also [5, 10, 11, 13]). They considered the following system4 x¨(t)+αtx˙(t)+β(t)ddt∇Φλ(t)(x(t))+b(t)∇Φλ(t)(x(t))=0,

where α≥1, t0>0, and β:[t0,+∞)↦[0,+∞) and b,λ:[t0,+∞)↦(0,+∞) are differentiable functions. On the one hand, the presence of the Hessian damping term is believed to help reducing the oscillations in the dynamical behaviour and provides the rates for the gradient of the objective function Φ. On the other hand, the time-scaling technique (which is considered to be an artificial way to speed up the convergence of values) affects the convergence rates while bringing more restrictions to the analysis. The following properties were establishedΦλ(t)(x(t))-Φ∗=o1t2b(t)and‖x˙(t)‖=o1tast→+∞,

from where through proximal mapping the convergence rates for the objective function Φ itself along the trajectory were obtainedΦ(proxλ(t)Φ(x(t)))-Φ∗=o1t2b(t)and‖proxλ(t)Φ(x(t))-x(t)‖=oλ(t)tb(t)ast→+∞.

Note that by taking b(·)≡1 we arrive at the well-known convergence rate of the values being of the order o1t2. In addition, the following rates for the gradient of the Moreau envelope were deduced‖∇Φλ(t)(x(t))‖=o1tb(t)λ(t),ast→+∞.

Finally, the weak convergence of the trajectories x(t) to a minimizer of Φ as t→+∞ was obtained.

In our analysis we borrow some ideas of [24] and develop them further in order to fit the new setting, namely, to adapt to a presence of the whole new term—Tikhonov regularization. The analysis becomes more involved and technical, some fundamental properties of Tikhonov regularization had to be proved for a nonsmooth setting. Its presence affects the set of conditions, which we have to impose on the system parameters: even though some of the conditions are formulated in the same spirit as in [24] (for instance, (11) and (14)), the other ones are completely new due to the presence of the Tikhonov term. Moreover, depending on how fast ε decays, two different setting arise providing different fundamental results (Sects. 3 and 4).

Tikhonov Regularization

It turned out that having additional term with specific properties in a system equation leads to improving the weak convergence of the trajectories to a minimizer of the objective function Φ to a strong one to the element of minimal norm of argminΦ. Such systems were studied, for instance, in [4, 6, 9, 12, 17, 23, 27]. The main goal of such a research is to show that these systems preserve all the typical properties of the second order in time dynamical system (fast convergence of the values, the rates for the gradient etc.) but moreover there is an improvement to the strong convergence of the trajectories to the minimal norm solution instead of a weak one to an arbitrary minimizer. One of the many examples of such systems is presented below (see [23])x¨(t)+αtx˙(t)+β∇2Φ(x(t))x˙(t)+∇Φ(x(t))+ε(t)x(t)=0fort≥t0,

where α≥3, t0>0, Φ:H↦R is twice continuously differentiable and convex and for the rest of the section the function ε:[t0,+∞)↦R+ is continuously differentiable and non-increasing with the property limt→+∞ε(t)=0. In that manuscript they provided two settings: one for the fast convergence of values obtainingΦ(x(t))-Φ∗=o1t2,ast→+∞

and the weak convergence of the trajectories to a minimizer of Φ and another setting for the strong convergence of x to x∗, as t→+∞.

Another fine example is given in [4]:5 x¨(t)+αε(t)x˙(t)+∇Φ(x(t))+ε(t)x(t)=0fort≥t0,

where α, t0>0 and Φ:H↦R is continuously differentiable and convex. In that paper authors obtained the rates for the function values Φ(x(t))-Φ∗, as well as for the quantity ‖x(t)-xε(t)‖, as t→+∞, where xε(t)=argminHΦ(x)+ε(t)‖x‖22. Thus, they assured the strong convergence of the trajectories to the minimal norm solution x∗=projargminΦ(0) under the appropriate assumptions and properly chosen energy functional, using the properties of Tikhonov regularization. The most important thing about this approach is that authors were able to establish fast convergence of values and strong convergence of the trajectories in the very same setting.

The next step was done in [6]:x¨(t)+αε(t)x˙(t)+βddt(∇φt(x(t))+(p-1)ε(t)x(t))+∇φt(x(t))=0fort≥t0,

where φt(x)=Φ(x)+ε(t)‖x‖22, Φ:H↦R is twice continuously differentiable and convex and p∈[0,1]. This system while preserving all the properties of (5), additionally provides the integral estimate for the norm of the gradient of φt.

Our Contribution

In that paper we will develop the ideas presented in [23] to cover the nonsmooth case with time scaling. We will obtain the fast convergence of the function values (as well as for the gradient of the Moreau envelope of the objective fucntion Φ) for the family of dynamical systems (1) governed by the Moreau envelope of the nonsmooth function Φ and having the Tiknonov term in their formulation:Φλ(t)(x(t))-Φ∗=o1t2b(t)ast→+∞;

in terms of the function itself:Φ(proxλ(t)Φ(x(t)))-Φ∗=o1t2b(t)ast→+∞,

where‖proxλ(t)Φ(x(t))-x(t)‖=oλ(t)tb(t)ast→+∞

and finally‖∇Φλ(t)(x(t))‖=o1tb(t)λ(t)ast→+∞.

We will also deduce (under some appropriate conditions) the following resultlim inft→+∞‖x(t)-x∗‖=0,

which under some restrictions will be improved to the full strong convergence of the trajectories of (1) to the minimal norm solution.

The paper is organized in the following way. Section 2 is devoted to some preliminary results, which we will need later. We will establish the fast rates of convergence of function values and its Moreau envelope, as well as the gradient of Moreau envelope along the trajectories of the dynamical system (Sect. 3). We will show that under some assumptions the strong convergence of the trajectories to the element of minimal norm from the set of all minimizers of the objective function takes place (Sect. 4). We will provide two settings for the polynomial choice of parameter functions to fulfill the assumptions made through the analysis (Sect. 5) and equip this manuscript with various numerical results (Sect. 6).

Preparatory Results

We start with the following lemma (see [21], Proposition 12.22, for the first term of the lemma and [18], Appendix, A1, for the second one).

Lemma 1

Let Φ:H↦R¯ be a proper, convex and lower semicontinuous function, λ,μ>0. Then (Φλ)μ=Φλ+μ.

proxμΦλ=λλ+μId+μλ+μprox(λ+μ)Φ.

Let us mention two key properties of the Tikhonov regularization, which we will use later in the analysis (see, for instance, [2] or [21] Theorem 23.44 for its classic analogue). First let us introduce the strongly convex function φε(t),λ(t):H↦R as φε(t),λ(t)(x)=Φλ(t)(x)+ε(t)‖x‖22 and denote the unique minimizer of φε(t),λ(t) as xε(t),λ(t)=argminHφε(t),λ(t). Thus, the first order optimality condition reads as6 ∇Φλ(t)(xε(t),λ(t))+ε(t)xε(t),λ(t)=0.

Now we are ready to formulate the following result:

Lemma 2

Suppose that7 limt→+∞λ(t)ε(t)=0.

Then the following properties of the mapping t↦xε(t),λ(t) are satisfied:8 forx∗=projargminΦ(0),‖xε(t),λ(t)‖≤‖x∗‖for allt≥t0

and9 limt→+∞‖xε(t),λ(t)-x∗‖=0.

Proof

By the monotonicity of ∇Φλ we deduce∇Φλ(t)(xε(t),λ(t))-∇Φλ(t)(x∗),xε(t),λ(t)-x∗≥0.

By (6) we obtain-ε(t)xε(t),λ(t),xε(t),λ(t)-x∗=ε(t)-‖xε(t),λ(t)‖2+xε(t),λ(t),x∗≥0.

Using Cauchy–Schwarz inequality we derive‖xε(t),λ(t)‖≤‖x∗‖.

This proves the first claim. For the second one consider (6) again and note that it is equivalent toxε(t),λ(t)=prox1ε(t)Φλ(t)(0)=proxλ(t)+1ε(t)Φ(0)λ(t)ε(t)+1

by the item 2. of Lemma 1. Note that λ(t)+1ε(t)→+∞, as t→+∞. Thus, the rest of the proof goes in line with Theorem 23.44 of [21]. □

Our nearest goal is to deduce the existence and uniqueness of the solutions of the dynamical system (1). Suppose β>0. Let us integrate (1) from t0 to t to obtainx˙(t)+β∇Φλ(t)(x(t))+∫t0tαsx˙(s)+b(s)∇Φλ(s)(x(s))+ε(s)x(s)ds-x˙(t0)+β∇Φλ(t0)(x(t0))=0.

Denoting z(t):=∫t0tαsx˙(s)+b(s)∇Φλ(s)(x(s))+ε(s)x(s)ds-(x˙(t0)+β∇Φλ(t0)(x0))) for every t≥t0 and noticing that z˙(t)=αtx˙(t)+b(t)∇Φλ(t)(x(t))+ε(t)x(t) we deduce, that (1) is equivalent tox˙(t)+β∇Φλ(t)(x(t))+z(t)=0,z˙(t)-αtx˙(t)-b(t)∇Φλ(t)(x(t))-ε(t)x(t)=0,x(t0)=x0,z(t0)=-x˙(t0)+β∇Φλ(t0)(x0).

Let us multiply the first line by the function b and the second one by the constant β and then sum them up to get rid of the gradient of the Moreau envelope in the second equationx˙(t)+β∇Φλ(t)(x(t))+z(t)=0,βz˙(t)+b(t)-αβtx˙(t)-βε(t)x(t)+b(t)z(t)=0,x(t0)=x0,z(t0)=-x˙(t0)+β∇Φλ(t0)(x0).

We denote now y(t)=βz(t)+b(t)-αβtx(t), and, after simplification, we obtain the following equivalent formulation for the dynamical systemx˙(t)+β∇Φλ(t)(x(t))+αt-b(t)βx(t)+1βy(t)=0,y˙(t)-b˙(t)+αβt2+βε(t)+b2(t)β-αb(t)tx(t)+b(t)βy(t)=0,x(t0)=x0,y(t0)=-βx˙(t0)+β∇Φλ(t0)(x0)+b(t0)-αβt0x0.

In case β=0 for every t≥t0, (1) can be equivalently written asx˙(t)-y(t)=0,y˙(t)+αty(t)+b(t)∇Φλ(t)(x(t))+ε(t)x(t)=0,x(t0)=x0,y(t0)=x˙(t0).

Based on the two reformulations of the dynamical system (1) we formulate the following existence and uniqueness result, which is a consequence of Cauchy-Lipschitz theorem for strong global solutions. The result can be proved in the lines of the proofs of Theorem 1 in [16] or of Theorem 1.1 in [19] with some small adjustments.

Theorem 3

Suppose that there exists λ0>0 such that λ(t)≥λ0 for all t≥t0. Then for every (x0,x˙(t0))∈H·H there exists a unique strong global solution x:[t0,+∞)↦H of the continuous dynamics (1) which satisfies the Cauchy initial conditions x(t0)=x0 and x˙(t0)=x˙0.

Fast Convergence Rates of the Function and Moreau Envelope Values

This chapter is devoted to obtaining the rates of convergence for the Moreau envelope values and for the values of function Φ itself. We will heavily rely on the tools and techniques provided by the Lyapunov analysis. We introduce a slightly modified energy function from [23]. For 2≤q≤α-1 we define10 Eq(t)=(t2b(t)-β(q+2-α)t)Φλ(t)(x(t))-Φ∗+t2ε(t)2‖x(t)‖2+12‖q(x(t)-x∗)+tx˙(t)+β∇Φλ(t)(x(t))‖2+q(α-1-q)2‖x(t)-x∗‖2.

The key assumptions which are essential to our analysis are the following: for all t≥t0

Theorem 4

Suppose α≥3 and assume that (11), (12), (13), (14) hold for all t≥t0. ThenΦλ(t)(x(t))-Φ∗=O1t2b(t),ast→+∞,‖x˙(t)+β∇Φλ(t)(x(t))‖=O1t,ast→+∞.

Moreover, one has for all a≥1tε(t)‖x(t)-x∗‖2,tε(t)‖x(t)‖2,(α-3)tb(t)-t2b˙(t)+β(2-α)Φλ(t)(x(t))-Φ∗andt2b(t)-βtλ˙(t)2-β2t+βt2b(t)-1a‖∇Φλ(t)(x(t))‖2∈L1([t0,+∞),R).

If, in addition, α>3 and (15) holds, then the trajectory x is bounded and∫t0+∞t‖x˙(t)‖2dt<+∞

and∫t0+∞tb(t)Φλ(s)(x(s))-Φ∗<+∞.

Proof

Let us compute the time derivative of the energy function. For every t≥t0 using (3) we deriveE˙q(t)=2tb(t)+t2b˙(t)-β(q+2-α)Φλ(t)(x(t))-Φ∗+t2ε(t)⟨x(t),x˙(t)⟩+q(α-1-q)⟨x˙(t),x(t)-x∗⟩+t2b(t)-β(q+2-α)t⟨∇Φλ(t)(x(t)),x˙(t)⟩-λ˙(t)2‖∇Φλ(t)(x(t))‖2+2tε(t)+t2ε˙(t)2‖x(t)‖2+〈q(x(t)-x∗)+tx˙(t)+β∇Φλ(t)(x(t)),(q+1)x˙(t)+β∇Φλ(t)(x(t))+tx¨(t)+βddt∇Φλ(t)(x(t))〉.

Define v(t)=q(x(t)-x∗)+tx˙(t)+β∇Φλ(t)(x(t)). Using (1) to replace x¨(t)+βddt∇Φλ(t)(x(t)) we obtain⟨v(t),v˙(t)⟩=〈q(x(t)-x∗)+tx˙(t)+β∇Φλ(t)(x(t)),(q+1-α)x˙(t)+β-tb(t)∇Φλ(t)(x(t))-tε(t)x(t)〉=q(q+1-α)x(t)-x∗,x˙(t)+(q+1-α)t‖x˙(t)‖2+β(q+2-α)t-t2b(t)∇Φλ(t)(x(t)),x˙(t)+β2t-βt2b(t)‖∇Φλ(t)(x(t))‖2-t2ε(t)⟨x(t),x˙(t)⟩-βt2ε(t)⟨x(t),∇Φλ(t)(x(t))⟩-qtb(t)-βt∇Φλ(t)(x(t))+ε(t)x(t),x(t)-x∗.

By (14) one has b(t)-βt>0 for all t≥t0, and thus for a strongly convex function φt(x)=b(t)-βtΦλ(t)(x)+ε(t)2‖x‖2 we haveφt(x∗)-φt(x)≥∇φt(x),x∗-x+ε(t)2‖x∗-x‖2

or-qtb(t)-βt∇Φλ(t)(x(t))+ε(t)x(t),x(t)-x∗≤-qtb(t)-βtΦλ(t)(x(t))-Φ∗-qtε(t)2‖x(t)‖2-qtε(t)2‖x(t)-x∗‖2+qtε(t)2‖x∗‖2.

Therefore, for every t≥t0E˙q(t)≤(2-q)tb(t)+t2b˙(t)-β(2-α)Φλ(t)(x(t))-Φ∗+(q+1-α)t‖x˙(t)‖2-t2b(t)-β(q+2-α)tλ˙(t)2-β2t+βt2b(t)‖∇Φλ(t)(x(t))‖2+(2-q)tε(t)+t2ε˙(t)2‖x(t)‖2-qtε(t)2‖x(t)-x∗‖2+qtε(t)2‖x∗‖2-βt2ε(t)⟨x(t),∇Φλ(t)(x(t))⟩.

Notice that for a≥1-βt2ε(t)⟨x(t),∇Φλ(t)(x(t))⟩≤βt2a‖∇Φλ(t)(x(t))‖2+aβt2ε2(t)4‖x(t)‖2,

which leads to16 E˙q(t)≤(2-q)tb(t)+t2b˙(t)-β(2-α)Φλ(t)(x(t))-Φ∗+(q+1-α)t‖x˙(t)‖2-(t2b(t)-β(q+2-α)t)λ˙(t)2-β2t+βt2b(t)-1a‖∇Φλ(t)(x(t))‖2+2(2-q)tε(t)+2t2ε˙(t)+aβt2ε2(t)4‖x(t)‖2-qtε(t)2‖x(t)-x∗‖2+qtε(t)2‖x∗‖2

for every t≥t0. Note that b(t)-1a>0 for all t≥t0. Then, due to the properties of b, there exists t∗≥t0 such that t2b(t)-β(q+2-α)t≥0 for all t≥t∗ and all q∈(2,α-1]. Therefore, since λ˙(t)≥0 for all t≥t0, there exists t∗∗, namely, t∗∗=maxt∗,βb(t0)-1a, such that(t2b(t)-β(q+2-α)t)λ˙(t)2-β2t+βt2b(t)-1a≥0for allt≥t∗∗.

Consider now two cases with t≥t∗∗. First, take q=α-1 to obtain from (16)17 E˙α-1(t)≤(3-α)tb(t)+t2b˙(t)-β(2-α)Φλ(t)(x(t))-Φ∗-(t2b(t)-βt)λ˙(t)2-β2t+βt2b(t)-1a‖∇Φλ(t)(x(t))‖2+2(3-α)tε(t)+2t2ε˙(t)+aβt2ε2(t)4‖x(t)‖2-(α-1)tε(t)2‖x(t)-x∗‖2+(α-1)tε(t)2‖x∗‖2

for every t≥t0. Under the assumptions (11) and (12) we conclude starting from t∗∗ that18 E˙α-1(t)≤(α-1)tε(t)2‖x∗‖2.

Under the assumption () using the fact that t↦Eα-1(t) is bounded from below we deduce the existence of the limit limt→+∞Eα-1(t) due to the Lemma A.1 and, therefore, t↦Eα-1(t) is bounded, which leads toΦλ(t)(x(t))-Φ∗=O1t2b(t),ast→+∞.

From the boundedness of t↦‖(α-1)(x(t)-x∗)+t(x˙(t)+β∇Φλ(t)(x(t))‖2 we obtain‖x˙(t)+β∇Φλ(t)(x(t))‖=O1t,ast→+∞,

using the following inequality, which is true for every t≥t0t2‖x˙(t)+β∇Φλ(t)(x(t))‖2≤2‖(α-1)(x(t)-x∗)+t(x˙(t)+β∇Φλ(t)(x(t))‖2+2(α-1)2‖x(t)-x∗‖2.

Moreover, integrating (17) one may obtain the integrability of tε(t)‖x(t)-x∗‖2 as well as the other terms in (17). Consider now q=α-1-δ, where δ is defined by (15). Thus, (16) becomes19 E˙α-1-δ(t)≤(3-α+δ)tb(t)+t2b˙(t)-β(2-α)Φλ(t)(x(t))-Φ∗-δt‖x˙(t)‖2-(t2b(t)-β(1-δ)t)λ˙(t)2-β2t+βt2b(t)-1a‖∇Φλ(t)(x(t))‖2+2(3-α+δ)tε(t)+2t2ε˙(t)+aβt2ε2(t)4‖x(t)‖2-(α-1-δ)tε(t)2‖x(t)-x∗‖2+(α-1-δ)tε(t)2‖x∗‖2.

Under the assumptions (12) and (15) we deduce E˙α-1-δ(t)≤(α-1-δ)tε(t)2‖x∗‖2 starting from t∗∗. Repeating the same argument we derive that t↦Eα-1-δ(t) is bounded. The function t↦‖x(t)-x∗‖ is also bounded and so is the trajectory x. Integrating (19) one may additionally obtain the integrability of t‖x˙(t)‖2. From the integrability of (α-3)tb(t)-t2b˙(t)-β(α-2)Φλ(t)(x(t))-Φ∗ and (15) we deduce∫t0+∞tb(t)Φλ(s)(x(s))-Φ∗<+∞.

□

The next theorem shows that we can actually improve the rates of convergence of the function values in case α>3.

Theorem 5

Assume that α>3 and (12), (13),(14) and (15) hold. Then20 tb(t)-βt∇Φλ(t)(x(t)),x(t)-x∗∈L1([t0,+∞),R).

In addition, limt→+∞ψ(t)=0, where for 2≤q≤α-1ψ(t)=(t2b(t)-β(q+2-α)t)Φλ(t)(x(t))-Φ∗+t2ε(t)2‖x(t)‖2+t22‖x˙(t)+β∇Φλ(t)(x(t))‖2,

which in particular means21 Φλ(t)(x(t))-Φ∗=o1t2b(t)ast→+∞,‖x˙(t)+β∇Φλ(t)(x(t))‖=o1tast→+∞

and moreover,Φ(proxλ(t)Φ(x(t)))-Φ∗=o1t2b(t)ast→+∞,‖proxλ(t)Φ(x(t))-x(t)‖=oλ(t)tb(t)ast→+∞

and‖∇Φλ(t)(x(t))‖=o1tb(t)λ(t)ast→+∞.

Proof

(i) Let us first prove an auxiliary estimate (20), which will allow us to obtain the rest of the desired results. We return toE˙q(t)≤2tb(t)+t2b˙(t)-β(q+2-α)Φλ(t)(x(t))-Φ∗+(q+1-α)t‖x˙(t)‖2-t2b(t)-β(q+2-α)tλ˙(t)2-β2t+βt2b(t)-1a‖∇Φλ(t)(x(t))‖2+4tε(t)+2t2ε˙(t)+aβt2ε2(t)4‖x(t)‖2-qtb(t)-βt∇Φλ(t)(x(t))+ε(t)x(t),x(t)-x∗.

Under condition (12) we deduce starting from t∗∗E˙q(t)≤2tb(t)+t2b˙(t)-β(q+2-α)Φλ(t)(x(t))-Φ∗+tε(t)‖x(t)‖2-qtb(t)-βt∇Φλ(t)(x(t))+ε(t)x(t),x(t)-x∗.

Integrating the last inequality on [t0,t] we obtain22 ∫t0tqsb(s)-βs∇Φλ(s)(x(s)),x(s)-x∗ds≤Eq(t0)-Eq(t)+∫t0tsε(s)‖x(s)‖2ds+∫t0t2sb(s)+s2b˙(s)-β(q+2-α)Φλ(s)(x(s))-Φ∗ds-∫t0tqs⟨ε(s)x(s),x(s)-x∗⟩.

Since the gradient ∇Φλ is monotone, we know that ∇Φλ(t)(x(t)),x(t)-x∗≥0. Moreover,23 -qt⟨ε(t)x(t),x(t)-x∗⟩≤qtε(t)2‖x(t)‖2+‖x(t)-x∗‖2.

Notice that by (15) we have(α-3-δ)tb(t)-t2b˙(t)+β(2-α)>0

or(α-3-δ)tb(t)-t2b˙(t)>β(α-2).

Obviuosly,(α-3-δ)tb(t)-t2b˙(t)β(α-2)>β(α-2-q)for everyq∈(2,α-1).

Introducing δ1=α-1-δ>0 (by the choice of δ) we obtain(δ1-2)tb(t)-t2b˙(t)>-β(q+2-α)

or2tb(t)+t2b˙(t)-β(q+2-α)<δ1tb(t).

From Theorem 4 we know that tb(t)Φλ(s)(x(s))-Φ∗ is integrable and therefore so is 2tb(t)+t2b˙(t)-β(q+2-α)Φλ(s)(x(s))-Φ∗. Since the function t↦Eq(t) is bounded and the rest of the right hand side of (22) belongs to L1([t0,+∞),R) by Theorem 4 and (23), we conclude with (20) due to (14).

(ii) In order to derive the convergence rates for the quantities of our interest we require some additional results. Our nearest goal is to establish the existence of the limitslimt→+∞‖x(t)-x∗‖andlimt→+∞tx˙(t)+β∇Φλ(t)(x(t)),x(t)-x∗.

Consider (as was done in [23, 24]) for two different q1,q2∈(2,α-1) and for every t≥t0 the differenceEq1(t)-Eq2(t)=(t2b(t)-β(q1+2-α)t)Φλ(t)(x(t))-Φ∗+t2ε(t)2‖x(t)‖2+12‖q1(x(t)-x∗)+t(x˙(t)+β∇Φλ(t)(x(t))‖2+q1(α-1-q1)2‖x(t)-x∗‖2-(t2b(t)-β(q2+2-α)t)Φλ(t)(x(t))-Φ∗-t2ε(t)2‖x(t)‖2-12‖q2(x(t)-x∗)+t(x˙(t)+β∇Φλ(t)(x(t))‖2-q2(α-1-q2)2‖x(t)-x∗‖2=(q1-q2)(-βtΦλ(t)(x(t))-Φ∗+tx˙(t)+β∇Φλ(t)(x(t)),x(t)-x∗+α-12‖x(t)-x∗‖2).

As we have established earlier in Theorem 4 the limits of Eq1(t)-Eq2(t) and tΦλ(t)(x(t))-Φ∗ exists (the latter is actually zero). Therefore, the limitlimt→+∞tx˙(t)+β∇Φλ(t)(x(t)),x(t)-x∗+α-12‖x(t)-x∗‖2also exists.

Let us introduce for every t≥t0 two auxiliary functionsk(t)=tx˙(t)+β∇Φλ(t)(x(t)),x(t)-x∗+α-12‖x(t)-x∗‖2

andr(t)=12‖x(t)-x∗‖2+β∫t0t∇Φλ(s)(x(s)),x(s)-x∗ds.

Noticing thatr˙(t)=⟨x(t)-x∗,x˙(t)⟩+β∇Φλ(t)(x(t)),x(t)-x∗

we may write for every t≥t0(α-1)r(t)+tr˙(t)=k(t)+β(α-1)∫t0t∇Φλ(s)(x(s)),x(s)-x∗ds.

From the fact that limt→+∞k(t) exists using (20) we obtain that limt→+∞(α-1)r(t)+tr˙(t) also exists. Applying Lemma A.2 we deduce the existence of the limit limt→+∞r(t). Using (20) again we obtain the existence of the limits limt→+∞‖x(t)-x∗‖ and limt→+∞tx˙(t)+β∇Φλ(t)(x(t)),x(t)-x∗.

(iii) Finally, we are in position to prove (21) and the rest of the convergence rates. The key idea is to show that the limitlimt→+∞(t2b(t)-β(q+2-α)t)Φλ(t)(x(t))-Φ∗+t2ε(t)2‖x(t)‖2+t22‖x˙(t)+β∇Φλ(t)(x(t))‖2

exists and is actually zero. Let us return to the definition of our energy functional and rewrite it asEq(t)=(t2b(t)-β(q+2-α)t)Φλ(t)(x(t))-Φ∗+t2ε(t)2‖x(t)‖2+t22‖x˙(t)+β∇Φλ(t)(x(t)‖2+qtx˙(t)+β∇Φλ(t)(x(t)),x(t)-x∗+q(α-1)2‖x(t)-x∗‖2.

Since the limitslimt→+∞Eq(t)andlimt→+∞qtx˙(t)+β∇Φλ(t)(x(t),x(t)-x∗+q(α-1)2‖x(t)-x∗‖2exist,

it follows thatlimt→+∞(t2b(t)-β(q+2-α)t)Φλ(t)(x(t))-Φ∗+t2ε(t)2‖x(t)‖2+t22‖x˙(t)+β∇Φλ(t)(x(t))‖2

exists as well. Denoteψ(t)=(t2b(t)-β(q+2-α)t)Φλ(t)(x(t))-Φ∗+t2ε(t)2‖x(t)‖2+t22‖x˙(t)+β∇Φλ(t)(x(t))‖2

and consider24 0≤ψ(t)t≤2tb(t)Φλ(t)(x(t))-Φ∗+tε(t)2‖x(t)‖2+t2‖x˙(t)+β∇Φλ(t)(x(t))‖2.

Let us show that the right hand side of (24) is integrable. Indeed, the first term is integrable by Theorem 4. As we have also established in Theorem 4, starting from t∗∗(tb(t)-β)λ˙(t)2+βtb(t)-1a≥β2,

where a≥1. Then, by (14) and λ˙(t)≥0 for all t≥t0, we deduce that there exists t1≥t∗∗ such that for all t≥t1(tb(t)-β)λ˙(t)2+βtb(t)-1a≥3β22

ort2b(t)λ˙(t)2+β1-1ab(t)≥3β+λ˙(t)βt2

ort2b(t)-βtλ˙(t)2-β2t+βt2b(t)-1a≥β2t2,

So, by Theorem 4 the right hand side of (24) belongs to L1([t1,+∞),R). Therefore, ψ(t)t also belongs to L1([t1,+∞),R) and since the limit limt→+∞ψ(t) exists we deduce that it should be actually zero, which gives us (21). To complete the proof notice that by the definition of the proximal mapping, we haveΦλ(t)(x(t))-Φ∗=Φ(proxλ(t)Φ(x(t)))-Φ∗+12λ(t)‖proxλ(t)Φ(x(t))-x(t)‖2∀t≥t0.

The conclusion follows immediately from (2) and (21). □

Strong Convergence of the Trajectories

In this chapter we will establish the strong convergence of the trajectories to the minimal norm element of argminΦ.

In order to do so, we will need to modify assumption (13) from the previous chapter:

Before moving to the main point of the section, let us prove an auxiliary result first.

Theorem 6

Suppose that α>3, the function λ is bounded for all t≥t0 and (11), (12), (14) and (25) hold. Thenlimt→+∞‖proxλ(t)Φ(x(t))-x(t)‖=0

andlimt→+∞Φproxλ(t)Φ(x(t))-Φ∗=0.

Proof

Let us return to (18):E˙α-1(t)≤(α-1)tε(t)2‖x∗‖2.

Let us integrate the last inequality on [T, t]Eα-1(t)≤Eα-1(T)+(α-1)‖x∗‖22∫Ttsε(s)ds.

On the other hand, for every t≥t0Eα-1(t)≥(t2b(t)-βt)Φλ(t)(x(t))-Φ∗.

Thus,Φλ(t)(x(t))-Φ∗≤Eα-1(T)t2b(t)-βt+(α-1)‖x∗‖22(t2b(t)-βt)∫Ttsε(s)ds.

We deduce due to (25) and the Lemma A.3 thatlimt→+∞1t2b(t)∫Tts2b(s)ε(s)sb(s)ds=0.

Therefore,limt→+∞(α-1)‖x∗‖22(t2b(t)-βt)∫Ttsε(s)ds=0

and clearlylimt→+∞Eα-1(T)t2b(t)-βt=0.

Thus, we establishlimt→+∞Φλ(t)(x(t))-Φ∗=0.

By the definition of the proximal mappingΦλ(t)(x(t))-Φ∗=Φproxλ(t)Φ(x(t))-Φ∗+12λ(t)‖proxλ(t)Φ(x(t))-x(t)‖2∀t≥t0.

Using the fact that λ is bounded for all t≥t0 we deducelimt→+∞‖proxλ(t)Φ(x(t))-x(t)‖=0

andlimt→+∞Φproxλ(t)Φ(x(t))-Φ∗=0.

□

For the remaining part of this section we will use a different energy functional. Inspired by [23] we introduce the following functional, which we will heavily rely on throughout this section26 Ep,q(t)=tp+1tb(t)+β(α-p-q-2)Φλ(t)(x(t))-Φ∗+ε(t)tp+22‖x(t)‖2-‖x∗‖2+tp2v(t)2,

where v(t)=q(x(t)-x∗)+t(x˙(t)+β∇Φλ(t)(x(t)) and p,q≥0.

The proof of the following theorem draws inspiration from [10, 15, 23].

Theorem 7

Suppose that λ is bounded for all t≥t0, α>3, b(t0)≥12+βt0 and (11), (12) and (25) are fulfilled. Suppose additionally that for all t≥t027 α3-1tb(t)-t2b˙(t)+αβ3≥0

and moreover that for all t≥t028 2α(α-3)-9t2ε(t)+6αβ≤0,

29 18βt+9βλ˙(t)-9tb(t)λ˙(t)+2β+3(α+3)β2+α2β≤0

and30 limt→+∞βtα3+1ε(t)∫t0tsα3+1ε2(s)ds=0.

If x:[t0,+∞)↦H is a solution to (1) and the trajectory x(t) stays either inside or outside the ball B(0,‖x∗‖), then x(t) converges to minimal norm solution x∗=projargminΦ(0), as t→+∞. Otherwise, lim inft→+∞‖x(t)-x∗‖=0.

Proof

As in [23] we will consider several cases with respect to the trajectory x staying either inside or outside the ball B0,‖x∗‖.

Case I.

Assume that the trajectory x stays in the complement of the ball B for all t≥t0. This means nothing but ‖x(t)‖≥‖x∗‖ for every t≥t0.

(i) Our nearest goal is to obtain the upper bound for the derivative of Ep,q. In order to do so, let us evaluate its time derivative for every t≥t0 first.31 ddtEp,q(t)=tp(p+2)tb(t)+t2b˙(t)+(p+1)β(α-p-q-2)Φλ(t)(x(t))-Φ∗+tp+1tb(t)+β(α-p-q-2)∇Φλ(t)(x(t)),x˙(t)-λ˙(t)2‖∇Φλ(t)(x(t))‖2+(p+2)tp+1ε(t)+tp+2ε˙(t)2‖x(t)‖2-‖x∗‖2+tp+2ε(t)⟨x˙(t),x(t)⟩+ptp-12‖q(x(t)-x∗)+t(x˙(t)+β∇Φλ(t)(x(t))‖2+tp⟨v˙(t),v(t)⟩.

Consider for every t≥t0 the inner product ⟨v˙(t),v(t)⟩:〈(q+1)x˙(t)+β∇Φλ(t)(x(t))+tx¨(t)+βddt∇Φλ(t)(x(t)),q(x(t)-x∗)+t(x˙(t)+β∇Φλ(t)(x(t))〉=〈(q+1-α)x˙(t)+β∇Φλ(t)(x(t))-t(b(t)∇Φλ(t)(x(t))+ε(t)x(t)),q(x(t)-x∗)+t(x˙(t)+β∇Φλ(t)(x(t))〉=q(q+1-α)⟨x˙(t),x(t)-x∗⟩+(q+1-α)t‖x˙(t)‖2+β∇Φλ(t)(x(t)),x˙(t)+βq⟨∇Φλ(t)(x(t)),x(t)-x∗⟩+βt⟨∇Φλ(t)(x(t)),x˙(t)⟩+β2t‖∇Φλ(t)(x(t))‖2-qtb(t)∇Φλ(t)(x(t))+ε(t)x(t),x(t)-x∗-t2b(t)∇Φλ(t)(x(t))+ε(t)x(t),x˙(t)-βt2〈b(t)∇Φλ(t)(x(t))+ε(t)x(t),∇Φλ(t)(x(t))〉,

where above we used (1). Consider now for every t≥t0,‖q(x(t)-x∗)+t(x˙(t)+β∇Φλ(t)(x(t)))‖2=q2‖x(t)-x∗‖2+2qt⟨x˙(t),x(t)-x∗⟩+2qβt⟨∇Φλ(t)(x(t)),x(t)-x∗⟩+t2‖x˙(t)‖2+2βt2⟨∇Φλ(t)(x(t)),x˙(t)⟩+β2t2‖∇Φλ(t)(x(t))‖2.

The two estimates that we made above lead to (31) becomingddtEp,q(t)=tp(p+2)tb(t)+t2b˙(t)+(p+1)β(α-p-q-2)Φλ(t)(x(t))-Φ∗+(p+2)tp+1ε(t)+tp+2ε˙(t)2‖x(t)‖2-‖x∗‖2+pq2tp-12‖x(t)-x∗‖2+(p+2)β2tp+12‖∇Φλ(t)(x(t))‖2+q+1-α+p2tp+1‖x˙(t)‖2+q(q+1-α+p)tp〈x˙(t),x(t)-x∗〉+qβ(p+1)tp〈∇Φλ(t)(x(t)),x(t)-x∗〉-qtp+1〈b(t)∇Φλ(t)(x(t))+ε(t)x(t),x(t)-x∗〉-βtp+2〈b(t)∇Φλ(t)(x(t))+ε(t)x(t),∇Φλ(t)(x(t))〉-λ˙(t)tp+1tb(t)+β(α-p-q-2)2‖∇Φλ(t)(x(t))‖2.

Let us apply the gradient inequality to the strongly convex function x↦b(t)Φλ(t)(x)+ε(t)‖x‖22:-〈b(t)∇Φλ(t)(x(t))+ε(t)x(t),x(t)-x∗〉+ε(t)‖x(t)-x∗‖22≤b(t)Φ∗+ε(t)‖x∗‖22-b(t)Φλ(t)(x(t))+ε(t)‖x(t)‖22

and thus-qtp+1〈b(t)∇Φλ(t)(x(t))+ε(t)x(t),x(t)-x∗〉≤-qtp+1b(t)Φλ(t)(x(t))-Φ∗-qtp+1ε(t)2‖x(t)‖2-‖x∗‖2-qtp+1ε(t)‖x(t)-x∗‖22

for every t≥t0. So, noticing that-βtp+2〈b(t)∇Φλ(t)(x(t))+ε(t)x(t),∇Φλ(t)(x(t))〉=-βtp+2b(t)‖∇Φλ(t)(x(t))‖2-βtp+2ε(t)〈x(t),∇Φλ(t)(x(t))〉

we deduceddtEp,q(t)≤tp(p+2-q)tb(t)+t2b˙(t)+(p+1)β(α-p-q-2)Φλ(t)(x(t))-Φ∗+(p+2-q)tp+1ε(t)+tp+2ε˙(t)2‖x(t)‖2-‖x∗‖2+pq2tp-12-qtp+1ε(t)2‖x(t)-x∗‖2+(p+2)β2tp+1-2βtp+2b(t)-λ˙(t)tp+1tb(t)+β(α-p-q-2)2‖∇Φλ(t)(x(t))‖2+q+1-α+p2tp+1‖x˙(t)‖2+q(q+1-α+p)tp〈x˙(t),x(t)-x∗〉+qβ(p+1)tp〈∇Φλ(t)(x(t)),x(t)-x∗〉-βtp+2ε(t)〈x(t),∇Φλ(t)(x(t))〉.

In order to proceed further we will need the following estimates:qβ(p+1)tp〈∇Φλ(t)(x(t)),x(t)-x∗〉≤qβ(p+1)tp+14c2‖∇Φλ(t)(x(t))‖2+qβ(p+1)c2tp-1‖x(t)-x∗‖2

and-βtp+2ε(t)〈x(t),∇Φλ(t)(x(t))〉≤βtp+2a‖∇Φλ(t)(x(t))‖2+aβtp+2ε2(t)4‖x(t)‖2

for every t≥t0, some c≥1 and a≥1. Thus,ddtEp,q(t)≤tp(p+2-q)tb(t)+t2b˙(t)+(p+1)β(α-p-q-2)Φλ(t)(x(t))-Φ∗+(p+2-q)tp+1ε(t)+tp+2ε˙(t)2+aβtp+2ε2(t)4‖x(t)‖2+pq2tp-12-qtp+1ε(t)2+qβ(p+1)c2tp-1‖x(t)-x∗‖2+((p+2)β2tp+1-2βtp+2b(t)-λ˙(t)tp+1tb(t)+β(α-p-q-2)2+qβ(p+1)tp+14c2+βtp+2a)·‖∇Φλ(t)(x(t))‖2+q+1-α+p2tp+1‖x˙(t)‖2+q(q+1-α+p)tp〈x˙(t),x(t)-x∗〉-(p+2-q)tp+1ε(t)+tp+2ε˙(t)2‖x∗‖2.

Let us fixq=2α3andp=α-33.

First of all, due to this choiceq+1-α+p=0

and thus we get rid of the term 〈x˙(t),x(t)-x∗〉. Secondly,32 q+1-α+p2=-p2≤0.

Then33 p+2-q=1-α3≤0.

So,ddtEp,q(t)≤tp(p+2-q)tb(t)+t2b˙(t)+(p+1)β(α-p-q-2)Φλ(t)(x(t))-Φ∗+(p+2-q)tp+1ε(t)+tp+2ε˙(t)2+aβtp+2ε2(t)4‖x(t)‖2+pq2tp-12-qtp+1ε(t)2+qβ(p+1)c2tp-1‖x(t)-x∗‖2+((p+2)β2tp+1-2βtp+2b(t)-λ˙(t)tp+1tb(t)+β(α-p-q-2)2+qβ(p+1)tp+14c2+βtp+2a)·‖∇Φλ(t)(x(t))‖2+q+1-α+p2tp+1‖x˙(t)‖2-(p+2-q)tp+1ε(t)+tp+2ε˙(t)2‖x∗‖2.

Obviously, for t large enough, say, t≥t2≥t0 the following expression is non-positive due to (27) and p+1=α3>0 and α-p-q-2=-1(p+2-q)tb(t)+t2b˙(t)+(p+1)β(α-p-q-2)=1-α3tb(t)+t2b˙(t)-αβ3≤0.

Moreover, from (28) it follows that for c=1pq2tp-12-qtp+1ε(t)2+qβ(p+1)c2tp-1=αtα-63272α(α-3)-9t2ε(t)+6αβ≤0

for all t≥t0. Furthermore,(p+2-q)tp+1ε(t)+tp+2ε˙(t)2+aβtp+2ε2(t)4‖x(t)‖2-(p+2-q)tp+1ε(t)+tp+2ε˙(t)2‖x∗‖2=(p+2-q)tp+1ε(t)+tp+2ε˙(t)2+aβtp+2ε2(t)4‖x(t)‖2-‖x∗‖2+aβtp+2ε2(t)4‖x∗‖2.

So, under the assumption (12) and the fact that ‖x(t)‖≥‖x∗‖ for all t≥t0 we deduce due to (33)(p+2-q)tp+1ε(t)+tp+2ε˙(t)2+aβtp+2ε2(t)4‖x(t)‖2-‖x∗‖2≤0.

Thus, under the assumptions (12), (27), (28) and (29) (the latest leads to the non-positivity of the coefficient of ‖∇Φλ(x)‖2) we conclude due to (32) that for every t≥t234 ddtEp,q(t)≤aβtα3+1ε2(t)4‖x∗‖2.

(ii) Let us obtain now the lower bound for Ep,q. Notice that for p=α-33 and q=2α3 we have α-p-q=1 and35 Ep,q(t)≥tp+1tb(t)+β(α-p-q-2)Φλ(t)(x(t))-Φ∗+ε(t)tp+22‖x(t)‖2-‖x∗‖2=tp+1tb(t)-βΦλ(t)(x(t))-Φ∗+ε(t)tp+22‖x(t)‖2-‖x∗‖2≥tp+22Φλ(t)(x(t))-Φ∗+ε(t)tp+22‖x(t)‖2-‖x∗‖2,

since tb(t)-β≥t2 for every t≥t0 by b(t0)≥12+βt0 and b being non-decreasing. On the other hand, applying the gradient inequality to the strongly convex function φε(t),λ(t)(x)=Φλ(t)(x)2+ε(t)2‖x‖2 we deduce for xε(t),λ(t)=argminHφε(t),λ(t)(x)φε(t),λ(t)(x)-φε(t),λ(t)(xε(t),λ(t))≥ε(t)2‖x-xε(t),λ(t)‖2for everyx∈H.

By the definition of φε(t),λ(t)(x) we deduceφε(t),λ(t)(xε(t),λ(t))-φε(t),λ(t)(x∗)=12Φλ(t)(xε(t),λ(t))-Φ∗+ε(t)2‖xε(t),λ(t)‖2-‖x∗‖2≥ε(t)2‖xε(t),λ(t)‖2-‖x∗‖2.

We may now add the last two inequalities to obtain36 φε(t),λ(t)(x)-φε(t),λ(t)(x∗)≥ε(t)2‖x-xε(t),λ(t)‖2+‖xε(t),λ(t)‖2-‖x∗‖2for everyx∈H.

Plugging (36) into (35) we conclude that for every t≥t237 Ep,q(t)≥tp+2ε(t)2‖x(t)-xε(t),λ(t)‖2+‖xε(t),λ(t)‖2-‖x∗‖2.

(iii) Finally, using the lower and upper bounds for Ep,q we can prove the strong convergence of the trajectories to a minimal norm solution. Integrating (34) on [t2,t] we obtainEp,q(t)≤Ep,q(t2)+aβ‖x∗‖24∫t2tsα3+1ε2(s)ds

and using (37) we deduce for every t≥t2‖x(t)-xε(t),λ(t)‖2≤‖x∗‖2-‖xε(t),λ(t)‖2+2Ep,q(t2)tα3+1ε(t)+aβ‖x∗‖22tα3+1ε(t)∫t2tsα3+1ε2(s)ds.

Note that due to (28)t2ε(t)≥2α(α-3)+6αβ9=C^≥0

andtα3+1ε(t)=t2ε(t)tα3-1≥C^tα3-1.

Since α>3 we deducelimt→+∞tα3+1ε(t)=+∞

and thuslimt→+∞2Ep,q(t2)tα3+1ε(t)=0.

Finally, by (9) and (30) we concludelimt→+∞x(t)=x∗.

Case II.

Assume now the opposite to the first case, namely, ‖x(t)‖<‖x∗‖ for every t≥t0. According to Theorem 6limt→+∞‖proxλ(t)Φ(x(t))-x(t)‖=0

andlimt→+∞Φproxλ(t)Φ(x(t))-Φ∗=0.

Denote ξ(t)=proxλ(t)Φ(x(t)). Considering a sequence {tk}k∈N such that {x(tk)}k∈N converges weakly to an element x^∈H as k→∞, we notice that {ξ(tk)}k∈N converges weakly to x^ as k→∞. Now, the function Φ being convex and lower semicontinuous in the weak topology, allows us to writeΦ(x^)≤lim infk→∞Φ(ξ(tk))=limt→+∞Φ(ξ(t))=Φ∗

and hence, x^∈argminΦ. The norm is weakly semicontinuous, so‖x^‖≤lim infk→∞‖ξ(tk)‖≤‖x∗‖,

which means that x^=x∗ by the uniqueness of the element of the minimum norm in argminΦλ. Therefore, the trajectory x converges weakly to x∗ and‖x∗‖≤lim inft→+∞‖x(t)‖≤lim supt→+∞‖x(t)‖≤‖x∗‖

and thuslimt→+∞‖x(t)‖=‖x∗‖.

From this and the weak convergence of the trajectory x follows the strong one: limt→+∞x(t)=x∗.

Case III.

Assume that for t≥t0 the trajectory x finds itself both inside and outside the ball B(0,‖x∗‖). Since x is continuous, there exists a sequence {tn}n∈N⊆[t0,+∞) such that tn→∞ as n→∞ and ‖x(tn)‖=‖x∗‖ for every n∈N. Consider again a weak sequential cluster point x^ of the sequence {x(tn)}n∈N. By repeating the same argument as in the previous case we deduce the weak convergence of {x(tn)}n∈N to x∗, as n→∞. Since ‖x(tn)‖→‖x∗‖, as n→∞, we obtain that ‖x(tn)-x∗‖→0, as n→∞, which means lim inft→+∞‖x(t)-x∗‖=0. □

Remark 1

In this section the condition b˙(t)≥0 for all t≥t0 is not necessary. Our conjecture is that we can weaken the setting by omitting this condition and thus widen the range for b, including the functions that decay not faster than 1t2 for the polynomial choice of parameters.

Remark 2

There is no setting which guarantees both fast rates for the values and strong convergence of the trajectories. One of the future goal would be to develop a new approach (based on [6]), which would help us deduce these two results simultaneously.

Strong Convergence of the Tajectories in Cse α=3

Throughout this section we no longer require that b is non-decreasing. In this case the analogue of Theorem 6 looks as follows.

Theorem 8

Suppose that for all t≥t0 the function λ is bounded, b(t)≡b>0 is a constant function and (12) and (14) hold. Suppose additionally that (25) holds for constant b, namely∫t0+∞ε(t)tdt<+∞.

Thenlimt→+∞‖proxλ(t)Φ(x(t))-x(t)‖=0

andlimt→+∞Φproxλ(t)Φ(x(t))-Φ∗=0.

Proof

In this case the energy functional becomesE2(t)=(bt2-βt)Φλ(t)(x(t))-Φ∗+t2ε(t)2‖x(t)+12‖2(x(t)-x∗)+tx˙(t)+β∇Φλ(t)(x(t))‖2.

Relation (16) thus becomes for all t≥t0E˙2(t)≤βΦλ(t)(x(t))-Φ∗-(bt2-βt)λ˙(t)2-β2t+βt2b-1a‖∇Φλ(t)(x(t))‖2+2t2ε˙(t)+aβt2ε2(t)4‖x(t)‖2-tε(t)‖x(t)-x∗‖2+tε(t)‖x∗‖2.

Thus, repeating the same arguments as in Theorem 4 we obtainE˙2(t)≤βΦλ(t)(x(t))-Φ∗+tε(t)‖x∗‖2.

Let us multiply this expression with t(bt-β) to obtaint(bt-β)E˙2(t)≤βt(bt-β)Φλ(t)(x(t))-Φ∗+t2(bt-β)ε(t)‖x∗‖2≤βE2(t)+t2(bt-β)ε(t)‖x∗‖2.

Now, we will divide by (bt-β)2 to concludet(bt-β)E˙2(t)≤β(bt-β)2E2(t)+t2(bt-β)ε(t)‖x∗‖2

orddttbt-βE2(t)≤t2(bt-β)ε(t)‖x∗‖2.

Integrating the last inequality on [T, t], where T≥t0, we deducetbt-βE2(t)≤TbT-βE2(T)+‖x∗‖2∫Tts2(bs-β)ε(s)ds.

By the definition of E2 we knowE2(t)≥(bt2-βt)Φλ(t)(x(t))-Φ∗.

Combining these two inequalities, we deduceΦλ(t)(x(t))-Φ∗≤Tt2(bT-β)E2(T)+‖x∗‖2t2∫Tts2(bs-β)ε(s)ds.

Now,limt→+∞Tt2(bT-β)E2(T)=0.

Applying Lemma A.3 we deduce due to (25)limt→+∞bt-βt3∫Tts3(bs-β)ε(s)sds=0

and thuslimt→+∞‖x∗‖2t2∫Tts2(bs-β)ε(s)ds=0.

Therefore, we establishlimt→+∞Φλ(t)(x(t))-Φ∗=0.

Again, by the definition of the proximal mappingΦλ(t)(x(t))-Φ∗=Φproxλ(t)Φ(x(t))-Φ∗+12λ(t)‖proxλ(t)Φ(x(t))-x(t)‖2∀t≥t0.

Using the fact that λ is bounded for all t≥t0 we deducelimt→+∞‖proxλ(t)Φ(x(t))-x(t)‖=0

andlimt→+∞Φproxλ(t)Φ(x(t))-Φ∗=0.

□

We are in position now to formulate the analogue of Theorem 7.

Theorem 9

Suppose that λ is bounded for all t≥t0, b(t)≡b≥12+βt0 and (12) and (25) hold. Assume, in addition, that38 limt→+∞t2ε(t)=+∞,

39 2βt+βλ˙(t)-btλ˙(t)+2β+2β2+β≤0for allt≥t0

and40 limt→+∞βt2ε(t)∫t0ts2ε2(s)ds=0.

If x:[t0,+∞)↦H is a solution to (1) and the trajectory x(t) stays either inside or outside the ball B(0,‖x∗‖), then x(t) converges to minimal norm solution x∗=projargminΦ(0), as t→+∞. Otherwise, lim inft→+∞‖x(t)-x∗‖=0.

Proof

The proof goes in line with the one of Theorem 7 by taking α=3, b(t)≡b>0, q=2, p=0 and referring to Theorem 8 instead of Theorem 6 in the second and third cases. □

Analysis of the Conditions

Since all the conditions cannot be satisfied simultaneously, let us treat them separately, namely: In order to obtain the fast convergence rates of the function values we require that for all t≥t0: α>3;

the existence of a≥1 such that 2ε˙(t)≤-aβε2(t),

b(t0)≥βt0andb(t0)>1a;

∫t0+∞tε(t)dt<+∞ and

the existence of 0<δ<α-3 such that (α-3)tb(t)-t2b˙(t)+β(2-α)≥δtb(t).

For the strong convergence of the trajectories we require the following for all t≥t0: α>3;

λ is bounded;

α-33b(t)-tb˙(t)+αβ3≥0;

(α-3)tb(t)-t2b˙(t)+β(2-α)≥0;

the existence of a≥1 such that 2ε˙(t)≤-aβε2(t), b(t0)>1aandb(t0)≥12+βt0;

∫t0+∞ε(t)tb(t)dt<+∞;

2α(α-3)-9t2ε(t)+6αβ≤0;

18βt+9βλ˙(t)-9tb(t)λ˙(t)+2β+3(α+3)β2+α2β≤0;

limt→+∞βtα3+1ε(t)∫t0tsα3+1ε2(s)ds=0.

We will analyse these conditions in details for the polynomial choice of functions b and ε, namely, b(t)=btn and ε(t)=εtd, where b is positive, n≥0 and ε,d>0.

Setting for the Fast Convergence Rates of the Function Values

The set of the conditions becomes for all t≥t0α>3;

there exists a≥1 such that -2dεtd+1≤-aβε2t2d,

b(t0)≥βt0andb(t0)>1a;

∫t0+∞εtd-1dt<+∞ and

there exists 0<δ<α-3 such that (α-3)btn+1-bntn+1+β(2-α)≥δbtn+1.

After some simple algebraic computations one may discover that in order to satisfy all the conditions at the same time it is enough to assumeα-3>n≥0(condition 5)

andd>2andd≥βε2(conditions 2 and 4),

since all the other inequalities could be fulfilled by taking the appropriate t0, namely,t0≥maxβbn+1,β(α-2)b(α-3-n)n+1andt0>1bn.

Setting for the Strong Convergence of the Trajectories

The set of the conditions becomes for all t≥t0α>3;

λ is bounded;

α-33btn+1-bntn+1+αβ3≥0;

(α-3)btn+1-bntn+1+β(2-α)≥0;

there exists a≥1 such that -2dεtd+1≤-aβε2t2d, b(t0)>1aandb(t0)≥12+βt0;

∫t0+∞εbtn+d+1dt<+∞;

2α(α-3)-9εtd-2+6αβ≤0;

18βt+9βλ˙(t)-9btn+1λ˙(t)+2β+3(α+3)β2+α2β≤0;

limt→+∞βεtα3-d+1∫t0tε2sα3-2d+1ds=0.

Again, analysis of the set of conditions leads to the following conclusion:λ is bounded (condition 2) ;

0≤n≤α-33 and α>3 (condition 3) ;

max1,βε2≤d≤2 (conditions 5, 7, 8, 9) .

As before, t0 should be chosen appropriately.

The Case α=3

In this case the following has to be assumed: there exists a≥1 such that for all t≥t0λ(t) is bounded;

2ε˙(t)≤-aβε2(t),b>1a and b≥12+βt0;

∫t0+∞ε(t)tdt<+∞;

limt→+∞t2ε(t)=+∞;

2βt+βλ˙(t)-btλ˙(t)+2β+2β2+β≤0;

limt→+∞βt2ε(t)∫t0ts2ε2(s)ds=0.

Essentially, for the polynomial choice of parameters that means b≥1 andλ(t) is bounded (condition 2) ;

max1,βε2≤d<2 (conditions 4, 5, 6) ,

so with the appropriate choice of t0 the whole set of conditions is fulfilled.

Numerical Examples

The Rates of Convergence of the Moreau Envelope Values

Consider the objective function Φ:R→R, Φ(x)=|x|+x22 and let us plot the values of its Moreau envelope as well as the gradient of its Moreau envelope for different polynomial functions λ, ε and b to illustrate the theoretical results with some numerical examples. We take λ(t)=tl, ε(t)=1td, b(t)=tn with x(t0)=x0=10, x˙(t0)=0, α=10 and t0=1.4.

First, let us take different time scaling parameter b with l=0 and d=3 and see how it affects the behaviour of the system (1) (see Fig. 1).Fig. 1 l=0 and d=3

As expected, the faster b grows, the faster the convergence is.

Consider now different Moreau envelope parameter λ with d=3 and n=0 (see Fig. 2).Fig. 2 d=3 and n=0

Note that the difference in the starting point comes from the fact that t0≠1, and for different exponents l the value t0l is also different. As predicted by theory, a faster growing function λ leads to faster convergence of not only the gradient of Moreau envelope of the objective function Φ, but also of the values of the Moreau envelope themselves.

Varying the Tikhonov function ε for n=0 and l=0 does not affect the system, which is illustrated by the following plot (see Fig. 3).Fig. 3 n=0 and l=0

Strong Convergence of the Trajectories

For a different objective function let us investigate the strong convergence of the trajectories of (1):Φ(x)=|x-1|,x>10,x∈[-1,1]|x+1|,x<-1.

The set argminΦ is nothing but the segment [-1,1] and 0 is its element of minimal norm. Let us fix α=6 and n=0.7. First we take constant lambda (λ(t)=1 for all t≥t0) and plot the behaviour of the trajectories of (1) with and without Tikhonov term (see Fig. 4.Fig. 4 The role of the Tikhonov term

As we see in case there is no Tikhonov regularization the trajectories converge to the minimizer 1 of Φ, but the Tikhonov term actually guarantees the convergence towards the minimal norm solution, which is 0.

Another comparison was made for non-constant lambda: λ(t)=1-1tl for l=1 (for different l’s the picture is the same), illustrating similar behaviour (see Fig. 5).Fig. 5 The role of the Tikhonov term

Finally, for the same choice of λ let us take different Tikhonov terms to figure out how changing them affects the trajectories of (1) (see Fig. 6).Fig. 6 n=0.7 and λ(t)=1-1t

We see, that the faster ε decays, the slower trajectories converge.

Appendix A

Let us state here some auxiliary lemmas which we used in our analysis. For the proof of the following lemma we refer to [3].

Lemma A.1

Suppose that f:[t0,+∞)→R is locally absolutely continuous and bounded from below and there exists g∈L1([t0,+∞),R) such that for almost all t≥t0ddtf(t)≤g(t).

Then there exists limt→+∞f(t)∈R.

For the proof of the next lemma we refer to [19].

Lemma A.2

Let H be a real Hilbert space and x:[t0,+∞)↦H be a continuously differentiable function satisfying x(t)+tαx˙(t)→L as t→+∞, with α>0 and L∈H. Then x(t)→L as t→+∞.

For the proof of the final Lemma we refer to [9].

Lemma A.3

Let δ>0 and f∈L1(δ,+∞),R be a non-negative and continuous function. Let g:[δ,+∞)→[0,+∞) be a non-decreasing function such that limt→+∞g(t)=+∞. Then it holdslimt→+∞1g(t)∫δtg(s)f(s)ds=0.

Acknowledgements

The authors are grateful to two anonymous reviewers for their remarks on this manuscript and for meaningful suggestions, which improved the quality of this paper.

Funding

Open access funding provided by University of Vienna.

Data availability

Data sharing not applicable to this article as no datasets were generated or analysed during the current study.

Declaration

Competing interests

The authors declare that they have no competing interests subject to the topic of this article.

Ernö Robert Csetnek: This work was supported by a Grant of the Ministry of Research (Romania), Innovation and Digitization, CNCS-UEFISCDI, project number PN-III-P1-1.1-TE-2021-0138, within PNCDI III. Mikhail A. Karapetyants: Research supported by the Doctoral Programme Vienna Graduate School on Computational Optimization (VGSCO) which is funded by FWF (Austrian Science Fund), project W 1260.

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Alvarez F Attouch H Bolte J Redont P A second-order gradient-like dissipative dynamical system with Hessian-driven damping J. de Mathématiques Pures et Appliquées 2002 81 8 747 779 10.1016/S0021-7824(01)01253-3
Alvarez, F., Attouch, H., Bolte, J., Redont, P.: A second-order gradient-like dissipative dynamical system with Hessian-driven damping. J. de Mathématiques Pures et Appliquées 81(8), 747–779 (2002)10.1016/S0021-7824(01)01253-3
2. Attouch H Viscosity solutions of minimization problems SIAM J. Optim. 1996 6 3 769 806 10.1137/S1052623493259616
Attouch, H.: Viscosity solutions of minimization problems. SIAM J. Optim. 6(3), 769–806 (1996)10.1137/S1052623493259616
3. Attouch H Abbas B Svaiter BF Newton-like dynamics and forward–backward methods for structured monotone inclusions in Hilbert spaces J. Optim. Theory Appl. 2014 161 2 331 360 10.1007/s10957-013-0414-5
Attouch, H., Abbas, B., Svaiter, B.F.: Newton-like dynamics and forward–backward methods for structured monotone inclusions in Hilbert spaces. J. Optim. Theory Appl. 161(2), 331–360 (2014)10.1007/s10957-013-0414-5
4. Attouch H Balhag A Chbani Z Riahi H Damped inertial dynamics with vanishing Tikhonov regularization: strong asymptotic convergence towards the minimum norm solution J. Differ. Equ. 2022 311 29 58 10.1016/j.jde.2021.12.005
Attouch, H., Balhag, A., Chbani, Z., Riahi, H.: Damped inertial dynamics with vanishing Tikhonov regularization: strong asymptotic convergence towards the minimum norm solution. J. Differ. Equ. 311, 29–58 (2022)10.1016/j.jde.2021.12.005
5. Attouch H Balhag A Chbani Z Riahi H Fast convex optimization via inertial dynamics combining viscous and Hessian-driven damping with time rescaling Evol. Equ. Control Theory 2022 11 2 487 514 10.3934/eect.2021010
Attouch, H., Balhag, A., Chbani, Z., Riahi, H.: Fast convex optimization via inertial dynamics combining viscous and Hessian-driven damping with time rescaling. Evol. Equ. Control Theory 11(2), 487–514 (2022)10.3934/eect.2021010
6. Attouch, H., Balhag, A., Chbani, Z., Riahi, H.: Accelerated gradient methods combining Tikhonov regularization with geometric damping driven by the Hessian. Appl. Math. Optim. 88(29), (2023)
7. Attouch H Cabot A Convergence of damped inertial dynamics governed by regularized maximally monotone operators J. Differ. Equ. 2018 264 7138 7182 10.1016/j.jde.2018.02.017
Attouch, H., Cabot, A.: Convergence of damped inertial dynamics governed by regularized maximally monotone operators. J. Differ. Equ. 264, 7138–7182 (2018)10.1016/j.jde.2018.02.017
8. Attouch H Chbani Z Peypouquet J Redont P Fast convergence of inertial dynamics and algorithms with asymptotic vanishing viscosity Math. Program. 2018 168 123 175 10.1007/s10107-016-0992-8
Attouch, H., Chbani, Z., Peypouquet, J., Redont, P.: Fast convergence of inertial dynamics and algorithms with asymptotic vanishing viscosity. Math. Program. 168, 123–175 (2018)10.1007/s10107-016-0992-8
9. Attouch H Chbani Z Riahi H Combining fast inertial dynamics for convex optimization with Tikhonov regularization J. Math. Anal. Appl. 2018 457 2 1065 1094 10.1016/j.jmaa.2016.12.017
Attouch, H., Chbani, Z., Riahi, H.: Combining fast inertial dynamics for convex optimization with Tikhonov regularization. J. Math. Anal. Appl. 457(2), 1065–1094 (2018)10.1016/j.jmaa.2016.12.017
10. Attouch H Chbani Z Riahi H Fast proximal methods via time scaling of damped inertial dynamics SIAM J. Optim. 2019 29 3 2227 2256 10.1137/18M1230207
Attouch, H., Chbani, Z., Riahi, H.: Fast proximal methods via time scaling of damped inertial dynamics. SIAM J. Optim. 29(3), 2227–2256 (2019)10.1137/18M1230207
11. Attouch H Chbani Z Riahi H Fast convex optimization via time scaling of damped inertial gradient dynamics Pure Appl. Funct. Anal. 2021 6 6 1081 1117
Attouch, H., Chbani, Z., Riahi, H.: Fast convex optimization via time scaling of damped inertial gradient dynamics. Pure Appl. Funct. Anal. 6(6), 1081–1117 (2021)
12. Attouch, H., Chbani, Z., Riahi, H.: Accelerated gradient methods with strong convergence to the minimum norm minimizer: a dynamic approach combining time scaling, averaging, and Tikhonov regularization (2022). arXiv:2211.10140v1
13. Attouch, H., Chbani, Z., Fadili, J., Riahi, H.: Convergence of iterates for first-order optimization algorithms with inertia and Hessian driven damping. J. Math. Program. Oper. Res. 72(5), (2023)
14. Attouch H Cominetti R A dynamical approach to convex minimization coupling approximation with the steepest descent method J. Differ. Equ. 1996 128 2 519 540 10.1006/jdeq.1996.0104
Attouch, H., Cominetti, R.: A dynamical approach to convex minimization coupling approximation with the steepest descent method. J. Differ. Equ. 128(2), 519–540 (1996)10.1006/jdeq.1996.0104
15. Attouch, H., Czarnecki, M.-O.: Asymptotic control and stabilization of nonlinear oscillators with non-isolated equilibria. J. Differ. Equ. 179, 278–310 (2002)
16. Attouch H László SC Continuous Newton-like inertial dynamics for monotone inclusions Set-valued Variat. Anal. 2021 29 555 581 10.1007/s11228-020-00564-y
Attouch, H., László, S.C.: Continuous Newton-like inertial dynamics for monotone inclusions. Set-valued Variat. Anal. 29, 555–581 (2021)10.1007/s11228-020-00564-y
17. Attouch, H., László, S. C.: Convex optimization via inertial algorithms with vanishing Tikhonov regularization: fast convergence to the minimum norm solution (2021). arXiv:2104.11987
18. Attouch H Peypouquet J Convergence of the inertial dynamics and proximal algorithms governed by maximally monotone operators Math. Program. 2019 174 391 432 10.1007/s10107-018-1252-x
Attouch, H., Peypouquet, J.: Convergence of the inertial dynamics and proximal algorithms governed by maximally monotone operators. Math. Program. 174, 391–432 (2019)10.1007/s10107-018-1252-x
19. Attouch H Peypouquet J Redont P Fast convex optimization via inertial dynamics with Hessian driven damping damping J. Differ. Equ. 2016 261 10 5734 5783 10.1016/j.jde.2016.08.020
Attouch, H., Peypouquet, J., Redont, P.: Fast convex optimization via inertial dynamics with Hessian driven damping damping. J. Differ. Equ. 261(10), 5734–5783 (2016)10.1016/j.jde.2016.08.020
20. Attouch H Peypouquet J Redont P Fast convergence of inertial dynamics and algorithms with asymptotic vanishing viscosity Math. Program. 2018 168 123 175 10.1007/s10107-016-0992-8
Attouch, H., Peypouquet, J., Redont, P.: Fast convergence of inertial dynamics and algorithms with asymptotic vanishing viscosity. Math. Program. 168, 123–175 (2018)10.1007/s10107-016-0992-8
21. Bauschke HH Combettes PL Convex Analysis and Monotone Operator Theory in Hilbert Spaces 2016 Springer CMS Books in Mathematics
Bauschke, H.H., Combettes, P.L.: Convex Analysis and Monotone Operator Theory in Hilbert Spaces. CMS Books in Mathematics, Springer (2016)
22. Boţ RI Csetnek ER Second order forward-backward dynamical systems for monotone inclusion problems SIAM J. Control. Optim. 2016 54 3 1423 1443 10.1137/15M1012657
Boţ, R.I., Csetnek, E.R.: Second order forward-backward dynamical systems for monotone inclusion problems. SIAM J. Control. Optim. 54(3), 1423–1443 (2016)10.1137/15M1012657
23. Boţ RI Csetnek ER László SC Tikhonov regularization of a second order dynamical system with Hessian driven damping Math. Program. 2021 189 151 186 10.1007/s10107-020-01528-8 34720194
Boţ, R.I., Csetnek, E.R., László, S.C.: Tikhonov regularization of a second order dynamical system with Hessian driven damping. Math. Program. 189, 151–186 (2021)34720194 10.1007/s10107-020-01528-8
24. Boţ, R. I., Karapetyants, M.A.: A fast continuous time approach with time scaling for nonsmooth convex optimization. Adv. Contin. Discrete Models: Theory Appl. 73 (2022)
25. Cabot A Engler H Gadat S On the long time behavior of second order differential equations with asymptotically small dissipation and insights Trans. Am. Math. Soc. 2009 361 5983 6017 10.1090/S0002-9947-09-04785-0
Cabot, A., Engler, H., Gadat, S.: On the long time behavior of second order differential equations with asymptotically small dissipation and insights. Trans. Am. Math. Soc. 361, 5983–6017 (2009)10.1090/S0002-9947-09-04785-0
26. Cabot A Engler H Gadat S Second order differential equations with asymptotically small dissipation and piecewise flat potentials Electron. J. Differ. Equ. 2009 17 33 38
Cabot, A., Engler, H., Gadat, S.: Second order differential equations with asymptotically small dissipation and piecewise flat potentials. Electron. J. Differ. Equ. 17, 33–38 (2009)
27. László SC On the strong convergence of the trajectories of a Tikhonov regularized second order dynamical system with asymptotically vanishing damping J. Differ. Equ. 2023 362 355 381 10.1016/j.jde.2023.03.014
László, S.C.: On the strong convergence of the trajectories of a Tikhonov regularized second order dynamical system with asymptotically vanishing damping. J. Differ. Equ. 362, 355–381 (2023)10.1016/j.jde.2023.03.014
28. May R Asymptotic for a second-order evolution equation with convex potential and vanishing damping term Turk. J. Math. 2017 41 3 681 685 10.3906/mat-1512-28
May, R.: Asymptotic for a second-order evolution equation with convex potential and vanishing damping term. Turk. J. Math. 41(3), 681–685 (2017)10.3906/mat-1512-28
29. Sell GR Dynamics of Evolutionary Equations 2002 New York Springer
Sell, G.R.: Dynamics of Evolutionary Equations. Springer, New York (2002)
30. Su W Boyd S Candès EJ A differential equation for modeling Nesterov’s accelerated gradient method: theory and insights J. Mach. Learn. Res. 2016 17 1 43
Su, W., Boyd, S., Candès, E.J.: A differential equation for modeling Nesterov’s accelerated gradient method: theory and insights. J. Mach. Learn. Res. 17, 1–43 (2016)
