
==== Front
MethodsX
MethodsX
MethodsX
2215-0161
Elsevier

S2215-0161(24)00373-X
10.1016/j.mex.2024.102922
102922
Statistic
Statistical inference on nonparametric regression model with approximation of Fourier series function: Estimation and hypothesis testing
Ramli Mustain
Budiantara I Nyoman i_nyoman_b@statistika.its.ac.id
⁎
Ratnasari Vita
Department of Statistics, Faculty of Science and Data Analytics, Institut Teknologi Sepuluh Nopember, Kampus ITS-Sukolilo, Surabaya 60111, Indonesia
⁎ Corresponding author. i_nyoman_b@statistika.its.ac.id
16 8 2024
12 2024
16 8 2024
13 10292226 6 2024
15 8 2024
© 2024 The Authors
2024
https://creativecommons.org/licenses/by/4.0/ This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).
Nonparametric regression is an approximation method in regression analysis that is not constrained by the assumption of knowing the regression curve. One of the functions to approximate the curve is a Fourier series function. The nonparametric regression model with approximation of a Fourier series function has been widely discussed by several researchers. However, discussions on statistical inference, particularly in partial hypothesis testing, has not been carried out previously. Therefore, the purpose of this research is to discuss the statistical inference on nonparametric regression model with approximation of a Fourier series function. The discussion includes parameter and model estimations, simultaneous and partial hypotheses testing. In the application, we use life expectancy data from East Java Province during 2022. Based on data analysis, we obtain a model estimation with an R-square value of 96.24 %. At a 5 % significance level, the parameters simultaneously have a significant influence on the model. Partially, four parameters are not significant. However, overall, the predictor variables significantly influence the life expectancy data.• The Fourier series function used is a Fourier series function introduced by Bilodeau (1992).

• The model estimation is obtained by selecting the optimal number of oscillation parameters.

• The statistical test is obtained using the LRT method.

Graphical abstract

Image, graphical abstract

Keywords

Nonparametric regression
Fourier series
Statistical inference
Likelihood ratio test
Life expectancy data
Method name

Fourier Series Function
==== Body
pmcSpecifications tableSubject area:	Mathematics and Statistics	
More specific subject area:	Statistics; Nonparametric Regression	
Name of your method:	Fourier Series Function	
Name and reference of original method:	Fourier series function developed by Bilodeau (1992), M. Bilodeau, Fourier smoother and additive models, The Canadian Journal of Statistics 20 (1992) 257–269. https://doi.org/10.2307/3315313	
Resource availability:	Life expectancy data and its predictor variables can be accessed through the website of the Badan Pusat Statistik (BPS) of East Java Province at https://jatim.bps.go.id/	

Background

Nonparametric regression is an approximation method in regression analysis that is not constrained by the assumption of knowing the regression curve. Moreover, nonparametric regression offers high flexibility, as the regression curve can be adapted to the local nature of the data [1]. However, a nonparametric regression curve cannot be determined arbitrarily without information from the data, such as examining data patterns based on a scatterplot [[2], [3]]. In nonparametric regression, there are several estimator approaches for the regression curve, such as Spline function [[4], [5], [6]], Kernel function [[7], [8], [9]], Fourier series function [[10], [11], [12], [13]], and many others. Among these approaches, Fourier series function is particularly intriguing to many researchers in the field of nonparametric regression due to its ability to handle data that exhibits repeated patterns at certain interval. According to Srivani et al. and Steinerberger, Fourier series is a trigonometric polynomial function that combines cos and sin functions [14,15]. However, Bilodeau introduced a new Fourier series function for a nonparametric regression model, which includes only the cos function and incorporates a trend to the function [10]. The Fourier series estimator introduced by Bilodeau has been widely adopted and further developed by numerous researchers for estimating the nonparametric regression curve [see, 13,[16], [17], [18], [19]]. In the application, the Fourier series estimator introduced by Bilodeau has been use for modelling various different data [[16], [17],[19], [20], [21]].

In regression analysis, statistical inference plays a crucial role in obtaining the best model estimates. Generally, statistical inference in regression analysis is divided into two main areas: estimation and hypothesis testing. Estimation in regression analysis involves estimating the parameters in the model. These parameter estimates provide a model estimation that is useful for drawing conclusions about the response variable based on the given predictor variables. Moreover, hypothesis testing in regression analysis aims to determine whether the parameters significantly influence the model. Generally, hypothesis testing in regression analysis is divided into two categories: simultaneous and partial hypotheses testing. Simultaneous hypothesis testing aims to determine the combined influence of all parameters on the model. If this testing indicates that all parameters simultaneously have a significant influence, then partial hypothesis testing is performed to identify which specific parameters influence the model. In the field of regression analysis, statistical inference has been widely discussed by several researchers [[22], [23], [24]].

The discussion related to statistical inference in nonparametric regression model with approximation of a Fourier series function has so far primarily focused on parameter and model estimation [10,12,[16], [17], [18],[20], [21]]. Ramli, et al. is the only researcher who discussed hypothesis testing in nonparametric regression model with approximation of a Fourier series function, focusing on simultaneous hypothesis testing [19]. Meanwhile, partial hypothesis testing in this model has not been carried out. The significant role of statistical inference in regression analysis, especially in nonparametric regression model with approximation of a Fourier series function makes this an interesting topic for discussion in this research. Although several aspects of statistical inference in this model, such as estimation and simultaneous hypothesis testing have been previously discussed both theoretically and practically, this research will only briefly discuss upon these areas. Therefore, the primary focus will be on the theoretical development of partial hypothesis testing, which has not been carried out previously. Furthermore, to apply statistical inference in nonparametric regression model with approximation of a Fourier series function, we use life expectancy data from 38 districts and cities in East Java Province during 2022.

Method details

The model and its estimation

In this section, we provide a general description of nonparametric regression model with approximation of a Fourier series function. The Fourier series function we utilize follows the formulation introduced by Bilodeau [10]. Parameter estimation in the nonparametric regression model with approximation of a Fourier series function is conducted using the Ordinary Least Squares (OLS) method, a technique employed by several previous researchers. Additionally, the selection of optimal parameter estimates for the model employs the Generalized Cross Validation (GCV) method.

The nonparametric regression model with approximation of a Fourier series function

Suppose yi is the response variable and xi1,xi2,...,xip are the p predictor variables, where i=1,2,...,n. Assuming the relationship between the response and predictor variables follows the nonparametric regression model as describe in Eq. (1).(1) yi=g(xi1,xi2,...,xip)+εi,εi∼N(0,σ2),

Assuming between the predictor variables are not correlated, the nonparametric regression model in Eq. (1) can be written as an additive model.(2) yi=∑j=1pg(xij)+εi.

Since g(xij) is the nonparametric regression curve which we assume as an unknown function, g(xij) can be approximated using one of the estimators in nonparametric regression model, specifically the Fourier series function [10], the Fourier series function is given in Eq. (3).(3) g(xij)=μj2+βjxij+∑t=1Tδtjcos(txij),

where μj is the constant parameter (representing the baseline level), βj is the trend parameter (representing a linear trend into the function), and δtj is the parameter for the cos terms (adding periodic component to the function), for j=1,2,...,p and t=1,2,...,T with T is the oscillation parameter.

Furthermore, since g(xij) is approximated with a Fourier series function then by substituting Eq. (3) into Eq. (2), we obtain the nonparametric regression model with approximation of a Fourier series function as follows.(4) yi=∑j=1p(μj2+βjxij+∑t=1Tδtjcos(txij))+εi.

The nonparametric regression model with approximation of a Fourier series function in Eq. (4) can be formulated in matrix and vector form as in Eq. (5), where the description holds for i=1,2,...,n,j=1,2,...,p and t=1,2,...,T [19].(5) y=Xδ+ε,ε∼N(0,Iσ2),

where y=[y1y2⋯yn]′ is a vector containing the response variable with the size of n×1, X=[X1(T)⋮X2(T)⋮⋯⋮Xp(T)] is a matrix containing the predictor variables with a Fourier series function with the size of n×p(T+2),δ=[δ1⋮δ2⋮⋯⋮δp]′ is a vector containing the parameters in the model with the size of p(T+2)×1, and ε=[ε1ε2⋯εn]′ is a vector containing the error in the model with the size of n×1. The elements of matrices X1(T),X2(T),...,Xp(T) and vectors δ1,δ2,...,δp are given as follows.Xj(T)=[12x1jcos(x1j)cos(2x1j)⋯cos(Tx1j)12x2jcos(x2j)cos(2x2j)⋯cos(Tx2j)⋮⋮⋮⋮⋱⋮12xnjcos(xnj)cos(2xnj)⋯cos(Txnj)],forj=1,2,...,p

δj=[μjβjδ1jδ2j⋯δTj]′,forj=1,2,...,p

Parameter and model estimations

In regression analysis, obtaining model estimates is equivalent to estimating the parameters in the model. In nonparametric regression model with approximation of a Fourier series function, parameter estimates can be obtained using several estimation methods, such as Penalized Least Square (PLS) method [10,12,18,20] and OLS method [[16], [17],19,21]. The choice between these two methods depends on the form of the curve g(xij). If the curve g(xij) is assumed to be a smooth function, then the Partial Least Squares (PLS) method is utilized. On the other hand, if we simply aim to obtain parameter estimates without assuming that the curve g(xij) must be smooth, we can use the OLS method. In general, the OLS method is a classic approach in regression analysis for obtaining parameter estimates by minimizing the sum of squared errors. In this research, we employ the OLS method for parameter estimation. Therefore, based on the nonparametric regression model with approximation of a Fourier series function in Eq. (5), the parameter estimation using the OLS method is obtained as follows [[16], [17],19,21].(6) δ^=argmin{ε′ε}=argminμ∈Rp(T+2){(y−Xδ)′(y−Xδ)}=(X′X)−1X′y,

where δ^ in Eq. (6) is the parameter estimation for δ in Eq. (5). Furthermore, based on δ^ in Eq. (6), we obtain the estimation of the nonparametric regression model with approximation of a Fourier series function as follows.(7) y^=Xδ^=X(X′X)−1X′y=Uy,

where U=X(X′X)−1X′ and y^ in Eq. (7) is the model estimation.

Model selection

In contrast to parametric regression models, where estimates are obtained directly from the parameters without the need for selecting an optimal model, nonparametric regression models require the selection of an optimal model estimate. Despite using approaches such as Spline, Kernel, or Fourier series functions to approximate the regression curve, there is always at least one parameter that necessitates choosing the optimal model estimate. The parameters in question are the knot points in a Spline function [[4], [5], [6]], the bandwidth in a Kernel function [[7], [8], [9]], and the oscillation parameters in a Fourier series function [16,19,21]. Therefore, a method is needed to select the optimal model. One such method is the GCV method. The GCV method was First introduced by Craven and Wahba [25]. This method has been widely used in nonparametric regression models for optimal model selection. Selecting an optimal nonparametric regression model with approximation of a Fourier series function involves choosing the optimal oscillation parameters based on the minimum GCV value. The GCV function for selecting optimal oscillation parameters is given in Eq. (8) [16,19,21].(8) GCV(TOptimum)=min{n−1y′(I−U)′(I−U)y(n−1trace[I−U])2}.

Simultaneous hypothesis testing

Simultaneous hypothesis testing is used to evaluate the influence of all parameters simultaneously in a regression model. This type of testing, particularly in nonparametric regression model with approximation of a Fourier series function was developed by Ramli et al. [19]. Therefore, explanations related to simultaneous hypothesis testing are provided briefly to support the application of the data in this research. The first step in simultaneous hypothesis testing is to formulate the hypothesis form. Based on the nonparametric regression model with approximation of a Fourier series function in Eq. (4), we can define the hypothesis form for simultaneous hypothesis testing as follows.(9) H0:μ1=μ2=...=μp=β1=β2=...=βp=δ11=δ21=...=δT1=...=δ1p=δ2p=...=δTp=0H1:Thereisminimumoneofμj≠0,βj≠0,δtj≠0,forj=1,2,...,pandt=1,2,...,T

The hypothesis form given in Eq. (9) is a specific hypothesis form used to evaluate the influence of all parameters on the model without testing the magnitude of their influence. Furthermore, to test the hypothesis form in Eq. (9), it is necessary to obtain a formula for the statistical test. One method to obtain a statistical test is the Likelihood Ratio Test (LRT) method. According to Ramli et al., using the LRT method, the formulation of the statistical test for simultaneous hypothesis testing with the hypothesis form given in Eq. (9) is as follows [19].(10) Λ=(Xδ^)′yp(T+1)+1(y−Xδ^)′(y−Xδ^)n−(p(T+1)+1),

where the statistical test of Λ in Eq. (10) is the statistical test for simultaneous hypothesis testing in nonparametric regression model with approximation of a Fourier series function. Furthermore, according to Ramli et al. [19], the statistical test of Λ follows an F distribution with degrees of freedom p(T+1)+1 and n−(p(T+1)+1). Given a significance level α for testing the hypothesis in Eq. (9), the null hypothesis will be rejected at significance level α if and only if the statistic value of Λ is greater than the statistic value of F(α,p(T+1)+1,n−(p(T+1)+1)), where F(α,p(T+1)+1,n−(p(T+1)+1)) denotes the statistic value of F distribution with significance level α and degrees of freedom p(T+1)+1 and n−(p(T+1)+1).

Partial hypothesis testing

In this section, we discuss partial hypothesis testing for parameters in nonparametric regression model with approximation of a Fourier series function. The partial hypothesis testing continues from the simultaneous hypothesis testing explained in the previous section. Following the same framework, the formula of the statistical test for partial hypothesis testing is derived using the LRT method. Furthermore, we will provide a more detailed discussion on the distribution of the statistical test and the rejection region of the null hypothesis. The initial step in partial hypothesis testing involves formulating the hypothesis form. Based on the nonparametric regression model with approximation of a Fourier series function described in Eq. (4), we can formulate the hypotheses form for partial hypothesis testing as presented in Table 1 below.Table 1 Partial hypotheses form.

Table 1Null Hypothesis	Alternative Hypothesis	Null Hypothesis	Alternative Hypothesis	Null Hypothesis	Alternative Hypothesis	
H0:μ1=0	H1:μ1≠0	H0:βp=0	H1:βp≠0	⋮	⋮	
H0:μ2=0	H1:μ2≠0	H0:δ11=0	H1:δ11≠0	H0:δ2p=0	H1:δ2p≠0	
⋮	⋮	H0:δ12=0	H1:δ12≠0	⋮	⋮	
H0:μp=0	H1:μp≠0	⋮	⋮	H0:δT1=0	H1:δT1≠0	
H0:β1=0	H1:β1≠0	H0:δ1p=0	H1:δ1p≠0	H0:δT2=0	H1:δT2≠0	
H0:β2=0	H1:β2≠0	H0:δ21=0	H1:δ21≠0	⋮	⋮	
⋮	⋮	H0:δ22=0	H1:δ22≠0	H0:δTp=0	H1:δTp≠0	

Table 1 presents a partial hypothesis formulation for all parameters in nonparametric regression model with approximation of a Fourier series function. The number of hypothesis formulations in Table 1 corresponds to the number of parameters in the model. Consequently, testing these hypotheses individually is time-consuming and involves lengthy mathematical elaboration. To address this, we combine the hypotheses in Table 1 into a simpler form to obtain the statistical test, its distribution, and the rejection region for the null hypothesis. This is achieved by formulating the hypotheses in Table 1 into the following vector form.(11) H0:alδ=0versusH1:alδ≠0

where al is a row vector with the lth column is 1 and otherwise is 0 with l=1,2,...,L, where L is the number of rows in vector δ, defined as L=p(T+2).

To obtain the formula for the statistical test in partial hypothesis testing, where the hypothesis form is given in Eq. (11), we use the LRT method. The LRT method has been widely used in hypothesis testing across various regression models [see, 19,[26], [27], [28]]. The main concept of the LRT method is to compare the likelihood function of the model under the null hypothesis with the likelihood function of the model under the alternative hypothesis. Based on the nonparametric regression model in Eq. (2), where εi is known to be normally distributed with mean 0 and variance σ2, the likelihood function is obtained as follows.(12) L(ε1,ε2,...,εn)=∏i=1n12πσ2exp(−εi22σ2)=(2πσ2)−n2exp(−12σ2∑i=1nεi2).

Furthermore, the likelihood function in Eq. (12) can be written as the following vector.(13) L(ε)=(2πσ2)−n2exp(−12σ2ε′ε).

Since the nonparametric regression curve g(xij) in Eq. (2) is approximated by the Fourier series function in Eq. (3), it follows from Eq. (5) that y=Xδ+ε. Therefore, we obtain ε=y−Xδ. Since ε=y−Xδ, the likelihood function in Eq. (13) can be written as follows.(14) L(δ,σ2)=(2πσ2)−n2exp(−12σ2(y−Xδ)′(y−Xδ)),

where the likelihood function in Eq. (14) is the likelihood function for y∼N(Xδ,Iσ2).

Furthermore, let ω be the parameter space under the null hypothesis and Ω be the parameter space under the hypothesis. Therefore, based on the hypothesis form in Eq. (11) and the likelihood function in Eq. (14), the parameter spaces ω and Ω can be determined as follows.(15) ω={δω=[δ1⋮δ2⋮⋯⋮δp]′,σω2|alδω=0}.

(16) Ω={δΩ=[δ1⋮δ2⋮⋯⋮δp]′,σΩ2}.

Based on the likelihood function in Eq. (14) and the parameter space under the null hypothesis (ω) in Eq. (15), the likelihood function under the parameter space ω is obtained as follows.(17) L(ω)=(2πσω2)−n2exp(−12σω2(y−Xδω)′(y−Xδω)).

where the likelihood function in Eq. (17) has a condition that alδω=0. Similarly, based on the likelihood function in Eq. (14) and the parameter space under the hypothesis (Ω) as defined in Eq. (16), the likelihood function under the parameter space Ω is obtained as follows.(18) L(Ω)=(2πσΩ2)−n2exp(−12σΩ2(y−XδΩ)′(y−XδΩ)).

The statistical test

The statistical test for partial hypothesis testing with the hypothesis form in Eq. (11) is obtained using the LRT method. Based on the LRT method, we obtain a likelihood ratio by comparing the maximum likelihood function under parameter space ω and the maximum likelihood function under parameter space Ω. Therefore, based on the likelihood ratio, we derive a formula for the statistical test. Furthermore, the statistical test is obtained based on Theorem 1. However, before presenting Theorem 1, we first obtain the likelihood ratio using the following equation.(19) w(xi1,xi2,...,xip;yi)=L(ω^)L(Ω^),0<w≤1,

where w in Eq. (19) is the likelihood ratio, L(ω^) is the maximum likelihood under the parameter space of ω, and L(Ω^) is the maximum likelihood under the parameter space of Ω. According to Ramli et al., if Ω is the parameter space under the hypothesis as given in Eq. (16) with the likelihood function under the parameter space Ω is provided in Eq. (18), then the maximum likelihood under the parameter space Ω is obtained as follows [19].(20) L(Ω^)=(2πσ^Ω2)−n2exp(−n2),

where L(Ω^) in Eq. (20) is the maximum likelihood under the parameter space Ω, with σ^Ω2=1n(y−Xδ^Ω)′(y−Xδ^Ω) is the variance under the parameter space Ω and δ^Ω=(X′X)−1X′y is the parameter under the parameter space Ω.

Furthermore, for L(ω^) can be obtained by estimating the parameters within the parameter space ω, namely δω and σω2. Since the parameter space ω has a specific constraint for the parameter δω, where alδω=0. Therefore, we can estimate δω using the Lagrange Multiplier (LM) method, which is commonly applied when parameters have specific constraints. This method is extensively discussed in the context of general linear hypotheses in parametric regression, as explained in Searle [29]. Based on the likelihood function under the parameter space ω in Eq. (17) with the parameter space ω defined in Eq. (15), the LM function is given as follows.(21) F(δω,θ)=M(δω)+2θalδω,

where M(δω)=(y−Xδω)′(y−Xδω)=y′y−2δ′ωX′y+δ′ωX′Xδω.

Since the LM function in Eq. (21) contains two parameters, namely δω dan θ. Therefore, the estimation of parameters δω dan θ are obtained by taking the partial derivatives OF F(δω,θ) with respect to δω and θ, and then setting these derivatives equal to 0. The estimation of parameter is obtained as follows.(22) ∂F(δω,θ)∂δω=y′y−2δ′ωX′y+δ′ωX′Xδω+2θalδω∂δω=−2X′y+2X′Xδω+2θa′l=0.

Based on Eq. (22), we obtain the estimation of δω as follows.(23) δ^ω=(X′X)−1(X′y−a′lθ).

Furthermore, for the estimation of θ is obtained as follows.(24) ∂F(δω,θ)∂θ=y′y−2δ′ωX′y+δ′ωX′Xδω+2θalδω∂θ=2alδω=0.

Based on Eq. (24), since δω is estimated by δ^ω in Eq. (23) then the estimation of θ is obtained as follows.(25) θ^=(al(X′X)−1a′l)−1al(X′X)−1X′y.

Since δ^ω in Eq. (23) still contains a parameter θ, and since parameter θ is estimated by θ^ in Eq. (25), and considering δ^Ω=(X′X)−1X′y. Therefore, the parameter δω, which is estimated by δ^ω in Eq. (23) can be written as follows.(26) δ^ω=(X′X)−1(X′y−a′l(al(X′X)−1a′l)−1al(X′X)−1X′y)=δ^Ω−(X′X)−1a′l(al(X′X)−1a′l)−1alδ^Ω

For the estimation of parameter σω2, we can directly use the Maximum Likelihood Estimation (MLE) method, by solving ∂lnL(ω)∂σω2=0. Based on the likelihood function under the parameter space ω in Eq. (17), we obtain lnL(ω) as follows.(27) lnL(ω)=−n2ln(2πσω2)−12σω2(y−Xδω)′(y−Xδω).

Therefore, based on Eq. (27), we obtain:(28) ∂lnL(ω)∂σω2=−n2σω2+12σω4(y−Xδω)′(y−Xδω)=0.

Furthermore, by elaborating on Eq. (28), and considering that δω is estimated by δ^ω in Eq. (26), the estimation of σω2 is obtained as follows.(29) σ^ω2=1n(y−Xδ^ω)′(y−Xδ^ω).

Therefore, based on (26), (29), we obtain the maximum likelihood under the parameter space ω as follows.(30) L(ω^)=(2πσ^ω2)−n2exp(−12σ^ω2(y−Xδ^ω)′(y−Xδ^ω))=(2πσ^ω2)−n2exp(−(y−Xδ^ω)′(y−Xδ^ω)2(y−Xδ^ω)′(y−Xδ^ω)n)=(2πσ^ω2)−n2exp(−n2),

where L(ω^) in Eq. (30) is the maximum likelihood under the parameter space ω. Furthermore, based on the maximum likelihood under the parameter space ω in Eq. (30) and the maximum likelihood under the parameter space Ω in Eq. (20), we can express the likelihood ratio in Eq. (19) as follows.(31) w(xi1,xi2,...,xip;yi)=L(ω^)L(Ω^)=(2πσ^ω2)−n2exp(−n2)(2πσ^Ω2)−n2exp(−n2)=(σ^ω2)−n2(σ^Ω2)−n2.

Given that σ^ω2 is provided in Eq. (29) and σ^Ω2 is provided in Eq. (20), Eq. (31) can be written as follows.(32) w(xi1,xi2,...,xip;yi)=((y−Xδ^ω)′(y−Xδ^ω)n(y−Xδ^Ω)′(y−Xδ^Ω)n)−n2=(M(δ^ω)(y−Xδ^Ω)′(y−Xδ^Ω))−n2,

where M(δ^ω)=(y−Xδ^ω)′(y−Xδ^ω)=y′y−2δ^′ωX′y+δ^′ωX′Xδ^ω. Furthermore, since δ^ω is the estimated parameter under the parameter space of ω, as given in Eq. (26), M(δ^ω) can be written as follows.M(δ^ω)=y′y−2(δ^Ω−(X′X)−1a′l(al(X′X)−1a′l)−1alδ^Ω)′X′y+(δ^Ω−(X′X)−1a′l(al(X′X)−1a′l)−1alδ^Ω)′X′X(δ^Ω−(X′X)−1a′l(al(X′X)−1a′l)−1alδ^Ω)=y′y−2δ^′ΩX′y+2δ^′Ωa′l(al(X′X)−1a′l)−1al(X′X)−1X′y+δ^′ΩX′Xδ^Ω−δ^′ΩX′X(X′X)−1a′l(al(X′X)−1a′l)−1alδ^Ω−δ^′Ωa′l(al(X′X)−1a′l)−1al(X′X)−1X′Xδ^Ω+δ^′Ωa′l(al(X′X)−1a′l)−1al(X′X)−1X′X(X′X)−1a′l(al(X′X)−1a′l)−1alδ^Ω=y′y−2δ^′ΩX′y+3δ^′Ωa′l(al(X′X)−1a′l)−1alδ^Ω+δ^′ΩX′Xδ^Ω−2δ^′Ωa′l(al(X′X)−1a′l)−1alδ^Ω=y′y−2δ^′ΩX′y+δ^′ΩX′Xδ^Ω+δ^′Ωa′l(al(X′X)−1a′l)−1alδ^Ω=(y−Xδ^Ω)′(y−Xδ^Ω)+(alδ^Ω)′(al(X′X)−1a′l)−1(alδ^Ω)

Therefore, by substituting M(δ^ω) into Eq. (32), we obtain the likelihood ratio as follows.(33) w(xi1,xi2,...,xip;yi)=((y−Xδ^Ω)′(y−Xδ^Ω)+(alδ^Ω)′(al(X′X)−1a′l)−1(alδ^Ω)(y−Xδ^Ω)′(y−Xδ^Ω))−n2=(1+(alδ^Ω)′(al(X′X)−1a′l)−1(alδ^Ω)(y−Xδ^Ω)′(y−Xδ^Ω))−n2

Based on the likelihood ratio in Eq. (33), the statistical test for partial hypothesis testing in nonparametric regression model with approximation of a Fourier series function with the hypothesis form given in Eq. (11) is presented in Theorem 1.

Theorem 1 Given the hypothesis form inEq. (11), by using the LRT method, the statistical test for testingH0:alδ=0againstH1:alδ≠0is as follows.(34) Z=alδ^Qr(al(X′X)−1a′l)

where Q=(y−Xδ^)′(y−Xδ^) and r is obtained in Theorem 2.

Proof. Noted that the partial hypothesis testing in nonparametric regression model with approximation of a Fourier series function employs the LRT method. According to Casella and Berger, the LRT method specifies a rejection region, denoted as {xi1,xi2,...,xip;yi|w(xi1,xi2,...,xip;yi)≤a}, where a is a constant satisfying 0≤a≤1 [30]. Therefore, based on the rejection region of the LRT method, the likelihood ratio in Eq. (33) can be expressed as:(35) (1+(alδ^Ω)′(al(X′X)−1a′l)−1(alδ^Ω)(y−Xδ^Ω)′(y−Xδ^Ω))−n2≤a(alδ^Ω)′(al(X′X)−1a′l)−1(alδ^Ω)(y−Xδ^Ω)′(y−Xδ^Ω)≥(a)−2n−1

By multiplying Eq. (35) by r, where r represents the degrees of freedom as obtained later in Theorem 2, Eq. (35) can be written as:(36) (alδ^Ω)′(al(X′X)−1a′l)−1(alδ^Ω)(y−Xδ^Ω)′(y−Xδ^Ω)r≥((a)−2n−1)r.

Since al is a row vector and δ^Ω is a column vector, the product of alδ^Ω dan al(X′X)−1a′l yield constants. Therefore, Eq. (36) can be written as:(37) (alδ^Ω)2(y−Xδ^Ω)′(y−Xδ^Ω)r(al(X′X)−1a′l)≥((a)−2n−1)r.

Furthermore, Eq. (37) is squared on both sides, and since δ^Ω=(X′X)−1X′y equals δ^ as of Eq. (6). Therefore, Eq. (37) can be written as:(38) |alδ^(y−Xδ^)′(y−Xδ^)r(al(X′X)−1a′l)|≥((a)−2n−1)r|Z|≥a*

where Z=alδ^Qr(al(X′X)−1a′l) with Q=(y−Xδ^)′(y−Xδ^) and a*=((a)−2n−1)r. Based on Eq. (38), the rejection region for testing the hypothesis H0:alδ=0 against H1:alδ≠0 is {xi1,xi2,...,xip;yi∥Z|≤a*}. Therefore, the statistical test for testing H0:alδ=0 against H1:alδ≠0 can be defined as follows.Z=alδ^Qr(al(X′X)−1a′l)

where Z is the statistical test for partial hypothesis testing with the hypothesis form given in Eq. (11).

Distribution of the statistical test

Based on Theorem 1, Z is the statistical test for partial hypothesis testing in nonparametric regression model with approximation of a Fourier series function using the hypothesis form in Eq. (11), where the rejection region for testing H0:alδ=0 against H1:alδ≠0 is {xi1,xi2,...,xip;yi∥Z|≤a*}, where a* contains a constant a. To determine the value of a we derive the distribution of the statistical test Z. The distribution of the statistical test Z is obtained based on Theorem 2. Before presenting Theorem 2, Lemma 1 is provided to facilitate the proof of Theorem 2.

Lemma 1 LetZis the statistical test given inTheorem 1thenZ2can be written in quadratic form as follows.Z2=y′Myy′(I−U)yr

where M=X(X′X)−1a′l(al(X′X)−1a′l)−1al(X′X)−1X′ and U is given in Eq. (7).

Proof. Noted that Z is the statistical test obtained based on Theorem 1. Furthermore, if the statistical test Z in Eq. (34) is squared, it can be written as follows.(39) Z2=(alδ^Qr(al(X′X)−1a′l))2=(alδ^)2(y−Xδ^)′(y−Xδ^)r(al(X′X)−1a′l)=(alδ^)′(al(X′X)−1a′l)−1(alδ^)y′y−2δ^′X′y+δ^′X′Xδ^r.

Since δ^=(X′X)−1X′y, Eq. (39) can be simplified in quadratic form as follows.Z2=y′X(X′X)−1a′l(al(X′X)−1a′l)−1al(X′X)−1X′yy′y−y′X(X′X)−1X′yr=y′Myy′(I−U)yr,

where Z2 describes the quadratic form of Z with y′My and y′(I−U)y are being quadratic in y.

Theorem 2 LetZdenote the statistical test for testingH0:alδ=0againstH1:alδ≠0. If the statisticZ2follows anFdistribution with degrees of freedom 1 andr, then the statistical testZfollows thetstudent distribution withrdegrees of freedom as follows.Z=alδ^ΩQr(al(X′X)−1a′l)∼t(r)

where r=n−p(T+1)+1.

Proof. Noted that Z is the statistical test for testing H0:alδ=0 against H1:alδ≠0, where Z is given in Theorem 1. Furthermore, based on Lemma 1 we have the statistic Z2 as follows.Z2=(alδ^ΩQr(al(X′X)−1a′l))2=y′Myy′(I−U)yr.

Let multiply the statistic Z2 with σ2σ2. Therefore, we have:Z2=y′Myσ2y′(I−U)yσ2r.

The statistic Z2 follows an F distribution with degrees of freedom 1 and r if and only if y′Myσ2 follows the chi-square distribution with 1 degree of freedom and y′(I−U)yσ2 follows the chi-square distribution with r degrees of freedom, where y′My and y′(I−U)y are independent. Furthermore, since y′My is being quadratic in y then to prove that y′Myσ2 follows the chi-square distribution with 1 degree of freedom is the same way as showing that matrix M is symmetry and idempotent with trace(M)=1.

For matrix M is symmetry then the product of M′=M as follows.(40) M′=(X(X′X)−1a′l(al(X′X)−1a′l)−1al(X′X)−1X′)′=X(X′X)−1a′l(al(X′X)−1a′l)−1al(X′X)−1X′=M

For matrix M is idempotent then the product of M2=M as follows.(41) M2=M′×M=M×M=X(X′X)−1a′l(al(X′X)−1a′l)−1al(X′X)−1X′X(X′X)−1a′l(al(X′X)−1a′l)−1al(X′X)−1X′=X(X′X)−1a′l(al(X′X)−1a′l)−1al(X′X)−1X′=M

For trace(M)=1, let U1=X(X′X)−1a′ and U2=al(X′X)−1X′, since al(X′X)−1a′l is a constant then trace(M) is obtained as follows.trace(M)=trace(U1(al(X′X)−1a′l)−1U2)=trace(U1U2al(X′X)−1a′l)=trace(U2U1al(X′X)−1a′l)=trace(al(X′X)−1a′al(X′X)−1a′l)=trace(1)=1

Based on (40), (41), matrix M is symmetry and idempotent with trace(M)=1. Therefore, y′Myσ2 follows the chi-square distribution with 1 degree of freedom as follows.(42) y′Myσ2∼χ(1)2.

According to Ramli, et al., y′(I−U)yσ2 follows the chi-square distribution with r degrees of freedom as follows [19].(43) y′(I−N)yσ2∼χ(r)2,

where r=n−p(T+1)+1. Furthermore, for y′My and y′(I−U)y are independent. This can be proved by showing the product of matrices M×(I−U)=0 as follows.M×(I−U)=M−M×U=X(X′X)−1a′l(al(X′X)−1a′l)−1al(X′X)−1X′−X(X′X)−1a′l(al(X′X)−1a′l)−1al(X′X)−1X′X(X′X)−1X′=X(X′X)−1a′l(al(X′X)−1a′l)−1al(X′X)−1X′−X(X′X)−1a′l(al(X′X)−1a′l)−1al(X′X)−1X′=0

Since the product of matrices M and (I−U) equals 0, it implies that y′My and y′(I−U)y are independent. Based on (42), (43), since y′Myσ2∼χ(1)2 and y′(I−U)yσ2∼χ(r)2, where y′My and y′(I−U)y are independent. Therefore, we obtain the statistic Z2 follows an F distribution with degrees of freedom 1 and r as follows.Z2=y′Myσ2y′(I−U)yσ2r=y′Myy′(I−U)yr∼F(1,r).

Since Z2∼F(1,r) and based on Lemma 1, where Z2=Z. Therefore, the statistical test Z follows the t student distribution with r degrees of freedom as follows.Z2=y′Myy′(I−U)yr=(alδ^Qr(al(X′X)−1a′l))2=alδ^Qr(al(X′X)−1a′l)∼t(r).

The rejection region for the null hypothesis

The purpose of determining the rejection region for the null hypothesis is to find out whether the null hypothesis is rejected or fails to be rejected at a significance level α. Noted that partial hypothesis testing in nonparametric regression model with approximation of a Fourier series function, with the hypothesis form in Eq. (11), uses the LRT method. Since the LRT method has the rejection region of {xi1,xi2,...,xip;yi|w(xi1,xi2,...,xip;yi)≤a}, where a is a constant. Therefore, based on Theorem 1, the critical region for testing H0:alδ=0 against H1:alδ≠0 can be written as follows.(44) C={xi1,xi2,...,xip;yi|w(xi1,xi2,...,xip;yi)≤a}={xi1,xi2,...,xip,yi||alδ^Qr(al(X′X)−1a′l)|≥((a)−2n−1)r}={xi1,xi2,...,xip,yi||Z|≥a*}

If a significance level α is given, then based on Eq. (44), the critical region for rejecting the null hypothesis for testing H0:alδ=0 against H1:alδ≠0 is given by:(45) α=P(RejectH0|H0True)α=P(|Z|≥a*|alδ=0)

Based on Theorem 2, we know that the statistical test Z follows a t student distribution with r degrees of freedom, where r=n−(p(T+1)+1). Furthermore, since a*=((a)−2n−1)r, where a is a constant, then based on Eq. (45), a* can be obtained by integrating the probability density function of the t student distribution with r degrees of freedom at a significance level α as follows.(46) ∫a*∞f(t)dt=∫−∞a*f(t)dt=α2,

where f(t) is the probability density function of the t student distribution with r degrees of freedom. Based on Eq. (46), a* represents a statistic value of the t student distribution with r degrees of freedom at the significance level α2, denoted as a*=t(α2,r). Therefore, based on Eq. (45), the critical region for rejecting the null hypothesis for testing H0:alδ=0 against H1:alδ≠0 can be written as Eq. (47).(47) α=P(|Z|≥t(α2,r)|alδ=0).

Based on Eq. (47), the null hypothesis is rejected at a significance level α if and only if |Z|≥t(α2,r) or the probability of P(|Z|≥t(α2,r)) is less than α.

Method validation

Real data application

Statistical inference on nonparametric regression model with approximation of a Fourier series function have been theoretically discussed in the previous section. To apply these concepts to real data, we gathered life expectancy data from all districts and cities in East Java Province. East Java has 29 districts and 9 cities, providing a total of n=38 observations for this research. The variables used in this research are divided into two categories: the response variable (y) and the predictor variables (x). The response variable is the life expectancy data of 38 districts and cities in East Java Province. The predictor variables used in the study include, poverty percentage (x1), open unemployment rate (x2), labor force participation rate (x3), average years of schooling (x4), and population percentage (x5). Life expectancy and predictor variables were sourced from the website of Badan Pusat Statistik (BPS) for East Java Province in 2022.

Before proceeding with modelling life expectancy data in East Java province, we first examined the relationship between life expectancy and each predictor variable. This preliminary investigation aims to discern the patterns exhibited by each predictor, thereby guiding the selection of an appropriate analytical approach. The relationship between life expectancy data and each predictor variable is expected to exhibit both unknown patterns and repeated patterns at certain intervals. Therefore, life expectancy data can be modelled using the nonparametric regression model with approximation of a Fourier series function. Furthermore, we visually explore the relationships between life expectancy and the predictor variables through scatter plots. The scatter plots illustrating the relationship between life expectancy data and variables poverty percentage (x1), open unemployment rate (x2), labor force participation rate (x3), average years of schooling (x4), and population percentage (x5) are presented in Fig. 1.Fig. 1 Scatter Plots between Response and Predictor Variables.

Fig. 1

Based on Fig. 1, the relationship between life expectancy data and each predictor variable does not exhibit a specific pattern such as linear or nonlinear, which are typically assumed in parametric regression models. Therefore, the relationship pattern between life expectancy data and each predictor variable is considered to be unknown. Furthermore, to model life expectancy data with each predictor variable, a suitable approach is to use a nonparametric regression model. One effective estimator within nonparametric regression, particularly for curve estimation of life expectancy data is a Fourier series function. Based on Eq. (4) with n=38 dan p=5, the general form of the nonparametric regression model with approximation of a Fourier series function for life expectancy data in East Java Province during 2022 can be written as follows.yi=∑j=15(μj2+βjxij+∑t=1Tδtjcos(txij))+εi,i=1,2,...,38

Since the parameter estimates in nonparametric regression model with approximation of a Fourier series function depend on the number of oscillation parameters T, achieving optimal estimation results involves selecting the optimal T. The optimal T is determined using the GCV method, where the optimal T is identified based on the smallest GCV value. In this research, the maximum number of T is for T=5. Furthermore, the selection of T in this research is approached in two scenarios, where T is consistent across all predictor variables and T varies for each predictor variable. Based on the analysis results, the GCV values obtained when T is consistent across all predictor variables are presented in Table 2.Table 2 The GCV values when T is consistent across all predictor variables.

Table 2Nilai T	GCV	MSE	MAE	RMSE	MAPE	R2 (%)	
1	1.992	1.006	0.830	1.003	0.012	73.27	
2	1.641	0.550	0.609	0.742	0.008	85.38	
3	1.748	0.350	0.445	0.592	0.006	90.70	
4	1.902	0.190	0.337	0.436	0.005	94.96	
5	3.181	0.108	0.265	0.329	0.004	97.13	

Table 2 shows the values of GCV, R2, Mean Square Error (MSE), Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Mean Absolute Percent Error (MAPE) for each T, where T is consistent across all predictor variables. According to Table 2, the smallest GCV value obtained is 1.641. Therefore, the optimal number of T is determined to be T=2 for all predictor variables. Furthermore, due to different behavioural changes observed in each predictor variable concerning the response variable, a second scenario was explored in this research where T varies for each predictor variable. By setting T=5 as the maximum, the combination of T for each predictor variable is selected when the maximum is T=2,T=3,T=4 and T=5. Based on the analysis results presented in Table 3, the minimum GCV values are obtained for each maximum T, specifically T=2,T=3,T=4 and T=5.Table 3 The minimum GCV values for the optimal combination of T.

Table 3Maximum T	Optimal Combinations of T	Minimum GCV	MSE	MAE	RMSE	MAPE	R2 (%)	
x1	x2	x3	x4	x5	
2	T=2	T=2	T=2	T=2	T=1	1.557	0.571	0.622	0.755	0.009	84.84	
3	T=2	T=2	T=3	T=3	T=2	1.345	0.373	0.466	0.611	0.007	90.10	
4	T=1	T=2	T=4	T=3	T=4	1.025	0.230	0.391	0.480	0.005	93.89	
5	T=1	T=2	T=4	T=3	T=5	0.706	0.141	0.320	0.376	0.004	96.24	

Table 3 only shows the minimum GCV values as well as R2, MSE, MAE, RMSE, and MAPE values for each optimal T combination. This is due to the substantial number of combinations for each maximum T, there are 32 combinations for T=2, 243 combinations for T=3, 1024 combinations for T=4, and 3125 combinations for T=5. For example, for a maximum of T=3, there are 243 combinations. From these 243 combinations, the smallest GCV value was 1.345, obtained with the combination T=2 for variable x1,T=2 for variable x2,T=3 for variable x3,T=3 for variable x4, and T=2 for variable x5. Based on Table 3, we obtain the smallest GCV value of 0.706, which corresponds to the maximum T=5 with the optimal combination of T being T=1 for variable x1,T=2 for variable x2,T=4 for variable x3,T=3 for variable x4, and T=5 for variable x5.

Overall, based on Table 2, Table 3, comparing the minimum GCV value when T is consistent across all predictor variables and when T varies for each predictor variable, the smallest GCV value is 0.706. This minimum GCV value is achieved with the combination T=1 for variable x1,T=2 for variable x2,T=4 for variable x3,T=3 for variable x4, and T=5 for variable x5. Therefore, we obtain the nonparametric regression model with approximation of a Fourier series function for life expectancy data in East Java Province during 2022 as follows.(48) yi=μ12+β1xi1+δ11cos(xi1)+μ22+β2xi2+∑t=12δt2cos(txi2)+μ32+β3xi3+∑t=14δt3cos(txi3)+μ42+β4xi4+∑t=13δt4cos(txi4)+μ52+β5xi5+∑t=15δt5cos(txi5)+εi,i=1,2,...,38

Furthermore, the estimated parameters in nonparametric regression model with approximation of a Fourier series function for life expectancy data in East Java Province during 2022 are presented in Table 4.Table 4 Parameter estimation.

Table 4Parameter	Estimation	Parameter	Estimation	Parameter	Estimation	
μ1	20.350	δ13	−0.334	μ5	20.350	
β1	0.242	δ23	1.390	β5	−0.171	
δ11	−1.480	δ33	−0.656	δ15	0.991	
μ2	20.350	δ43	0.902	δ25	−0.072	
β2	0.148	μ4	20.350	δ35	−1.425	
δ12	1.438	β4	1.014	δ45	−0.478	
δ22	−0.554	δ14	−0.498	δ55	0.688	
μ3	20.350	δ24	−0.200			
β3	0.140	δ34	0.471			

Noted that life expectancy data in East Java Province is modelled using the nonparametric regression model with approximation of a Fourier series function as given in Eq. (48). Furthermore, to determine whether the parameters in Eq. (48) simultaneously influence the model, simultaneous hypothesis testing is conducted. Based on Eq. (48), the hypothesis form for simultaneous hypothesis testing in nonparametric regression model with approximation of a Fourier series function for life expectancy data in East Java Province can be expressed as follows.H0:μ1=...=μ5=β1=...=β5=δ11=δ12=δ22=δ13=...=δ43=δ14=...=δ34=δ15=...=δ55=0

H1: there is at least one parameter that is not equal to 0.

Since Λ is the statistical test for simultaneous hypothesis testing in nonparametric regression model with approximation of a Fourier series function, as specified in Eq. (10), where Λ follows an F distribution with degrees of freedom p(T+1)+1 and n−(p(T+1)+1). Therefore, based on analysis results for p=5 and n=38 with the optimal combination of T given in Table 3, we obtain the statistical test Λ value of 29,767.82 with a P−value of 7.384×10−35. Furthermore, using a significance level α of 5 %, we obtain an F(0.05,21,17) value of 2.219. Since the statistic value of Λ is greater than the statistic value of F(0.05,21,17) or alternatively, since the P−value is less than 5 %, we reject the null hypothesis. Therefore, it can be concluded that at a significance level of 5 %, there is evidence that at least one parameter in Eq. (48) is not equal to 0. In other words, the parameters in the nonparametric regression model with approximation of a Fourier series function for life expectancy data in East Java Province simultaneously have a significant influence.

Furthermore, following the simultaneous hypothesis testing which concluded that at least one parameter is significant in the model, the next step is to determine which specific parameters are significant through partial hypothesis testing. Based on the nonparametric regression model with approximation of a Fourier series function for life expectancy data in East Java Province as described in Eq. (48), the hypotheses form for partial hypothesis testing are provided in Table 5.Table 5 Partial Hypotheses Form.

Table 5Null Hypothesis	Alternative Hypothesis	Null Hypothesis	Alternative Hypothesis	Null Hypothesis	Alternative Hypothesis	
H0:μ1=0	H1:μ1≠0	H0:β5=0	H1:β5≠0	H0:δ24=0	H1:δ24≠0	
H0:μ2=0	H1:μ2≠0	H0:δ11=0	H1:δ11≠0	H0:δ34=0	H1:δ34≠0	
H0:μ3=0	H1:μ3≠0	H0:δ12=0	H1:δ12≠0	H0:δ15=0	H1:δ15≠0	
H0:μ4=0	H1:μ4≠0	H0:δ22=0	H1:δ22≠0	H0:δ25=0	H1:δ25≠0	
H0:μ5=0	H1:μ5≠0	H0:δ13=0	H1:δ13≠0	H0:δ35=0	H1:δ35≠0	
H0:β1=0	H1:β1≠0	H0:δ23=0	H1:δ23≠0	H0:δ45=0	H1:δ45≠0	
H0:β2=0	H1:β2≠0	H0:δ33=0	H1:δ33≠0	H0:δ55=0	H1:δ55≠0	
H0:β3=0	H1:β3≠0	H0:δ43=0	H1:δ43≠0			
H0:β4=0	H1:β4≠0	H0:δ14=0	H1:δ14≠0			

Table 5 describes the hypothesis forms for partial hypothesis testing in nonparametric regression model with approximation of a Fourier series function for life expectancy data in East Java Province during 2022, where the general model is given in Eq. (48). Furthermore, according to Theorem 1, Z is the statistical test for partial hypothesis testing in nonparametric regression model with approximation of a Fourier series function, with Z defined by Eq. (34). Furthermore, based on Theorem 2, it is known that the statistical test Z follows a t student distribution with r degrees of freedom. Based on the analysis results for l=1,2,...25,p=5,n=38 and the optimal combination of T given in Table 3, we obtain the statistic values of |Z| and the P−value as presented in Table 6.Table 6 Partial Hypothesis Testing.

Table 6Parameter	Estimation	Statistic |Z|	P-value	Decision	
μ1	20.350	12.331	0.000	Reject H0	
β1	0.242	3.581	0.002	Reject H0	
δ11	−1.480	6.067	0.000	Reject H0	
μ2	20.350	12.331	0.000	Reject H0	
β2	0.148	1.945	0.069	Failed to Reject H0	
δ12	1.438	5.672	0.000	Reject H0	
δ22	−0.554	3.488	0.003	Reject H0	
μ3	20.350	12.331	0.000	Reject H0	
β3	0.140	2.925	0.009	Reject H0	
δ13	−0.334	2.079	0.053	Failed to Reject H0	
δ23	1.390	6.212	0.000	Reject H0	
δ33	−0.656	3.554	0.002	Reject H0	
δ43	0.902	4.538	0.000	Reject H0	
μ4	20.350	12.331	0.000	Reject H0	
β4	1.014	5.839	0.000	Reject H0	
δ14	−0.498	2.337	0.032	Reject H0	
δ24	−0.200	1.104	0.285	Failed to Reject H0	
δ34	0.471	2.780	0.013	Reject H0	
μ5	20.350	12.331	0.000	Reject H0	
β5	−0.171	2.130	0.048	Reject H0	
δ15	0.991	3.653	0.002	Reject H0	
δ25	−0.072	0.279	0.783	Failed to Reject H0	
δ35	−1.425	4.593	0.000	Reject H0	
δ45	−0.478	2.399	0.028	Reject H0	
δ55	0.688	3.263	0.005	Reject H0	

Table 6 describes the partial hypothesis testing for all parameters in Eq. (48) and shows the estimation, the statistical test, the P−value, and the decision of each parameter. Furthermore, based on Eq. (47), the null hypothesis in partial hypothesis testing in nonparametric regression model with approximation of a Fourier series function is rejected at the significance level α if and only if the statistic value |Z| is greater than the statistic value of t(α2,r), or alternatively, if the P−value is less than α. Furthermore, using a significance level α of 5 % and for r=17, we obtain the statistic value t(0.025,17) of 2.11. Based on Table 6, the decisions to reject the null hypothesis or fail to reject the null hypothesis are made by comparing each statistic value |Z| with the statistic value t(0.025,17) of 2.11 or alternatively, by comparing each P−value with α of 5 %. For example, for parameter β1, we obtain the statistic value |Z| of 3.581. This value is greater than 2.11, or the P−value of 0.002 is <5 %. Therefore, the decision is to reject the null hypothesis, indicating that parameter β1 has a significant influence on the model. Overall, at a significance level of 5 %, all parameters have a significant influence on the model, except for parameters β2,δ13,δ24 and δ25, which do not have a statistically significant influence on the model. Although several parameters are not significant, overall, the predictor variables significantly influence the life expectancy data at a 5 % significance level.

Model interpretation

The life expectancy data was modelled using the nonparametric regression model with approximation of a Fourier series function, where theoretically the model depends on the optimal number of oscillation parameter T. In this research, selecting the optimal number of oscillation parameter T is based on the smallest GCV value. The smallest GCV value was obtained as 0.706, with the optimal combinations of T given in Table 3. Therefore, the nonparametric regression model with approximation of a Fourier series function for the life expectancy data is written in Eq. (48), where the parameters estimation provided in Table 4. Furthermore, based on Eq. (48) and Table 4, the model estimation for the life expectancy data is given as follows.y^i=50.875+0.242xi1−1.480cos(xi1)+0.148xi2+1.438cos(xi2)−0.554cos(2xi2)+0.140xi3−0.334cos(xi3)+1.390cos(2xi3)−0.656cos(3xi3)+0.902cos(4xi3)+1.014xi4−0.498cos(xi4)−0.200cos(2xi4)+0.471cos(3xi4)−0.171xi5+0.991cos(xi5)−0.072cos(2xi5)−1.425cos(3xi5)−0.478cos(4xi5)+0.688cos(5xi5)

where y^i is a model estimation for the life expectancy data in East Java Province during 2022 with i=1,2,...,38. The model estimation produces an R2 value of 96.24 %, MSE value of 0.141, and other values as MAE, RMSE, MAPE are given in Table 3. This indicates that 96.24 % of the predictor variables are able to explain the response variable or simultaneously the predictor variables have an effect of 96.24 % on the response variable, with a residual variance of 0.141, indicating the variability in the response variable that is not explained by the model. This is a strong indication that the model fits the data well. Moreover, the model estimation can also be interpreted graphically by plotting the actual values of the response variable against the predicted values.

Fig. 2 shows the comparison between the actual and predicted values of the response variable. Based on Fig. 2, the predicted values follow the actual values of the response variable closely. This indicates that the model estimation fits well for predicting the response variable. Moreover, in term of efficiency of the model estimation, we have checked for the presence or absence of heteroscedasticity in the residuals of the model estimation. This can be seen using the residuals plot and the Glejser test, where the residuals plot is given in Fig. 3.Fig. 2 Comparison between Actual and Predicted Values of the Response Variable.

Fig. 2

Fig. 3 Residuals plot.

Fig. 3

Fig. 3 shows the residuals plot of the model estimation with the predicted values, where the residual values are obtained by the difference between the actual and predicted values of the response variable. Based on Fig. 3, the residuals are randomly scattered without any clear pattern, indicating that the model estimation fits well and there is no evident heteroscedasticity. In addition to the graphical analysis, we also conducted a Glejser test. The results showed the simultaneous P−value of 0.738, which is greater than the 5 % significance level. Furthermore, at a 5 % significance level, no individual parameters were found to be significant when regressing the absolute errors against all predictor variables (including predictor variables for the cos functions). Therefore, based on the residuals plot and the Glejser test, we conclude that there is no heteroscedasticity present in the residuals of the model.

Limitations

Not applicable.

Ethics statements

The data used in this research is secondary data collected from the Badan Pusat Statistik (BPS) of East Java Province during 2022. This data is available upon request or can be accessed directly from the website of the Badan Pusat Statistik (BPS) of East Java Province at https://jatim.bps.go.id/.

CRediT authorship contribution statement

Mustain Ramli: Conceptualization, Methodology, Software, Writing – original draft. I Nyoman Budiantara: Conceptualization, Methodology, Writing – review & editing, Validation, Supervision. Vita Ratnasari: Conceptualization, Methodology, Writing – review & editing, Validation, Supervision.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Data availability

Data will be made available on request.

Acknowledgments

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

Related research article: None

For a published article: None
==== Refs
References

1 Eubank R.L. Nonparametric Regression and Spline Smoothing 2nd ed. 1999 Marcel Dekker Inc. New York 10.1201/9781482273144
2 Hardle W. Applied Nonparametric Regression 1990 Cambridge University Press Cambridge
3 Cizek P. Sadikoglu S. Robust nonparametric regression: a review WIREs Comput. Stat. 12 2020 1 16 10.1002/wics.1492
4 Kuter S. Akyurek Z. Weber G.W. Estimation of subpixel snow-covered area by nonparametric regression splines Int. Arch. Photogramm. Remote Sens. Spatial Inf. Sci. XLII-2/W1 2016 31 36 10.5194/isprs-archives-XLII-2-W1-31-2016
5 Ahmed S.E. Aydin D. Yilmaz E. Estimating the nonparametric regression function by using pade approximation based on total least squares Numer. Funct. Anal. Optim. 41 2020 1827 1870 10.1080/01630563.2020.1794891
6 Putra R. Fadhlurrahman M.G. Determination of the best knot and bandwidth in geographically weighted truncated spline nonparametric regression using generalized cross validation MethodsX 10 2023 1 17 10.1016/j.mex.2022.101994
7 Aydin D. Guneri O.I. Fit A. Choice of bandwidth for nonparametric regression models using kernel smoothing: a simulation study Int. J. Sci. 26 2016 47 61
8 Lee A.B. Izbicki R. A spectral series approach to high-dimensional nonparametric regression Electron. J. Stat. 10 2016 423 463 10.1214/16-EJS1112
9 Zhu T. Politis D.N. Kernel estimates of nonparametric function autoregression models and their bootstrap approximation Electron. J. Stat. 11 2017 2876 2906 10.1214/17-EJS1303
10 Bilodeau M. Fourier smoother and additive models Canad. J. Stat. 20 1992 257 269 10.2307/3315313
11 Canditiis D.D. Feis I.D. Pointwise convergence of Fourier regularization for smoothing data J. Comput. Appl. Math. 196 2006 540 552 10.1016/j.cam.2005.10.009
12 Pane R. Budiantara I.N. Zain I. Otok B.W. Fourier series estimator in nonparametric multivariable regression and its characteristics Proc. Int. Conf. Appl. Stat. 1 2013 34 43
13 Mardianto M.F.F. Gunardi H.U. An analysis about Fourier series estimator in nonparametric regression for longitudinal data Math. Stat. 9 2021 501 510 10.13189/ms.2021.090409
14 Srivani A. Arunkumar T. Ashok S.D. Fourier harmonic regression method for bearing condition monitoring using vibration measurements Mater. Today 5 2018 12151 12160 10.1016/j.matpr.2018.02.193
15 Steinerberger S. Wasserstein distance, Fourier series and applications Monatsh. Math. 194 2021 305 338 10.1007/s00605-020-01497-2
16 Octavanny M.A.D. Budiantara I.N. Kuswanto H. Rahmawati D.P. Modeling of children ever born in Indonesia using Fourier series nonparametric regression J. Phys. Conf. Ser. 1752 2021 1 7 10.1088/1742-6596/1752/1/012019
17 Zoumb P.A.W. Li X. Fourier regression model predicting train-bridge interactions under wind and wave actions Struct. Infrastruct. Eng. 19 2023 1489 1503 10.1080/15732479.2022.2033281
18 Iriany A. Fernandes A.A.R. Hybrid Fourier series and smoothing spline path non-parametric estimation model Front. Appl. Math. Stat. 8 2023 1 8 10.3389/fams.2022.1045098
19 Ramli M. Budiantara I.N. Ratnasari V. A method for parameter hypothesis testing in nonparametric regression with Fourier series approach MethodsX 11 2023 1 13 10.1016/j.mex.2023.102468
20 Saputro D.R.S. Sukmayanti A. Widyaningsih P. The nonparametric regression model using Fourier series approximation and Penalized Least Squares (PLS) (case on data proverty in East Java) J. Phys. Conf. Ser. 1188 2019 1 7 10.1088/1742-6596/1188/1/012019
21 Prahutama A. Suparti T.W.U. Modelling Fourier regression for time series data- a case study: modelling inflation in foods sector in Indonesia J. Phys. Conf. Ser. 974 2018 1 9 10.1088/1742-6596/974/1/012067
22 Zhang J.T. Statistical inferences for linear models with functional responses Stat. Sin. 21 2011 1431 1451
23 Wei Z. Fan Y. Zhang J. Statistical inference on restricted regression models with partial distortion measurement errors Braz. J. Probab. Stat. 30 2016 464 484 10.1214/15-BJPS289
24 Guo X. Li R. Liu J. Zeng M. Reprint: statistical inference for linear mediation models with high-dimensional mediators and application to studying stock reaction to COVID-19 pandemic J. Econom. 239 2024 1 13 10.1016/j.jeconom.2023.105650
25 Craven P. Wahba G. Smoothing noisy data with spline functions Numer. Math. 31 1979 377 403 10.1007/BF01404567
26 Fonseca M. Mexia J.T. Sinha B.K. Zmyslony R. Likelihood ratio tests in linear models with linear inequality restrictions on regression coefficients REVSTAT-Stat. J. 13 2015 103 118 10.57805/revstat.v13i2.166
27 He Y. Jiang T. Wen J. Xu G. Likelihood ratio test in multivariate linear regression: from low to high dimension Stat. Sin. 31 2021 1215 1238 10.5705/ss.202019.0056
28 Liu Y. Zhu J. Lin D.K.J. A generalized likelihood ratio test for monitoring profile data J. Appl. Stat. 48 2021 1402 1415 10.1080/02664763.2021.1880555 35706466
29 Searle S.R. Linear Models 1997 John Wiley and Sons New York 10.1002/9781118491782
30 Casella G. Berger R.L. Statistical Inference 2nd ed. 2002 Duxbury Press Pacific Grove
