
==== Front
Heliyon
Heliyon
Heliyon
2405-8440
Elsevier

S2405-8440(24)13265-3
10.1016/j.heliyon.2024.e37234
e37234
Research Article
An optimal estimation approach in stratified random sampling utilizing two auxiliary attributes with application in agricultural, demography, finance, and education sectors
Almulhim F.A. a
Iqbal Kanwal kanwaliqbal3110@gmail.com
kanwal.iqbal@math.uol.edu.pk
b⁎
Al Samman Fathia M. c
Ali Asad d
Almazah Mohammed M.A. e
a Department of Mathematical Sciences, College of Science, Princess Nourah bint Abdulrahman University, P.O. Box 84428, Riyadh, 11671, Saudi Arabia
b Department of Mathematics and Statistics, University of Lahore, Sargodha-Campus, Sargodha, 40100, Pakistan
c Department of Mathematics, College of Sciences, Northern Border University, Arar, Saudi Arabia
d Department of Statistics, Govt Graduate College Abdullahpur, Faisalabad, Pakistan
e Department of Mathematics, College of Sciences and Arts (Muhyil), King Khalid University, Muhyil, 61421, Saudi Arabia
⁎ Corresponding author. kanwaliqbal3110@gmail.comkanwal.iqbal@math.uol.edu.pk
30 8 2024
15 9 2024
30 8 2024
10 17 e3723421 4 2024
23 7 2024
29 8 2024
© 2024 The Authors
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
In the contemporary era of information technology, copious amounts of data are ubiquitous, generated across various sectors on a daily basis. Analyzing every unit of data is impractical due to constraints such as limited resources in terms of time, labor, and cost. In such scenarios, survey sampling becomes a recommended approach for extracting information about population parameters. The primary goal of this study is to devise an estimation method for acquiring information about population parameters. We propose an optimal estimator for an improved estimation of the population mean in stratified random sampling by leveraging the information from two auxiliary attributes. The proposed estimator's bias, mean squared error (MSE), and minimum mean squared error are determined up to the first-order approximation. It is demonstrated that, under the derived conditions, the proposed estimator theoretically outperforms existing estimators. Four population are utilized to evaluate both the performance and applicability of the proposed estimator. The percentage relative efficiency (PRE) of proposed estimator for all the populations is 178.389, 142.881, 181.383, and 152.679 respectively. The suggested estimator superior to existing estimators, as demonstrated by the numerical examples.

Keywords

Auxiliary attributes
Optimal estimation
Stratified random sampling
Mean square error
Percentage relative efficiency
==== Body
pmc1 Introduction

Cochran [1] provides detailed information on the use of supplementary data from auxiliary variables or attributes in sample surveys in order to improve the accuracy of estimators. This involves leveraging the correlation between the main study variable and auxiliary variables. Examples of estimators that capitalize on this approach include regression, ratio, and product estimators. While Cochran [2] established that the ratio estimator is most appropriate when there is a strong positive correlation between the study and auxiliary variables, the product method of estimation is preferred for highly negatively correlated variables. However, the research variable is not always quantitative in many real-world scenarios; instead, it frequently consists of qualitative reactions that are documented as characteristics. A number of research, including those by Shabbir and Gupta [3] and Abd-Elfattah et al. [4] have concentrated on improving estimator precision by the incorporation of auxiliary attributes, with applications in the fields of power engineering, fisheries, agriculture, and health science, and more. Research indicates that the use of multiple auxiliary attributes enhances estimator efficiency. For the homogenous population, Singh et al. [5] employed auxiliary attribute information to establish ratio estimators. They computed the MSE and bias of the estimator using existing datasets. Then, they devised a modified estimator inspired by the work of Koyuncu and Kadilar [6], demonstrating its superiority over existing estimators. By incorporating qualitative auxiliary data and considering non-response in estimation, Singh and Kumar [7] developed an enhanced regression estimator for Simple Random Sampling (SRS) utilizing auxiliary information. By introducing improved estimators, pioneered the utilization of two auxiliary variables in population mean estimation. Their estimator outperformed the basic regression estimator in the current case because the auxiliary information was qualitative.

Ekpenyong and Enang [8] proposed enhanced exponential to estimate the mean of population using a SRS approach. The concept of SRS without replacement served as the foundation for the estimator's development. Lu [9] looked into how estimators created with auxiliary data were used in the fields of power engineering and agriculture. The efficiency of the estimator was demonstrated by a comparison with regression estimators and other available techniques. Zaman and Kadilar [10] created a new set of exponential estimators by leveraging auxiliary data. Generalized exponential ratio estimators with the use of auxiliary information were developed by Ahmad et al. [11].

The estimator suggested in our study is a mixed-type estimator, incorporating simple, ratio, and product estimators, whereas their estimator was exponential-based. The uses of sample surveys and estimation in the fields of agricultural and health sciences were investigated by Mahajan et al. [12]. A predictive method for determining the population mean under auxiliary attributes was proposed by Kumar and Saini [13]. Auxiliary variables were employed by Yunusa et al. [14] to create novel regression-type estimators to compute the population mean estimate. For population mean estimation, auxiliary data were employed by Rather et al. [15] to create a miscellaneous exponential ratio-style estimator. Samples were chosen using double sampling and simple random sampling methods. An exponential-type estimator was designed by Zaman et al. [16] to evaluate and estimate the COVID-19 risk in numerous nations. They used data from two auxiliary variables to propose two multivariate sets of exponential-type estimators. A number of authors have covered several data analysis methods for the examination of agricultural information, including Wayangkau et al. [17], Waheeb et al. [18], Rajak [19], and Jabal et al. [20]. Saini et al. [21] introduced an optimal estimator for an improved estimation under SRS method to estimate the mean of population by leveraging the information from two auxiliary attributes. The aforementioned literature serves as motivation for investigating the suitability of two auxiliary attributes in determining the average of varied populations related to the agricultural, demography, finance, and education sectors.

In many studies, estimating parameters for the heterogeneous population has been a keen interest for survey statisticians. Therefore, for the heterogeneous population, Neyman [22] introduced Stratified Random Sampling (StRS), in which the population is partitioned into groups called “strata.” Then a sample is chosen by some pattern within each stratum, and independent selections are made in different groups. Stratification is a probability sampling design used to increase the precision of estimation. The sampling procedure of StRS ensures that every section of the population gets an appropriate representation in the sample. For small sample sizes, adequate precision and accuracy may be achieved by using the procedure of StRS. In sample selection, the StRS design helps to minimize bias. From an organizational point of view, stratified sampling is very convenient. Moreover, In StRS, the procedure of estimating population parameters by utilizing auxiliary information is also suggested and is useful for comparing estimates among several population groups. For more recent studies on StRS, see Bhushan, Kumar, Shahab, Lone, Akhtar [23], Ahmad, Hussain, Aamir, Yasmeen, Shabbir, Ahmad [24], and Javed, Irfan, Bhatti, Onyango [25] references therein. Ping Ma [26] presented a straightforward yet effective framework for evaluating the statistical properties of algorithmic leveraging in the context of parameter estimation for linear regression models with a fixed number of predictors. Specifically, we derive results for the bias and variance both conditional and unconditional on the observed data for several variants of leverage-based sampling. Our findings reveal that, from a statistical perspective, in terms of bias and variance, leverage-based sampling does not outperform uniform sampling, nor does uniform sampling dominate leverage-based sampling. This is particularly notable given that, from an algorithmic perspective focused on worst-case analysis, leverage-based sampling consistently offers superior worst-case performance compared to uniform sampling.

Ping Ma [27] developed an asymptotic analysis to derive the distribution of RandNLA sampling estimators for the least-squares problem. In particular, we derive the asymptotic distribution of a general sampling estimator with arbitrary sampling probabilities in a fixed design setting. The analysis is conducted in two complementary settings, i.e., when the objective of interest is to approximate the full sample estimator, and when it is to infer the underlying ground truth model parameters. For each setting, we show that the sampling estimator is asymptotically normally distributed under mild regularity conditions. Moreover, the sampling estimator is asymptotically unbiased in both settings. Based on our asymptotic analysis, we use two criteria, the Asymptotic Mean Squared Error (AMSE) and the Expected Asymptotic Mean Squared Error (EAMSE), to identify optimal sampling probabilities.

The primary aim of the current study is to introduce a new estimator/method for estimating to estimate the mean of population for heterogenous environment, incorporating information from two auxiliary attributes. Below, we highlight the foremost contribution of the present manuscript. The study involves deriving the bias and MSE of the novel estimator/method and comparing them with those of existing estimators to assess theoretical efficiency. The practical applicability of the proposed estimator is emphasized through its application to real datasets from diverse sectors. By conducting both theoretical analyses and numerical evaluations, we illustrate that the proposed estimator exhibits greater efficiency compared to existing estimators.

The article is structured into several sections. Section 2 covers notation, methodology, and existing works, while Section 3 introduces the proposed estimators. Efficiency conditions are developed in Section 4, Section 5 demonstrates the application of the proposed methods, and finally, Section 6 provides the conclusion of the manuscript.

2 Materials and methods

2.1 Notations and existing estimators

Let δ be the population size, δ=δ1+δ2+δ3+.....+δh+....+δN, where δh is the number of units in stratum h, δˆ=δˆ1+δˆ2+δˆ3+.....+δˆh+....+δˆn, where δˆh in the number of sampling units in the hth stratum. Let Yi refers the study variable and τi be the characteristics of the auxiliary attributes i.e. τi=1 if the ith unit possess attribute and τi=0, otherwise. Let G=∑i=1Nτi represent the total number of units in the population possessing attributes τi and g=∑i=1nτi represent the total number of units in the sample possessing attributes τi. Let D=GN be the proportion of units in the population and D=gn be the population of units in the sample. The Y‾h refers to the population mean of the variable of interest Y in stratum h, D1handD2h is the population proportion in auxiliary attribute in stratum h. y‾h are the sample means of Y in stratum h. d1handd2h is the sample proportion in auxiliary attribute in stratum h. Wh=NhN, Wh2=Nh2N2 is the stratum weight. The Syh2 is the population variance of study variable in stratum h, Sd1h2andSd2h2 are the population variance of auxiliary attribute in stratum h.

Cyh is the coefficient of variation in stratum h of the study variable while Cd1handCd2h are the coefficients of variations in stratum h of the auxiliary attributes. Let Sydj be the population covariance between the study variable y and the auxiliary attributes dj=(j=1,2). Let Sd1hd2h be the population covariance between the auxiliary attributes d1h and d1h. ρydjh is the correlation coefficient between study variable and auxiliary attributes in stratum h.

To obtain biases and mean square error (MSE) of the proposed estimator, we consider the following relations under stratified random sampling:

e0=y‾h−Y‾hY‾h,y‾h=Y‾h(1+e0),e1=d1h−D1hD1h,d1h=D1h(1+e1),e2=d2h−D2hD2h,D2h=d2h(1+e2) be the error terms such that E(e0)=E(e3)=E(e4)=0.E(e02)=∑h=1nWh2λhCyh2,E(e12)=∑h=1nWh2λhCd1h2,E(e22)=∑h=1nWh2λhCd2h2, E(e0e1)=∑h=1nWh2λhρyhd1hCyhCd1h,E(e0e2)=∑h=1nWh2λhρyhd2hCyhCd2h,E(e1e2)=∑h=1nWh2λhρd1hd2hCd1hCd2h,andλh={1nh−1Nh}..

The proposed method is outlined in a stepwise framework as follows.Stage 1 Start with the δ (size of finite population).

Stage 2 Choose a random sample of size δˆh from the h strata's

δˆ=δˆ1+δˆ2+δˆ3+.....+δˆh+....+δˆn from total population δ using stratified random sampling.Stage 3 Detect Yi and τi. from the selected sampling units.

Stage 4 Formulate expressions for the characteristic's population and sample.

Stage 5 Develop and evaluate the properties of an estimator that uses two auxiliary attributes to estimate the population mean.

Stage 6 Through both theoretically and numerically evaluation the suggested estimator will be compared with current estimators.

Further, Table 1 discusses some well-stablished existing estimators and their bias and mean square error.Table 1 Some existing estimators along with bias and mean square error.

Table 1Source	Estimator	Bias	MSE	
Cochran [2]	μˆC=y‾		V(μˆC)=V(y‾)=λY‾2Cy2	
Naik and Gupta [28]	(μˆNG1)=y‾(D1d1)	Bias(μˆNG1)≅λY‾[Cd12−ρyd1CyCyd1]	MSE(μˆNG1)≅λY‾2[Cy2+Cd12−2ρyd1CyCyd1]	
Naik and Gupta [28]	(μˆNG2)=y‾(d1D1)	Bias(μˆNG2)≅λY‾[ρyd1CyCyd1]	MSE(μˆNG2)≅λY‾2[Cy2+Cd12+2ρyd1CyCyd1]	
Singh et al. [5]	(μˆS1)=y‾exp[D1−d1D1+d1]	Bias(μˆS1)≅λY‾[38Cd12−12ρyd1CyCyd1]	MSE(μˆS1)≅λY‾2[Cy2+14Cd12−2ρyd1CyCyd1]	
Singh et al. [5]	(μˆS2)=y‾exp[d1−D1d1+D1]	Bias(μˆS2)≅12λY‾[ρyd1CyCyd1−14Cd12]	MSE(μˆS2)≅12λY‾2[Cy2+14Cd12+ρyd1CyCyd1]	
Kumar and Bhougal [29]	(μˆKB)=y‾[αexp(D1−d1D1+d1)+(1−α)exp(D1−d1D1+d1)]	Bias(μˆKB)≅λY‾[18(4α−1)Cd12−(α−12)ρyd1CyCyd1]	MSE(μˆKB)≅λY‾2Cy2[1−ρyd12]	
Singh and Kumar [7]	(μˆSK)=y‾[(d1D1)+(d1D2)]	Bias(μˆSK)≅λY‾[2ρd1d2Cd1Cd2−ρyd1CyCd1+ρyd2CyCd2]	MSE(μˆSK)≅λY‾2[Cy2+Cd12+Cd22+2(2ρd1d2Cd1Cd2−ρyd1CyCd1+ρyd2CyCd2)]	
Ahmed et al. [8]	(μˆA)=y‾[exp(S1−M1S1+M1)+(S2−M2S2+M2)]	Bias(μˆA)≅λY‾[σ1ρyd1CyCd1+σ2ρyd2CyCd1Cd2+18Cd12(σ12−2σ1v1)+18Cd22(σ22−2σ2v2)]	MSE(μˆA)≅Y‾2Cy2[1−Ryd1d22]	
Saini et al. [21]	(μˆS)=[ω1y‾+ω2(D1−d1)+ω3(D2−d2)4+(D1d1+d1D1)(D2d2+d2D2)]	Bias(μˆS)≅Y‾[(ω1−1)+12λω1Cd12+12λω1Cd22]	MSE(μˆSK)≅λY‾2[Cy2+Cd12+Cd22+2(2ρd1d2Cd1Cd2−ρyd1CyCd1+ρyd2CyCd2)]	

2.2 Proposed estimator and its properties

Inspired by the work of Saini et al. [21], following a similar methodology, we introduce an estimator for heterogenous population to estimate the population mean, leveraging two auxiliary attributes.(1) (μˆK,StRS)=∑h=1Nδhδ[ω1y‾h+ω2(D1h−d1h)+ω3(D2h−d2h)+Y‾h4(D1hd1h+d1hD1h)(D2hd2h+d2hD2h)]

However, for the hth stratum, let μˆK,StRS be an optimum estimator of the population mean given as follows:(2) (μˆK(h),StRS)=[ω1y‾h+ω2(D1h−d1h)+ω3(D2h−d2h)+Y‾h4(D1hd1h+d1hD1h)(D2hd2h+d2hD2h)]

To derive the suggested estimator's bias and mean square error expression, we incorporate the error approximation into (20) and expressing the proposed estimator in terms of e′s, we obtain:(3) (μˆK(h),StRS)=[ω1Y‾h+ω1Y‾he0−ω2D1he1+ω3D2he2+Y‾h4(1(1+e1)+(1+e1))(1(1+e2)+(1+e2))]

Using the Taylor series and neglecting the high-order terms of (21), we get,(4) (μˆK(h),StRS)=[ω1Y‾h+ω1Y‾he0−ω2D1he1+ω3D2he2+Y‾h4(4+2e12+2e22)]

(5) (μˆK(h),StRS−Y‾h)=[ω1Y‾h+ω1Y‾he0−ω2D1he1+ω3D2he2+12Y‾he12+12Y‾he22]

We need to take the expectation on both sides of (23), in order to determine the bias of the novel estimator. As a result, the suggested estimator's bias will be revealed up to the first level of approximation (see Fig. 1).(6) (μˆK(h),StRS−Y‾h)=[ω1Y‾h+ω1Y‾h(e0)−ω2D1h(e1)+ω3D2h(e2)+12Y‾hE(e12)+12Y‾hE(e22)]

Fig. 1 Mean square error (MSE) of all estimators.

Fig. 1

The bias of the μˆK(h),StRS is expressed as:Bias(μˆK(h),StRS)=ω1Y‾h+12Y‾hλh[Cd1h2+Cd2h2]

Hence, the Bias expression of μˆK,StRS is obtained as follows:Bias(μˆK(h),StRS)=∑h=1Nδh2δ2Bias[μˆK(h),StRS]

Bias(μˆK(h),StRS)=∑h=1Nδh2δ2(ω1Y‾h+12Y‾h=λh[Cd1h2+Cd2h2])

To derive the Mean Square Error of the suggested estimator, let's make both sides of equation (23) square, ignoring terms involving e′s with powers larger than two. After taking the expectation on both sides and simplifying, we obtain:(7) MSE(μˆK(h),StRS)=[ω12Y‾h2+ω1Y‾h2λCd2h2+ω12Y‾h2λhCyh2−2Y‾hω1ω2D1hλρyhd1hCyhCd1h−2Y‾hω1ω3D2hλhρyhd2hCyhCd2h+ω22D1h2λhCd1h2+2ω2ω3D1hD2hλhρd1hd2hCd1hCd2h+ω32D2h2λhCd2h2+Y‾h2ω1λhCd1h2]

Setting the partial derivatives with respect to ω1,ω2,andω3 equal to zero, the optimum values of ω1,ω2,andω3 are determined by:ω1*=Cd1h2−1+Cyh2λhαh

ω2*=Y‾hCyhCd1hγhD1βh(−1+Cyh2λhαh)

Where γh=ρyhd1h−ρyhd1hd2hρyhd2h.ω3*=Y‾hCyhCd1h2ηhD2hCd2hβh(−1+Cyh2λhαh)

Where ηh=ρyhd2h−ρyhd1hd2hρyhd1h.

The minimum Mean Squared Error (MSE) of μˆK(h),StRS can be expressed as:(8) MSE(μˆK(h),StRS)=Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+σhλhCyh2Cd1h2+CyhCd1h2Cd2hλhξh]

Where σh=−2βh(ρyhd1hγh−ρyhd2hηh)+(γh2−ηh)2,ξ=2γhηhρyhd1hd2h.

Hence, the mean square error expression of μˆK,StRS is given by,(9) MSE(μˆK,StRS)=∑h=1Nδ2δ2MSE(μˆK(h),StRS)

(10) MSE(μˆK,StRS)=∑h=1Nδh2δ2Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+σhλhCyh2Cd1h2+CyhCd1h2Cd2hλhξh]

2.3 Theoretical comparison of estimators

In this area, efficiency conditions are established by comparing the MSE of the suggested estimators with numerous existing estimators:

Referring to Table (10), (2),

MSE(μˆC)>MSE(μˆK,StRS)min if and only ifCy2>∑h=1Nδh2δ2Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+σhλhCyh2Cd1h2+CyhCd1h2Cd2hλhξh]

Referring to Table (10), (4),

MSE(μˆNG1)>MSE(μˆK,StRS)min if and only if[Cy2+Cd12−2ρyd1CyCyd1]>∑h=1Nδh2δ2Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+σhλhCyh2Cd1h2+CyhCd1h2Cd2hλhξh]

Referring to Table (10), (6),

MSE(μˆNG2)>MSE(μˆK,StRS)min if and only if[Cy2+Cd12+2ρyd1CyCyd1]>∑h=1Nδh2δ2Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+σhλhCyh2Cd1h2+CyhCd1h2Cd2hλhξh]

Referring to Table (10), (8),

MSE(μˆS1)>MSE(μˆK,StRS)min if and only if[Cy2+14Cd12−2ρyd1CyCyd1]>∑h=1Nδh2δ2Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+σhλhCyh2Cd1h2+CyhCd1h2Cd2hλhξh]

Referring to Table (1) and Equation (28),

MSE(μˆS2)>MSE(μˆK,StRS)min if and only if[Cy2+14Cd12+ρyd1CyCyd1]>∑h=1Nδh2δ2Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+σhλhCyh2Cd1h2+CyhCd1h2Cd2hλhξh]

Referring to Table (1) and Equation(28),

MSE(μˆKB)>MSE(μˆK,StRS)min if and only if[1−ρyd12]>∑h=1Nδh2δ2Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+σhλhCyh2Cd1h2+CyhCd1h2Cd2hλhξh]

Referring to Table(1) and Equation (28),

MSE(μˆSK)>MSE(μˆK,StRS)min if and only if[Cy2+Cd12+Cd22+2(2ρd1d2Cd1Cd2−ρyd1CyCd1+ρyd2CyCd2)]>∑h=1Nδh2δ2Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+σhλhCyh2Cd1h2+CyhCd1h2Cd2hλhξh]

Referring to Table (1) and Equation (28),

MSE(μˆA)>MSE(μˆK,StRS)min if and only if[1−Ryd1d22]>∑h=1Nδh2δ2Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+σhλhCyh2Cd1h2+CyhCd1h2Cd2hλhξh]

Referring to Table (1) and Equation (28),

MSE(μˆS)>MSE(μˆK,StRS)min if and only if[4Cy2α−βλ(Cd12+Cd22)2]4[β+Cy2λα+βλ(Cd12+Cd22)]>∑h=1Nδh2δ2Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+σhλhCyh2Cd1h2+CyhCd1h2Cd2hλhξh]

Under these conditions, the suggested estimators are expected to exhibit efficiency comparable to that of the ordinary estimators.

3 Results and discussion

3.1 Applications

To observe the performance of suggested estimator with respect to numerous existing estimators for heterogenous environment, below are some real datasets used to demonstrate the application of the proposed estimators.

3.1.1 Data- I: agriculture data

The population I is nominated from Kadilar and Cingi [30]. In this case, y represents the quantity of apples produced in 1999; d1 and d2 indicate the percentages of apple trees that produce more than 20,000 and 25,000 apples respectively in 1999 and 1998. With respect to the proportional arrangement among various strata.

Table 2 presents the summary statistics of agriculture data. Population I indicate that there are 854 people in the population and 200 people in the sample. Furthermore, correlations, sample mean, and coefficient of variation are computed, and it is discovered that all of the correlations are positive.Table 2 Summary statistics of population I.

Table 2δ1=106	δ2=106	δ3=94	δ4=171	δ5=204	δ6=173	
δˆ1=13	δˆ2=24	δˆ3=55	δˆ4=95	δˆ5=10	δˆ6=3	
Y‾1=1536.774	Y‾2=212.594	Y‾3=9384.309	Y‾4=5588.02	Y‾5=966.96	Y‾6=404.39	
Cy1=4.181	Cy2=5.221	Cy3=3.187	Cy4=5.126	Cy5=2.472	Cy6=2.339	
Cd11=1.762	Cd12=1.563	Cd13=1.095	Cd14=1.069	Cd15=1.358	Cd16=2.773	
Cd21=1.857	Cd22=1.677	Cd23=1.220	Cd24=1.265	Cd25=1.483	Cd26=2.942	
ρy1d11=0.35	ρy2d12=0.27	ρy3d13=0.33	ρy4d14=0.19	ρy5d15=0.39	ρy6d16=0.67	
ρy1d21=0.38	ρy2d22=0.28	ρy3d23=0.36	ρy4d24=0.23	ρy5d25=0.42	ρy6d26=0.68	
ρd11d21=0.95	ρd12d22=0.93	ρd13d23=0.89	ρd14d24=0.85	ρd15d25=0.89	ρd16d26=0.88	

3.1.2 Data-II: demography data

The population II is selected from Sarndal et al. [31]. In this case, y signifies the population in thousands in 1985, d1 the population proportion is less than 60 in 1975, and d2 the proportion of less than 100 (one hundred) seats in the municipal council. Using the proportionate distribution among the various strata,

Table 3 shows the summary statistics of demography data. Population II indicate that there are 284 people in the population and 68 people in the sample. Furthermore, the sample mean, coefficient of variation, and correlations are computed, revealing that certain correlations between the variables are positive and others are negative.Table 3 Summary statistics of population II.

Table 3δ1=25	δ2=48	δ3=32	δ4=38	δ5=56	δ6=41	
δˆ1=6	δˆ2=11	δˆ3=8	δˆ4=9	δˆ5=13	δˆ6=10	
Y‾1=62.44	Y‾2=29.6	Y‾3=24.06	Y‾4=31	Y‾5=29.41	Y‾6=20.83	
Cy1=1.99	Cy2=1.22	Cy3=0.87	Cy4=1.26	Cy5=1.92	Cy6=0.85	
Cd11=0.37	Cd12=0.41	Cd13=0.26	Cd14=0.39	Cd15=0.24	Cd16=0.22	
Cd21=3.46	Cd22=1.09	Cd23=1.15	Cd24=1.32	Cd25=0.97	Cd26=1.14	
ρy1d11=−0.61	ρy2d12=−0.93	ρy3d13=−0.77	ρy4d14=−0.78	ρy5d15=−0.72	ρy6d16=−0.77	
ρy1d21=−0.12	ρy2d22=−0.52	ρy3d23=−0.56	ρy4d24=−0.36	ρy5d25=−0.35	ρy6d26=−0.56	
ρd11d21=0.10	ρd12d22=0.38	ρd13d23=0.22	ρd14d24=0.29	ρd15d25=0.24	ρd16d26=0.20	

3.1.3 Data III: education data

The population III is obtained from Koyuncu and Kadilar [6]. In this case, y represents the number of teachers, d1 the percentage of Turkish primary and secondary school students in 2007 across 923 districts in six regions where there are fewer than 1000 students, and d2 the percentage of Turkish primary and secondary school students in 2008 across 923 districts in six regions where there are less than 200 students.

Table 4 uses a proportionate allocation across various strata to yield a sample size of 180 from a population of 923. In addition, calculations are made for the sample mean, coefficient of variation, and correlations; all correlations are determined to be positive.Table 4 Summary statistics of population III.

Table 4δ1=127	δ2=117	δ3=103	δ4=170	δ5=205	δ6=201	
δˆ1=31	δˆ2=21	δˆ3=29	δˆ4=38	δˆ5=22	δˆ6=39	
Y‾1=703.74	Y‾2=413	Y‾3=573.17	Y‾4=424.67	Y‾5=267.03	Y‾6=393.84	
Cy1=1.25	Cy2=1.56	Cy3=1.8	Cy4=1.9	Cy5=1.52	Cy6=1.8	
Cd11=0.35	Cd12=0.32	Cd13=0.34	Cd14=0.52	Cd15=0.45	Cd16=0.33	
Cd21=0.82	Cd22=1.05	Cd23=0.93	Cd24=1.26	Cd25=1.39	Cd26=0.97	
ρy1d11=0.27	ρy2d12=0.19	ρy3d13=0.18	ρy4d14=0.25	ρy5d15=0.25	ρy6d16=0.16	
ρy1d21=0.56	ρy2d22=0.48	ρy3d23=0.43	ρy4d24=0.53	ρy5d25=0.58	ρy6d26=0.40	
ρd11d21=0.43	ρd12d22=0.30	ρd13d23=0.37	ρd14d24=0.41	ρd15d25=0.33	ρd16d26=0.34	

3.1.4 Data IV: finance data

The population IV is chosen from Gerard et al. [32]. In this case, y denotes the total taxes in euros for 2001; d1 and d2 represent the percentage of total taxable income in euros for 2001 that is below the mean and the percentage of total average income in euros for 2001 that is below the mean, respectively.

Table 5 shows the summary statistics of finance data which explain that the population size (589) and sample size (150) using the proportionate distribution among the various strata. Furthermore, the sample mean, coefficient of variation, and correlations are computed, revealing that certain correlations between the variables are positive and others are negative.Table 5 Summary statistics of population IV.

Table 5δ1=70	δ2=111	δ3=64	δ4=65	δ5=69	δ6=84	
δˆ1=18	δˆ2=28	δˆ3=16	δˆ4=17	δˆ5=18	δˆ6=21	
Y‾1=81845269	Y‾2=77637833	Y‾3=53163090	Y‾4=72678044	Y‾5=45901248	Y‾6=33367892	
Cy1=2.05	Cy2=0.88	Cy3=1.20	Cy4=1.40	Cy5=1.35	Cy6=1.68	
Cd11=0.89	Cd12=0.90	Cd13=0.60	Cd14=0.93	Cd15=0.53	Cd16=0.39	
Cd21=0.97	Cd22=1.72	Cd23=0.29	Cd24=1.05	Cd25=0.31	Cd26=0.52	
ρy1d11=−0.30	ρy2d12=−0.68	ρy3d13=−0.68	ρy4d14=−0.42	ρy5d15=−0.63	ρy6d16=−0.58	
ρy1d21=0.10	ρy2d22=0.12	ρy3d23=−0.16	ρy4d24=0.15	ρy5d25=0.04	ρy6d26=0.07	
ρd11d21=0.11	ρd12d22=−0.18	ρd13d23=0.35	ρd14d24=−0.16	ρd15d25=−0.16	ρd16d26=−0.11	

The results for these datasets are computed in terms of MSE and PRE. The Percent Relative Efficiency (PRE) of the estimators is calculated against the basic mean estimator by considering the following expression.(11) PRE=V(μˆC)MSE(μˆith)×100

Where i=C,NG1,NG2,S1,S2,KB,SK,A,Sandk,StRS.

Table 6 shows the necessary calculations (MSE) for proposed and comparative estimators with respect to StRS for those four real-life datasets mentioned. In dataset I, μˆk,StRS has the lowest MSE, coming in at 4177.948. In a similar vein, for data set II, III, and IV μˆk,StRS additionally displays 380.036, 770.7447, and 424661186.551 as dominant MSE values, respectively. This implies that when compared to comparative estimators, the suggested estimator performs remarkably well and efficiently. In Fig. 1, the suggested estimator's performance is further illustrated graphically.Table 6 Mean square error (MSE) of all estimators.

Table 6Estimator	Population-I	Population-II	Population-III	Population-IV	
μˆC	7453.537	543.738	1398.378	648368453.278	
μˆNG1	6797.763	490.346	1345.628	505440101.273	
μˆNG2	7183.752	500.603	1081.298	545847395.201	
μˆS1	7105.403	517.251	1279.188	493472401.308	
μˆS2	6218.502	455.197	1353.483	559488163.482	
μˆKB	6456.223	415.172	1190.881	626551915.338	
μˆSK	6691.806	425.885	1268.637	639927805.229	
μˆA	7265.619	443.272	1237.902	499163493.297	
μˆS	6911.292	413.906	1278.007	534159755.677	
μˆk,StRS	4177.948	380.036	770.7447	424661186.551	

Table 7 show the results of PRE for all estimators with respect to μˆC based on all four populations discussed earlier. In dataset I, μˆk,StRS has the highest PREE, coming in at 178.389. In a similar vein, for data set II, III, and IV μˆk,StRS additionally displays 142.881, 181.383, and 152.679 as dominant PREE values, respectively. Results shows that our proposed estimator outperformed its all-competitor estimators discussed above. In population 2 and 4 PRE of our proposed estimator decreases due to some negative correlations in data sets. Overall, proposed estimator performed most efficiently in population 3 with 181.383 percent more than estimator proposed by Cochran [2]. In Fig. 2, the suggested estimator's performance is further illustrated graphically.Table 7 Percent Relative Efficiency (PRE) with respect to μˆC.

Table 7Estimator	Population-I	Population-II	Population-III	Population-IV	
μˆC	100	100	100	100	
μˆNG1	109.639	110.738	103.892	128.278	
μˆNG2	103.748	108.469	129.289	118.782	
μˆS1	104.892	104.978	109.288	131.389	
μˆS2	119.852	119.289	103.289	115.886	
μˆKB	115.439	130.789	117.392	103.482	
μˆSK	111.375	127.499	110.197	101.319	
μˆA	102.579	122.498	112.933	129.891	
μˆS	107.838	131.189	109.389	121.381	
μˆk,StRS	178.389	142.881	181.383	152.679	

Fig. 2 Percent Relative Efficiency (PRE) of estimators with respect to μˆC.

Fig. 2

4 Conclusion and future recommendations

In different scenarios, using auxiliary information is a very useful strategy to improve the estimator's efficiency. This article introduces an optimal method for estimating the population mean under heterogenous environment by leveraging information from two auxiliary attributes inspired by the work of Saini et al. [21]. We have derived the Mean Squared Error (MSE) of the proposed class of estimators. Efficiency conditions are established by comparing the MSEs of the proposed and conventional estimators. Furthermore, we analyze several widely-used real datasets to showcase the performance of the proposed estimators. The numerical findings from the real data analysis affirm that the suggested estimators outperform existing ones, exhibiting lower MSE and higher percentage relative efficiency (PRE). We suggest employing this estimator in sample surveys across domains such as agriculture, health sciences, education and fisheries. Ultimately, we conclude that the efficiency of the suggested estimator holds both theoretically and in practical settings. Moreover, we suggest extending the scope of this study to include other types of estimators, such as ratio, product, power, difference, exponential, and regression estimators, under stratified random sampling. Additionally, the method can be expanded to incorporate multiple auxiliary sources under various sampling frameworks.

Data availability statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.

CRediT authorship contribution statement

F.A. Almulhim: Funding acquisition, Conceptualization. Kanwal Iqbal: Writing – original draft, Validation, Software, Methodology, Data curation, Conceptualization. Fathia M. Al Samman: Validation, Investigation. Asad Ali: Visualization, Project administration, Investigation, Formal analysis. Mohammed M.A. Almazah: Writing – review & editing, Supervision, Resources.

Declaration of competing interest

The authors declare the following financial interests/personal relationships which may be considered as potential competing interests:

F.A. Almulhim reports financial support was provided by Princess Norah bint Abdulrahman University,. F.A. Almulhim reports a relationship with Princess Norah bint Abdulrahman University, that includes: employment and funding grants. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Appendix A.1 Definition of the estimator μˆK,StRS

(μˆK,StRS)=∑h=1Nδh2δ[ω1y‾h+ω2(D1h−d1h)+ω3(D2h−d2h)+Y‾h4(D1hd1h+d1hD1h)(D2hd2h+d2hD2h)]

However, for the hth stratum, let μˆK,StRS be an optimum estimator of the population mean given as follows:(μˆK(h),StRS)=[ω1y‾h+ω2(D1h−d1h)+ω3(D2h−d2h)+Y‾h4(D1hd1h+d1hD1h)(D2hd2h+d2hD2h)]

A.2 Derivation ofBias(μˆK,StRS)andMSE(μˆK,StRS)

To derive the suggested estimator's bias expression of (μˆK,StRS) , we incorporate the error approximation into the above equation by using the notation and expressing the proposed estimator in terms of e′s, we get(μˆK(h),StRS)=[ω1Y‾h(1+e0)+ω2(D1h−D1h(1+e1))+ω3(D2h−D2h(1+e2))+Y‾h4(D1hD1h(1+e1)+D1h(1+e1)D1h)(D2hD1h(1+e1)+D1h(1+e1)D2h)]

(μˆK(h),StRS)=[ω1Y‾h+ω1Y‾he0−ω2D1he1+ω3D2he2(−D2h(1+e2))+Y‾h4(D1hD1h(1+e1)+D1h(1+e1)D1h)(D2hD2h(1+e2)+D2h(1+e2)D2h)]

(μˆK(h),StRS)=[ω1Y‾h+ω1Y‾he0−ω2D1he1+ω3D2he2+Y‾h4(1(1+e1)+(1+e1))(1(1+e2)+(1+e2))]

(μˆK(h),StRS)=[ω1Y‾h+ω1Y‾he0−ω2D1he1+ω3D2he2+Y‾h4((1+e1)−1+(1+e1))((1+e2)−1+(1+e2))]

(μˆK(h),StRS)=[ω1Y‾h+ω1Y‾he0−ω2D1he1+ω3D2he2+Y‾h4((1−e1+e12……)+(1+e1))((1−e2+e22……)+(1+e2))]

Using the Taylor series and neglecting the high-order terms we get,(μˆK(h),StRS)=[ω1Y‾h+ω1Y‾he0−ω2D1he1+ω3D2he2+Y‾h4(4+2e12+2e22)]

(μˆK(h),StRS)=[ω1Y‾h+ω1Y‾he0−ω2D1he1+ω3D2he2+Y‾h+12Y‾he12++12Y‾he22]

(μˆK(h),StRS−Y‾h)=[ω1Y‾h+ω1Y‾he0−ω2D1he1+ω3D2he2+12Y‾he12+12Y‾he22]

Taking expectations on both sidesE(μˆK(h),StRS−Y‾h)=[ω1Y‾h+ω1Y‾h(e0)−ω2D1h(e1)+ω3D2h(e2)+12Y‾hE(e12)+12Y‾hE(e22)]

Using notations, we getBias(μˆK(h),StRS)=[ω1Y‾h+12Y‾hλhCd1h2+12Y‾hλhCd2h2]

Bias(μˆK(h),StRS)=ω1Y‾h+12Y‾hλh[Cd1h2+Cd2h2]

Hence, the Bias expression of μˆK,StRS is obtained as follows:Bias(μˆK(h),StRS)=∑h=1Nδhδ2Bias[μˆK(h),StRS]

Bias(μˆK(h),StRS)=∑h=1Nδh2δ2(ω1Y‾h+12Y‾hλh[Cd1h2+Cd2h2])

Derivation of MSE[μˆK(h),StRS] is as follows:

Taking expectation and square on both sidesE(μˆK(h),StRS−Y‾h)2=[ω1Y‾h+ω1Y‾hE(e0)−ω2D1E(e1)+ω3D2E(e2)+12Y‾hE(e12)+12Y‾hE(e22)]=[(ω1Y‾h)2+2(ω1Y‾h)(Y‾hω1e0)+2(ω1Y‾h)(−ω2D1e1)+2(ω1Y‾h)(−ω3D2e2)+2(ω1Y‾h)(12Y‾he22)+(Y‾hω1e0)2+2(ω1Y‾he0)(−ω2D1e1)+2(ω1Y‾he0)(−ω3D2e2)+2(ω1Y‾he0)(12Y‾he22)+(−ω2D1e1)2+2(−ω2D1e1)(−ω3D2e2)+2(−ω2D1e1)(12Y‾he22)+(−ω3D2e2)2+(−ω3D2e2)(12Y‾he22)+(12Y‾he22)2+2(12Y‾he12)(ω1Y‾h)+2(12Y‾he12)(Y‾hω1e0)+2(12Y‾he12)(−ω2D1e1)+2(12Y‾he12)(−ω3D2e2)+2(12Y‾he12)(12Y‾he22)+(12Y‾he12)2]

Apply expectations and ignoring high order terms, we get=[ω12Y‾h2+2ω12Y‾h2E(e0)−2Y‾hω1ω2D1E(e1)−2Y‾hω1ω3D2E(e2)+ω1Y‾h2E(e22)+ω12Y‾h2E(e02)+2Y‾hω1ω2D1E(e0e1)+2Y‾hω1ω3D2E(e0e2)+ω22D12E(e12)+2ω2ω3D1D2E(e1e2)+ω32D22E(e22)+Y‾h2ω1E(e12)]

By using notation, we get(A) MSE(μˆK(h),StRS)=[ω12Y‾h2+ω1Y‾h2λCd2h2+ω12Y‾h2λhCyh2−2Y‾hω1ω2D1hλρyhd1hCyhCd1h−2Y‾hω1ω3D2hλhρyhd2hCyhCd2h+ω22D1h2λhCd1h2+2ω2ω3D1hD2hλhρd1hd2hCd1hCd2h+ω32D2h2λhCd2h2+Y‾h2ω1λhCd1h2]

Partially differentiating MSE(μˆK(h),StRS) and with respect to " ω1" and equating the result to zero gives the maximum value of " ω1". We obtain the equations for ω1:∂MSE(μˆK(h),StRS)∂ω1=0

gives,Y‾h22ω1+Y‾hλhCd2h2+Y‾h2Cyh2λh2ω1−2Y‾hω2D1hλhρyhd1hCyhCd1h−2Y‾hω3D2hλhρyhd2hCyhCd2h+Y‾h2λhCd1h2=0

2Y‾h2ω1(1+Cyh2λh)=2Y‾hω2D1hλhρyhd1hCyhCd1h+2Y‾hω3D2hλhρyhd2hCyhCd2h−Y‾h2λhCd1h2

ω1=Y‾h[ω2D1hλhρyhd1hCyhCd1h+2ω3D2hλhρyhd2hCyhCd2h−Y‾hλhCd1h2]Y‾h2ω1(1+Cyh2λh)

(1) ω1=λh[ω2D1hρyhd1hCyhCd1h+ω3D2hρyhd2hCyhCd2h−Y‾hCd1h2]Y‾h(1+Cyh2λh)

Partially differentiating MSE(μˆK(h),StRS) and with respect to " ω2" and equating the result to zero gives the maximum value of " ω2". We obtain the equations for ω2:∂MSE(μˆK(h),StRS)∂ω2=0

gives,−2Y‾hω1D1hλhρyhd1hCyhCd1h+D1h2λhCd1h22ω2+2ω3D1hD2hλhρyhd1hd2hCd1hCd2h=0

2λhD1h2Cd1h2ω2=2Y‾hω1D1λhρyhd1hCyhCd1h−2ω3D1hD2hλhρyhd1hd2hCd1hCd2h

ω2=2λhD1hCd1h(Y‾hω1ρyhd1hCyh−2ω3D2hρyhd2hCd2h)2λhD1h2Cd1h2

(2) ω2=Y‾hω1ρyhd1hCyh−2ω3D2hρyhd2hCd2hD1hCd1h

Partially differentiating MSE(μˆK(h),StRS) and with respect to " ω3" and equating the result to zero gives the maximum value of " ω3". We obtain the equations for ω3:∂MSE(μˆK(h),StRS)∂ω3=0

gives,−2Y‾hω1D2λhρyhd2hCyhCd2h+2ω2D1hD2hλhρyhd2hCd1hCd2h+2ω3D2h2Cd2h2=0

2ω3D2h2Cd2h2=2Y‾hω1D2hλhρyhd2hCyhCd2h−2ω2D1hD2hλhρyhd2hCd1hCd2h

ω3=2D2hλhCd2h(Y‾hω1ρyhd2hCyh−2ω2D1hλhρyhd2hCd1h)2D2h2Cd2h2

(3) ω3=Y‾hω1ρyhd2hCyh−2ω2D1hλhρyhd2hCd1hD2Cd2

Now find the optimum values using equation (3) in equation 2ω2=1D1hCd1h[Y‾hω1D1hρyhd1hCyhCd1h−D1hD2hρyhd1hd2hCd1hCd2h[1D2hCd2h(Y‾hω1ρyhd2hCyh−ω2D1hλhρyhd2hCd1h)]]

ω2=1D1hCd1hY‾hρyhd1hCyhCd1h−1D1hCd1hD2hρyhd2hCd1hCd2h−Y‾hD2hCd2hω1ρyhd2hCyh+D2hρyhd1hd2hCd2hD1hD2hCd1hCd2hD1hρyhd1hd2hCd1hω2

[1−(ρyhd1hd2h)2]ω2=ω1Y‾hCyhD1hCd1h[ρyhd1h−ρyhd1hd2hρyhd2h]

(4) ω2=Y‾hCyhD1hCd1h[ρyhd1h−ρyhd1hd2hρyhd2h][1−(ρyhd1hd2h2)]ω1

Now using equation (2) in equation 3ω3=1D2hCd2h[Y‾hω1ρyhd2hCyh−D1hλhρyhd2hCd1h[1D1hCd1h(Y‾hCyhρyhd1hω1−D2hρyhd1hd2hCd2hω3)]]

ω3=Y‾hρyhd2hCyhω1D2hCd2h−D1hρyd1hCd1hD1hCd1hD2hCd2hY‾hρyhd1hCyhω1+D1hρyhd1hd2hCd1hD1hD2hCd1hCd2hD2hρyhd1hd2hCd2hω3

[1−(ρyhd1hd2h2)]ω3=Y‾hCyhD2hCd2h[ρyhd2h−ρyhd1hd2hρyhd1h]ω1

(5) ω3=Y‾hCyhD2hCd2hρyhd2h−ρyhd1hd2hρyhd1h[1−(ρyhd1hd2h2)]ω1

By using equations (4), (5)) in equation 1ω1hY‾h(1+Cyh2λh)=λhD1hρyhd1hCyhCd1h[Y‾hCyhD1hCd1hρyhd1h−ρd1hd2hρyhd2h1−ρyhd1hd2h2]ω1+λhD2hρyhd2hCyhCd2h[Y‾hCyhD2Cd2ρyhd2h−ρd1hd2hρyhd1h1−ρyhd1hd2h2]ω1−Y‾hCd1h2

ω1Y‾h(1+Cyh2λh)=λhρyhd1hCyh21−ρyhd1hd2h2[ρyhd1h(ρyhd1h−ρd1hd2hρyhd2h)+ρyhd2h(ρyhd2h−ρd1hd2hρyhd1h)]ω1−Y‾hCd1h2

ω1Y‾h(1+Cyh2λh)=λhρyhd1hCyh21−ρydh1hd2h2[ρyhd1h2−2ρyhd1hρyhd2hρd1hd2h+ρyhd2h2]ω1−Y‾hCd1h2

ω1Y‾h(1+Cyh2λh)−λhρyhd1hCyh21−ρyhd1hd2h2[ρyhd1h2−2ρyhd1hρyhd2hρd1hd2h+ρyhd2h2]ω1=−Y‾hCd1h2

ω1Y‾h[1+Cyh2λh−λhCyh2βh(ρyhd1h2−2ρyhd1hρyhd2hρd1hd2h+ρyhd2h2)]=−Y‾hCd1h2

Where βh=1−ρyhd1hd2h2.Y‾hCd12=ω1Y‾h[1+Cyh2λh(1−1βh(ρyhd1h2+ρyhd2h2−2ρyhd1hρyhd2hρd1hd2h))]

Y‾hCd1h2=ω1Y‾h[−1+Cyh2λhαh]

Where αh=−1+(ρyhd1h2+ρyhd2h2−2ρyhd1hρyhd2hρd1hd2h)βh.(6) ω1*=Cd1h2−1+Cyh2λhαh

By using equation (6) in equation 4ω2=Y‾hCyhD1hCd1h[ρyhd1h−ρyhd1hd2hρyhd2h][1−(ρyhd1hd2h2)]Cd1h2−1+Cyh2λhαh

ω2=Y‾hCyhCd1hD1hβh(−1+Cyh2λhαh)ρyhd1h−ρyhd1hd2hρyhd2h

(7) ω2*=Y‾hCyhCd1hγhD1βh(−1+Cyh2λhαh)

Where γh=ρyhd1h−ρyhd1hd2hρyhd2h.

By using equation (6) in equation 5ω3=Y‾hCyhD2hCd2h[ρyhd2h−ρyhd1hd2hρyhd1h][1−(ρyhd1hd2h2)]Cd1h2−1+Cyh2λhαh

(8) ω3*=Y‾hCyhCd1h2ηhD2hCd2hβh(−1+Cyh2λhαh)

Where ηh=ρyhd2h−ρyhd1hd2hρyhd1h.

Now putting the optimum values of ω1*,ω2*,andω3* in equation AMSE(μˆK(h),StRS)=[ω12Y‾h2+ω12Y‾h2λhCyh2+Y‾h2ω1λhCd1h2+ω1Y‾h2λhCd2h2−2Y‾hω1ω2D1hλhρyhd1hCyhCd1h−2Y‾hω1ω3D2hλhρyhd2hCyhCd2h+2ω2ω3D1hD2hλhρd1hd2hCd1hCd2h+ω22D1h2λCd1h2+ω32D2h2λhCd2h2]

Where MSE(μˆK(h),StRS)=[A+B+C−D−E+F+G+H+I].

Now A and Bω12Y‾h2+ω12Y‾h2λhCyh2

Y‾h2(1+λhCyh2)ω12

Y‾h2(1+λhCyh2)Cd1h2−1+Cyh2λhαh

Now C and DY‾h2ω1λhCd1h2+ω1Y‾h2λhCd2h2

Y‾h2λh(Cd1h2+Y‾h2λhCd2h2)Cd1h2−1+Cyh2λhαh

Now E2Y‾hD1hλhρyhd1hCyhCd1hω1ω2

2Y‾hD1hλhρyhd1hCyhCd1hCd1h2−1+Cyh2λhαhY‾hCyhCd1hγhD1hβh(−1+Cyh2λhαh)

2Y‾hλhCyh2Cd1h4βh(Cyh2λhαh−1)2ρyhd1hγh

Now F2Y‾hD2hλhρyhd2hCyhCd2hω1ω3

2Y‾hD2hλhρyhd2hCyhCd2hCd1h2−1+Cyh2λhαhY‾hCyhCd1hηhD2hCd2hβh(−1+Cyh2λhαh)

2Y‾h2λhCyh2Cd1h4βh(Cyh2λhαh−1)2ρyhd2hCd1h2ηh

Now G2D1hD2hλhρd1hd2hCd1hCd2hω2ω3

2D1hD2hλhρd1hd2hCd1hCd2hY‾hCyhCd1hγhD1hβh(−1+Cyh2λhαh)Y‾hCyhCd1hηhD2hCd2hβh(−1+Cyh2λhαh)

2Y‾h2λhCyh2Cd1h4Cd2hβh2(Cyh2λhαh−1)2ρyhd1hd2hCd1h2γhηh

Now HD1h2λhCd1h2ω2h2

D1h2λhCd1h2(Y‾hCyhCd1hγhD1hβh(−1+Cyh2λhαh))2

D1h2λhCd1h2Y‾h2Cyh2Cd1h2γh2D1h2βh2(Cyh2λhαh−1)2

λhCd1h2Y‾h2Cyh2γh2βh2(Cyh2λhαh−1)2

Now ID2h2λhCd2h2ω32

D2h2λhCd2h2(Y‾hCyhCd1hηhD2hCd2hβh(−1+Cyh2λhαh))2

D2h2λhCd2h2Y‾h2Cyh2Cd1h4ηh2D2hCd2hβh(Cyh2λhαh−1)2

λhY‾h2Cyh2Cd1h4ηh2βh2(Cyh2λhαh−1)2

Therefore, equation AMSE(μˆK(h),StRS)=Y‾h2(1+λhCyh2)Cd1h2−1+Cyh2λhαh+Y‾h2λh(Cd1h2+Y‾h2λhCd2h2)Cd1h2−1+Cyh2λhαh−2Y‾hλhCyh2Cd1h4βh(Cyh2λhαh−1)2ρyhd1hγh+2Y‾h2λhCyh2Cd1h4βh(Cyh2λhαh−1)2ρyhd2hCd1h2ηh+2Y‾h2λhCyh2Cd1h4Cd2hβh2(Cyh2λhαh−1)2ρyhd1hd2hCd1h2γhηh+λhCd1h2Y‾h2Cyh2γh2βh2(Cyh2λhαh−1)2+λhY‾h2Cyh2Cd1h4ηh2βh2(Cyh2λhαh−1)2

MSE(μˆK(h),StRS)=Y‾h2Cd1h2Cyh2λhαh[1+λhCyh2+λhCd1h2+λhCd2h2]−2Y‾h2λhCyh2Cd1h4βh(Cyh2λhαh−1)2⌈ρyhd1hγh−ρyhd2hηh⌉+Y‾h2λhCyh2Cd1h4(γh2−ηh)2βh2(Cyh2λhαh−1)2+2Y‾h2λhCyhCd1h4Cd2hγhηhβh2(Cyh2λhαh−1)2ρyhd1hd2h

MSE(μˆK(h),StRS)=Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)−2λhCyh2Cd1h2(ρyhd1hγh−ρyhd2hηh)+λhCyh2Cd1h2(γh2−ηh)2+2CyhCd1h2Cd2hλhγhηhρyhd1hd2h]

MSE(μˆK(h),StRS)=Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+λhCyh2Cd1h2(−2βh(ρyhd1hγh−ρyhd2hηh)+(γh2−ηh)2)+2CyhCd1h2Cd2hλhγhηhρyhd1hd2h]

MSE(μˆK(h),StRS)=Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+σhλhCyh2Cd1h2+CyhCd1h2Cd2hλhξh]

Where σh=−2βh(ρyhd1hγh−ρyhd2hηh)+(γh2−ηh)2,ξ=2γhηhρyhd1hd2h.

Hence, the mean square error expression of μˆK,StRS is given by,MSE(μˆK,StRS)=∑h=1Nδ2δ2MSE(μˆK(h),StRS)

MSE(μˆK,StRS)=∑h=1Nδh2δ2Y‾h2Cd1h2βh2(Cyh2λhαh)2[βh2(Cyh2λhαh−1)(1+λhCyh2+λhCd1h2+λhCd2h2)+σhλhCyh2Cd1h2+CyhCd1h2Cd2hλhξh]

Acknowledgments

10.13039/501100004242 Princess Nourah bint Abdulrahman University Researchers Supporting Project number (PNURSP2024R515 ), 10.13039/501100004242 Princess Nourah bint Abdulrahman University , Riyadh, Saudi Arabia, their authors extend Their appreciation to the 10.13039/501100023674 Deanship of Scientific Research at King Khalid University for funding this work through Large Research Groups Project under grant number (RGP.2/41/45 ) and the authors extend their appreciation to the Deanship of Scientific Research at Northern Border University, Arar, KSA for funding this research work through the project number NBU-FFR-2024-1324-03 .
==== Refs
References

1 Cochran W.G. Sampling Techniques 1977 john wiley & sons
2 Cochran W. The estimation of the yields of cereal experiments by sampling for the ratio of grain to total produce J. Agric. Sci. 30 2 1940 262 275
3 Shabbir J. Gupta S. Estimation of the finite population mean in two phase sampling when auxiliary variables are attributes Hacettepe Journal of Mathematics and Statistics 39 1 2010 121 129
4 Abd-Elfattah A. El-Sherpieny E. Mohamed S. Abdou O. Improvement in estimating the population mean in simple random sampling using information on auxiliary attribute Appl. Math. Comput. 215 12 2010 4198 4202
5 Singh R. Chauhan P. Sawan N. Smarandache F. Ratio estimators in simple random sampling using information on auxiliary attribute Auxiliary Information and a priori Values in Construction of Improved Estimators 1 2007 7
6 Koyuncu N. Kadilar C. Family of estimators of population mean using two auxiliary variables in stratified random sampling Commun. Stat. Theor. Methods 38 14 2009 2398 2417
7 Singh H.P. Kumar S. A regression approach to the estimation of the finite population mean in the presence of non‐response Aust. N. Z. J. Stat. 50 4 2008 395 408
8 Ekpenyong E.J. Enang E.I. Efficient exponential ratio estimator for estimating the population mean in simple random sampling Hacettepe Journal of Mathematics and Statistics 44 3 2015 689 705
9 Lu J. Efficient estimator of a finite population mean using two auxiliary variables and numerical application in agricultural, biomedical, and power engineering Math. Probl Eng. 2017 2017
10 Zaman T. Kadilar C. Novel family of exponential estimators using information of auxiliary attribute J. Stat. Manag. Syst. 22 8 2019 1499 1509
11 Ahmad S. Arslan M. Khan A. Shabbir J. A generalized exponential-type estimator for population mean using auxiliary attributes PLoS One 16 5 2021 e0246947
12 Mahajan J. Banal K. Mahajan S. Estimation of crop production using machine learning techniques: a case study of J&K Int. J. Inf. Technol. 13 4 2021 1441 1448
13 Kumar A. Saini M. A predictive approach for finite population mean when auxiliary variables are attributes Thailand Statistician 20 3 2022 575 584
14 Audu A. Abdulazeez S. Danbaba A. Ahijjo Y. Gidado A. Yunusa M. Modified classes of regression-type estimators of population mean in the presence of auxiliary attributes Asian Research Journals of Mathematics 18 1 2022 65 89
15 Rather K.U.I. Jeelani M.I. Rizvi S. Sharma M. A new exponential ratio type estimator for the population means using information on auxiliary attribute J Stat Appl Pro Lett 9 2022 31 42
16 Zaman T. Sagir M. Şahin M. A new exponential estimators for analysis of COVID‐19 risk Concurrency Comput. Pract. Ex. 34 10 2022 e6806
17 Wayangkau I.H. Mekiuw Y. Rachmat R. Suwarjono S. Hariyanto H. Utilization of IoT for soil moisture and temperature monitoring system for onion growth Emerg Sci J 4 2020 102 115
18 Waheeb R, Wheib K, Andersen B, Alsuhili R. The Prospective of Artificial Neural Network (ANN's) Model Application to Ameliorate Management of Post Disaster Engineering Projects. Available at: SSRN 4208375 2022..
19 Rajak A.A. Emerging technological methods for effective farming by cloud computing and IoT Emerging Science Journal 6 5 2022 1017 1031
20 Jabal Z.K. Khayyun T.S. Alwan I.A. Impact of climate change on crops productivity using MODIS-NDVI time series Civil Engineering Journal 8 6 2022 1136 1156
21 Saini M. Jitendrakumar B.R. Kumar A. Optimum estimator in simple random sampling using two auxiliary attributes with application in agriculture, fisheries and education sectors MethodsX 9 2022 101915
22 Neyman J. On the two different aspects of the representative method: the method of stratified sampling and the method of purposive selection Breakthroughs in Statistics 1992 Springer 123 150
23 Bhushan S. Kumar A. Shahab S. Lone S.A. Akhtar M.T. On efficient estimation of the population mean under stratified ranked set sampling J. Math. 2022 2022
24 Ahmad S. Hussain S. Aamir M. Yasmeen U. Shabbir J. Ahmad Z. Dual use of auxiliary information for estimating the finite population mean under the stratified random sampling scheme J. Math. 2021 2021 1 12
25 Javed M. Irfan M. Bhatti S.H. Onyango R. A simulation-based study for progressive estimation of population mean through traditional and nontraditional measures in stratified random sampling J. Math. 2021 2021 1 16
26 Ma P. Mahoney M. Yu B. A statistical perspective on algorithmic leveraging International Conference on Machine Learning 2014 PMLR 91 99
27 Ma P. Chen Y. Zhang X. Xing X. Ma J. Mahoney M.W. Asymptotic analysis of sampling estimators for randomized numerical linear algebra algorithms J. Mach. Learn. Res. 23 177 2022 1 45
28 Naik V. Gupta P. A note on estimation of mean with known population proportion of an auxiliary character Jour Ind Soc Agr Stat 48 2 1996 151 158
29 Kumar S. Bhougal S. Estimation of the population mean in presence of non-response Communications for Statistical Applications and Methods 18 4 2011 537 548
30 Kadilar C. Cingi H. Ratio estimators in stratified random sampling Biom. J. 45 2 2003 218 225
31 Särndal C.-E. Swensson B. Wretman J. Model Assisted Survey Sampling 2003 Springer Science & Business Media
32 Gérard M. Jayet H. Paty S. Tax interactions among Belgian municipalities: do interregional differences matter? Reg. Sci. Urban Econ. 40 5 2010 336 342
