
==== Front
PLoS One
PLoS One
plos
PLOS ONE
1932-6203
Public Library of Science San Francisco, CA USA

10.1371/journal.pone.0307607
PONE-D-24-07265
Research Article
Physical Sciences
Mathematics
Statistics
Physical Sciences
Mathematics
Probability Theory
Probability Distribution
Physical Sciences
Mathematics
Probability Theory
Probability Distribution
Normal Distribution
Research and Analysis Methods
Mathematical and Statistical Techniques
Mathematical Functions
Research and Analysis Methods
Research Design
Survey Research
Surveys
Biology and Life Sciences
Computational Biology
Population Modeling
Biology and Life Sciences
Population Biology
Population Modeling
Social Sciences
Economics
Labor Economics
Salaries
Medicine and Health Sciences
Oncology
Cancers and Neoplasms
Exploring mixture estimators in stratified random sampling
Exploring Mixture Estimators in Stratified Random Sampling
Iqbal Kanwal Conceptualization Data curation Formal analysis Investigation Methodology Software Validation Visualization Writing – original draft Writing – review & editing 1 2
Raza Syed Muhammad Muslim Conceptualization Investigation Methodology Project administration Resources Software Supervision Writing – original draft Writing – review & editing 1 3
https://orcid.org/0000-0002-8748-5949
Mahmood Tahir Conceptualization Funding acquisition Supervision Validation Visualization Writing – original draft Writing – review & editing 4 *
Riaz Muhammad Conceptualization Investigation Methodology Validation Writing – review & editing 5
1 Department of Economics and Statistics, Dr Hasan Murad School of Management (HSM), University of Management and Technology, Lahore, Pakistan
2 Department of Mathematics and Statistics, University of Lahore, Sargodha Campus, Sargodha, Pakistan
3 Department of Statistics, Virtual University of Pakistan, Lahore, Pakistan
4 School of Computing, Engineering and Physical Sciences, University of the West of Scotland, Paisley, United Kingdom
5 Department of Mathematics, College of Computing and Mathematics, King Fahd University of Petroleum and Minerals, Dhahran, Saudi Arabia
Abonazel Mohamed R. Editor
Cairo University, EGYPT
Competing Interests: The authors have declared that no competing interests exist.

* E-mail: tahir.mahmood@uws.ac.uk
17 9 2024
2024
19 9 e030760722 2 2024
9 7 2024
© 2024 Iqbal et al
2024
Iqbal et al
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Advancements in sensor technology have brought a revolution in data generation. Therefore, the study variable and several linearly related auxiliary variables are recorded due to cost-effectiveness and ease of recording. These auxiliary variables are commonly observed as quantitative and qualitative (attributes) variables and are jointly used to estimate the study variable’s population mean using a mixture estimator. For this purpose, this work proposes a family of generalized mixture estimators under stratified sampling to increase efficiency under symmetrical and asymmetrical distributions and study the estimator’s behavior for different sample sizes for its convergence to the Normal distribution. It is found that the proposed estimator estimates the population mean of the study variable with more precision than the competitor estimators under Normal, Uniform, Weibull, and Gamma distributions. It is also revealed that the proposed estimator follows the Cauchy distribution when the sample size is less than 35; otherwise, it converges to normality. Furthermore, the implementation of two real-life datasets related to the health and finance sectors is also presented to support the proposed estimator’s significance.

The author(s) received no specific funding for this work. Data AvailabilityAll the relevant data are within the manuscript.
Data Availability

All the relevant data are within the manuscript.
==== Body
pmcIntroduction

Sampling is a procedure of selecting a representative fraction of a population so that one may observe and estimate something about the characteristics of interest for the entire population [1, 2]. A primary objective of survey sampling is to achieve a practical design for the survey to attain an adequate sample size for estimating the parameters of interest for the population under study. Survey sampling has several advantages over a full population study, including lower resource consumption, shorter turnaround times, and lower costs [3]. Moreover, it also provides a basis for acquiring precise and useful parameter estimations.

Additionally, survey sampling makes generating accurate and efficient parameter estimates easier. Survey statisticians can improve an estimator’s efficiency by enhancing the sampling technique, boosting the sample size, or employing auxiliary data. Auxiliary information about the population can consist of a known variable to which the study variable is approximately related. Typically, this auxiliary information is easy to quantify, whereas measuring the study variable itself can be costly [4]. By using this additional data, which includes characteristics and variables that are highly correlated with the variable of interest, the estimation process can improve the accuracy of the study variable’s mean [5, 6]. While measuring the study variable can often be expensive, the auxiliary information is usually easy to quantify. The ith element of Y may be correlated with one or more auxiliary variables (Xi) in addition to the study variable Y. The average elevation, area, and type of vegetation in a cattle field are examples of auxiliary variables that may be included if the study variable is the number of animals in the field. Auxiliary data is used in survey sampling for three main purposes: pre-selection, selection (i.e., selecting units that correspond to the study variable based on the auxiliary variable), and estimation (i.e., creating estimators of the ratio, product, and regression type).

In many studies, survey statisticians have been keenly interested in estimating parameters for the heterogeneous population. Therefore, for the heterogeneous population, Neyman [7] introduced Stratified Random Sampling (StRS), in which the population is partitioned into groups called "strata." Then a sample is chosen by some pattern within each stratum, and independent selections are made in different groups. Stratification is a probability sampling design used to increase the precision of estimation [8]. StRS’s sampling methodology ensures that every demographic group is adequately represented in the sample. For small sample sizes, adequate precision and accuracy may be achieved by using StRS’s procedure. In sample selection, the StRS design helps to minimize bias. From an organizational point of view, stratified sampling is very convenient. Moreover, it is also recommended that auxiliary information be used in the StRS population parameter estimation process, which is useful for comparing estimates among several population groups. For more recent studies on StRS, see studies [9–11] and references therein.

Historically, numerous researchers have independently utilized auxiliary variables and attributes, proposing diverse estimators within the StRS framework to enhance efficiency [12]. Cochran [13] developed classical ratio and regression methods to calculate the study variable’s population mean. Graunt [14] was the first known user of the ratio estimator. When estimating using a ratio, Cochran first used auxiliary information. To calculate the population mean, Robson [15] suggested product type estimators by incorporating the ancillary data. Kadilar and Cingi [16] proposed the estimator when the population coefficient of skewness and kurtosis are unknown in stratified random sampling. In the estimation of population mean three cases of using auxiliary information have been suggested by Samiuddin and Hanif [17], such as no information, full information, and partial information. Ahmad, Hanif [18] suggested a modified and efficient estimator of the population mean using two auxiliary variables in survey sampling. Ahmad, Hanif [19] established a generalized multi-phase multivariate regression estimator with the help of several auxiliary variables. To find the population mean, Moeen, Shahbaz [20] suggested mixture estimators by simultaneously utilizing auxiliary variables and attributes. Double-phase and multi-phase sampling of a vector of variables of interest are used to calculate the population mean. Malik and Singh [21] also worked on the stratified sampling estimator. In single-phase sampling, Verma, Sharma [22] suggested some modified regression-cum-ratio, and exponential ratio type estimators. The improved version of the exponential estimator of the mean under StRS, initially proposed by Zaman [23], is presented by Singh, Ragen [24], and its properties for big sample approximation. Singh, Ragen [24] also suggested a class of estimators of population variance along with their asymptotic properties. Zaman [25] established a class of ratio-type estimators with the help of auxiliary attributes to calculate the population’s mean. Under the StRS design, mixture regression cum ratio estimators in a single-phase scheme were established by Moeen [26]. Yadav and Zaman [27] suggested a class of estimators using both conventional and non-conventional auxiliary variable parameters. Under StRS design, using an auxiliary attribute, Zaman and Kadilar [28] suggested exponential ratio estimators for stratified two-phase sampling. Lawson and Thongsak [29] introduce a novel set of population mean estimators designed for use in stratified random sampling. Their study examines the bias and mean square error of these estimators using Taylor series approximation. Through simulation and application to air pollution data in northern Thailand, they evaluate the performance of the estimators. Results from the air pollution data show that the proposed estimators outperform others in terms of efficiency.

Many sampling survey investigations seek to devise an efficient estimator for the population mean. Numerous studies have been designed to pursue this aim, incorporating various adaptations to classical ratio, product, and regression estimators utilizing SRS and StRS. Kadilar and Cingi [16] and Zaman [25] employed auxiliary variables and attributes, respectively, proposing different estimators within the StRS framework. However, these estimators prove beneficial only in specific scenarios, such as when auxiliary variables or attributes are solely utilized to estimate the population mean of the study variable. Consequently, a gap exists in the literature concerning the simultaneous utilization of auxiliary attributes and auxiliary variables alongside the study variable to estimate the population parameter. For example, in a household survey, income, expenditures, family size, number of employed, and number of literate persons are related variables that can be used as auxiliary information for estimating any characteristic. So, from the above example, we can use income (a quantitative variable) and family size (a qualitative variable) simultaneously to estimate the expenditure (a study variable). Therefore, the current article aims to propose a class of generalized mixture ratio estimators to estimate the population mean of the study variable by simultaneously incorporating the auxiliary attributes (qualitative) and variables (quantitative) in stratified random sampling (StRS). Therefore, the suggested estimators could be used in various sampling surveys.

The subsequent sections of the paper are organized as follows: Section: “Notations under Stratified Random Sampling” offers an overview of stratified random sampling. Section: “A Family of Proposed Estimators under Stratified Random Sampling” introduces the proposed estimators. In Section: “A Simulation Study”, a comparative analysis is conducted based on the simulation study. Section: “Illustrative Examples” showcases two real-life examples. Finally, Section: “Conclusion and Recommendations” outlines the conclusions drawn from the study and provides recommendations for future research.

Notations under stratified random sampling

Let N denote the size of the population, N = N1 + N2 + N3 +…‥ + Nh+.… + Nk, where the number of units in stratum h is represented by Nh, n = n1 + n2 + n3+…‥ +nh +….+nk, where the number of sampling units in stratum h is represented by nh, Y stands for the study variable, X refers to the auxiliary variable, P is the population proportion auxiliary attribute, and X and Y have a strong correlation. The Y¯h refers to the population mean of the variable of interest Y in stratum h, X¯h is the population mean of auxiliary variable X in stratum h, Ph is the population proportion in auxiliary attribute in stratum h. X¯⌢h and Y¯h are the sample means of Y and X in stratum h. P^h is the sample proportion in auxiliary attribute in stratum h Wh=NhN,Wh2=Nh2N2 is the stratum weight. The Syh2 and Sxh2 are the variances of stratum h, Cyh and Cxh are the coefficients of variations in stratum h. ρxyh is the correlation coefficient between X and Y in stratum h, β2(xh) and β1(xh) are the kurtosis and skewness in stratum h and Cxyh=ρxyhCxhCyh Let′s say that an attribute j = 1,2,3,…,m is dichotomized in the population based on its presence or absence; the attribute’s values should be "0" and "1" correspondingly.

τijh=1,0,ifhthunitofthepopulationpossessesattributeotherwise

Under stratified random sampling, we take into account the following notations in order to determine the biases and mean square error (MSE) of the suggested estimator: e0=y¯h−Y¯hY¯h,y¯h=Y¯h1+e0,e3=pUh−PUhPUh,pUh=PUh1+e3,e4=x¯Uh−x¯UhX¯Uh,x¯Uh=X¯Uh(1+e4), Ee0=Ee3=Ee4=0,Ee02=∑h=1kωh2λhSyh2Y¯h2,Ee32=∑h=1kωh2λhSph2PUh2,Ee42=∑h=1kωh2λhSxh2X¯Uh2, Ee0e3=∑h=1kωh2λhSyphY¯h⋅PUh,Ee0e4=∑h=1kωh2λhSyxhY¯h⋅X¯Uh,Ee3e4=∑h=1kωh2λhsxphPUhX¯Uh,andλh=1nh−1Nh.

Relevant estimators under StRS

This section discusses some existing estimators and their bias and mean square error (MSE). Kadilar and Cingi [16] proposed the following estimators: μ^yKS1,StRS=y¯st∑h=1kωhX¯h+Cxh∑h=1kωhX¯^h+Cxh. (1)

MSE(μ^yKS1,StRS)=∑h=1kωh2λh[Syh2−2RSsyxh+R2Sxh2]. (2)

μ^yKS2,StRS=y¯st∑h=1kωhX¯h+β(2h)x∑h=1kωhX¯^h+β(2h)x. (3)

MSE(μ^yKS2,StRS)=∑h=1kωh2λh[Syh2−2RSsyxh+R2Sxh2]. (4)

μ^yKS3,StRS=y¯st∑h=1kωhX¯hCxh+β(2h)x∑h=1kωhX¯^hCxh+β(2h)x. (5)

MSE(μ^yKS3,StRS)=∑h=1kωh2λh[Syh2−2RSsyxh+R2Sxh2]. (6)

μ^yKS4,StRS=y¯st∑h=1kωhX¯hβ(2h)x+Cxh∑h=1kωhX¯^hβ(2h)x+Cxh. (7)

MSE(μ^yKS4,StRS)=∑h=1kωh2λh[Syh2−2RSsyxh+R2Sxh2]. (8)

In stratified random sampling by using the auxiliary attribute, Zaman [25] proposed the following estimators: μ^yZ1,StRS=y¯st∑h=1kωhPh+Cph∑h=1kωhP^h+Cph. (9)

MSE(μ^yZ1,StRS)=∑h=1kωh2λh[Syh2−2RSsyph+R2Sph2]. (10)

μ^yZ2,StRS=y¯st∑h=1kωhPh+β2h(ϑ)∑h=1kωhP^h+β2h(ϑ). (11)

MSE(μ^yZ2,StRS)=∑h=1kωh2λh[Syh2−2RSsyph+R2Sph2]. (12)

μ^yZ3,StRS=y¯st∑h=1kωhPh+ρyph∑h=1kωhP^h+ρyph. (13)

MSE(μ^yZ3,StRS)=∑h=1kωh2λh[Syh2−2RSsyph+R2Sph2]. (14)

μ^yZ4,StRS=y¯st∑h=1kωhPhCph+β2h(ϑ)∑h=1kωhP^hCph+β2h(ϑ). (15)

MSE(μ^yZ4,StRS)=∑h=1kωh2λh[Syh2−2RSsyph+R2Sph2]. (16)

μ^yZ5,StRS=y¯st∑h=1kωhPhβ2h(ϑ)+Cph∑h=1kωhP^hβ2h(ϑ)+Cph. (17)

MSE(μ^yZ5,StRS)=∑h=1kωh2λh[Syh2−2RSsyph+R2Sph2]. (18)

A family of proposed estimators under stratified random sampling

Within this section, we introduce a novel class of generalized mixture estimator, building upon the framework proposed by Zaman [25]; this estimator will be suitable for the estimation of the population mean Y¯, incorporating the concurrent utilization of auxiliary variables and attributes within a stratified random sampling context. TKMst=∑h=1kNhNaTph+(1−a)Txh, (19)

where Tph=y¯hPU,Txh=y¯hX¯U, (20)

PU=PhK1+K2andX¯U=X¯hK3+K4pU=P^hK1+K2andX¯^U=X¯^hK3+K4. (21)

Given that where a and K are constant, with a taking on the values 0 and 1 and K ∈ ℝ. Consequently, K1,K2,K3,and K4 may consist of any real number. Eq (19) can be rewritten as follows using Eq (20): TKMst=∑h=1kNhNay¯hPUP^U+(1−a)y¯hX¯UX¯^U. (22)

However, for the hthstratum, let TKMh be a mixture estimator of the population mean given as follows: TKMh=ay¯hPUP^U+(1−a)y¯hX¯UX¯^U. (23)

Derivation of BiasTKMh and MSETKMh

Similarly, we can rewrite Eq (23) as follows by using y¯h as a common factor: TKMh=y¯haPUP^U+(1−a)X¯UX¯^U, (24)

Using notations from preceding section, the bias of TKMh is derived as follows, and after simplification, we obtain, TKMh=Y¯h1+e0aPUhP^Uh1+e3+(1−a)X¯UhX¯^Uh1+e4.

Similarly, using notations from preceding section and after simplification, we get, Bias[TKMh]=Y¯hSxh2X¯Uh2−SyhxhX¯Uh+aY¯hSph2PUh2−SyhphPUh−Y¯hSxh2X¯Uh2+SyhxhX¯Uh. (25)

Hence, the Bias expression of TKMst is obtained as follows: BiasTKMst=∑h=1kNh2N2BiasTKMh

BiasTKMst=∑h=1kNh2N2Y¯hSxh2X¯Uh2−SyhxhX¯Uh+a∑h=1kNh2N2Y¯hSph2PUh2−SyhphPUh−Y¯hSxh2X¯Uh2+SyhxhX¯Uh. (26)

and MSETKMh=λhSyh2+Y¯h2Sxh2X¯Uh2−2Y¯hSyxhX¯Uh+a2λhY¯h2Sp2PUh2+Y¯h2Sx2X¯Uh2−2Y¯h2SxpX¯Uh.PUh+2aλhY¯hSyxhX¯Uh−Y¯hSyphPUh+Y¯h2SxphX¯Uh.PUh−Y¯h2Sxh2X¯Uh2. (27)

Moreover, the following simplification can be used to determine the maximum value of " a ": a=X¯Uh.PUh−SyxhPUh+SyphX¯Uh+Y¯hPUh−SxphX¯Uh+Sxh2PUhY¯h[Sph2X¯Uh2+Sxh2PUh2+2SxphX¯Uh.PUh]. (28)

Adding the value of a in Eq (27) and after simplification, we obtained, MSETKMh=λhSyh2+Y¯h2Sxh2X¯Uh2−2Y¯hSyxhX¯Uh+λh−SyxhPUh+SyphX¯Uh+Y¯hX¯Uh−SxphX¯Uh+Sxh2PUh2Sph2X¯Uh2+Sxh2PUh2+2SxphX¯Uh.PUh. (29)

Hence, the mean square error (MSE) expression of TKMst is obtained as follows: MSETKMst=∑h=1kNh2N2MSETKMh

MSETKMst=∑h=1kNh2N2λhSyh2+Y¯h2Sxh2X¯Uh2−2Y¯hSyxhX¯Uh+λh−SyxhPUh+SyphX¯Uh+Y¯hX¯Uh−SxphX¯Uh+Sxh2PUh2Sph2X¯Uh2+Sxh2PUh2+2SxphX¯Uh.PUh. (30)

The complete derivation of Bias(TKMh) and MSE(TKMh) given in Eqs (25–29) are provided in S1 Appendix.

Unique cases of the proposed mixture estimator

In this section, we will delve into specific cases of the proposed mixture estimator, examining various combinations of constants. While Tables 1 and 2 highlight certain special cases of the estimator, exploring additional combinations of constants can reveal further special cases not covered in the tables.

10.1371/journal.pone.0307607.t001 Table 1 Particular situations for the proposed mixture estimator for a = 0.

K 1	K 2	K 3	K 4	Estimators	
β2(ph)	C ph	β2(xh)	β1(ph)	Ts1=y¯h∑h=1kWhX¯hβ2xh+β1phX¯^hβ2xh+β1ph	
β2(ph)s	C ph	β2(xh)	σ h	Ts2=y¯h∑h=1kWhX¯hβ2xh+σhX^¯hβ2xh+σh	
β2(ph)	C ph	β2(xh)	ρ h	Ts3=y¯h∑h=1kWhX¯hβ2xh+ρhX¯^hβ2xh+ρh	
β2(ph)	C ph	β2(xh)	β1(xh)	Ts4=y¯h∑h=1kWhX¯hβ2xh+β1xhX¯¯hβ2xh+β1xh	
β2(ph)	C ph	Cxh	β2(xh)	Ts5=y¯h∑h=1kWhX¯Cxh+β2xhX¯^hCxh+β2xh	
β2(ph)	C ph	β1(ph)	β2(xh)	Ts6=y¯h∑h=1kWhX¯β1ph+β2xhX¯^hβ1ph+β2xh	
β2(ph)	C ph	σ h	β2(xh)	Ts7=y¯h∑h=1kWhX¯σh+β2xhX^¯hσh+β2xh	
β2(ph)	C ph	β1(xh)	β2(xh)	TS8=y¯h∑h=1kWhX¯1xh+β2xhX¯^hβ1xh+β2xh	

10.1371/journal.pone.0307607.t002 Table 2 Particular situations for the proposed mixture estimator for a = 1.

K 1	K 2	K 3	K 4	Estimators	
Cxh	β2(xh)	β2(ph)	C ph	Ts9=y¯h∑h=1kWhPhCxh+β2xhP^Cxh+β2xh	
β1(ph)	β2(xh)	β2(ph)	C ph	TS10=y¯h∑h=1kWhPhβ1ph+β2xhP^β1ph+β2xh	
σh	β2(xh)	β2(ph)	C ph	Ts11=y¯h∑h=1kWhPhσh+β2xhP^σh+β2xn	
ρh	β2(xh)	β2(ph)	C ph	TS12=y¯h∑h=1kWhPρhρh+β2xhP^ρh+β2xh	
β1(xh)	β2(xh)	β2(ph)	C ph	Ts13=y¯h∑h=1kWhPhβ1xh+β2xnP^β1xh+β2xh	
β2(xh)	C x	β2(ph)	C ph	Ts14=y¯h∑h=1kWhPβ2xh+CxhP^β2xn+Cxh	
β2(xh)	β1(p)	β2(ph)	C ph	TS15=y¯h∑h=1kWhPβ2xh+β1phP^β2xn+β1ph	
β2(xh)	σ	β2(ph)	C ph	TS16=y¯h∑h=1kWhPβ2xh+σhP^β2xh+σh	
β2(xh)	β1(x)	β2(ph)	C ph	TS17=y¯h∑h=1kWhPβ2xh+β1xhP^β2xh+β1xh	

Theoretical comparison

In this area, efficiency conditions are established by comparing the MSE of the suggested estimators with numerous existing estimators:

Referring to Eqs (2) and (29), MSEμ^yKS1,StRS>MSETKMst if and only if ∑h=1kωh2λh[Syh2−2RSsyxh+R2Sxh2]>∑h=1kNh2N2λhSyh2+Y¯h2Sxh2X¯Uh2−2Y¯hSyxhX¯Uh+λh−SyxhPUh+SyphX¯Uh+Y¯hX¯Uh−SxphX¯Uh+Sxh2PUh2Sph2X¯Uh2+Sxh2PUh2+2SxphX¯Uh.PUh. (31)

Referring to Eqs (4) and (29), MSEμ^yKS2,StRS>MSETKMst if and only if ∑h=1kωh2λh[Syh2−2RSsyxh+R2Sxh2]>∑h=1kNh2N2λhSyh2+Y¯h2Sxh2X¯Uh2−2Y¯hSyxhX¯Uh+λh−SyxhPUh+SyphX¯Uh+Y¯hX¯Uh−SxphX¯Uh+Sxh2PUh2Sph2X¯Uh2+Sxh2PUh2+2SxphX¯Uh.PUh. (32)

Referring to Eqs (6) and (29), MSEμ^yKS3,StRS>MSETKMst if and only if ∑h=1kωh2λh[Syh2−2RSsyxh+R2Sxh2]>∑h=1kNh2N2λhSyh2+Y¯h2Sxh2X¯Uh2−2Y¯hSyxhX¯Uh+λh−SyxhPUh+SyphX¯Uh+Y¯hX¯Uh−SxphX¯Uh+Sxh2PUh2Sph2X¯Uh2+Sxh2PUh2+2SxphX¯Uh.PUh. (33)

Referring to Eqs (8) and (29), MSEμ^yKS4,StRS>MSETKMst if and only if ∑h=1kωh2λh[Syh2−2RSsyxh+R2Sxh2]>∑h=1kNh2N2λhSyh2+Y¯h2Sxh2X¯Uh2−2Y¯hSyxhX¯Uh+λh−SyxhPUh+SyphX¯Uh+Y¯hX¯Uh−SxphX¯Uh+Sxh2PUh2Sph2X¯Uh2+Sxh2PUh2+2SxphX¯Uh.PUh. (34)

Referring to Eqs (10) and (29), MSEμ^yZ1,StRS>MSETKMst if and only if ∑h=1kωh2λh[Syh2−2RSsyph+R2Sph2]>∑h=1kNh2N2λhSyh2+Y¯h2Sxh2X¯Uh2−2Y¯hSyxhX¯Uh+λh−SyxhPUh+SyphX¯Uh+Y¯hX¯Uh−SxphX¯Uh+Sxh2PUh2Sph2X¯Uh2+Sxh2PUh2+2SxphX¯Uh.PUh. (35)

Referring to Eqs (12) and (29), MSEμ^yZ2,StRS>MSETKMst if and only if ∑h=1kωh2λh[Syh2−2RSsyph+R2Sph2]>∑h=1kNh2N2λhSyh2+Y¯h2Sxh2X¯Uh2−2Y¯hSyxhX¯Uh+λh−SyxhPUh+SyphX¯Uh+Y¯hX¯Uh−SxphX¯Uh+Sxh2PUh2Sph2X¯Uh2+Sxh2PUh2+2SxphX¯Uh.PUh. (36)

Referring to Eqs (14) and (29), MSEμ^yZ3,StRS>MSETKMst if and only if ∑h=1kωh2λh[Syh2−2RSsyph+R2Sph2]>∑h=1kNh2N2λhSyh2+Y¯h2Sxh2X¯Uh2−2Y¯hSyxhX¯Uh+λh−SyxhPUh+SyphX¯Uh+Y¯hX¯Uh−SxphX¯Uh+Sxh2PUh2Sph2X¯Uh2+Sxh2PUh2+2SxphX¯Uh.PUh. (37)

Referring to Eqs (16) and (29), MSEμ^yZ4,StRS>MSETKMst if and only if ∑h=1kωh2λh[Syh2−2RSsyph+R2Sph2]>∑h=1kNh2N2λhSyh2+Y¯h2Sxh2X¯Uh2−2Y¯hSyxhX¯Uh+λh−SyxhPUh+SyphX¯Uh+Y¯hX¯Uh−SxphX¯Uh+Sxh2PUh2Sph2X¯Uh2+Sxh2PUh2+2SxphX¯Uh.PUh. (38)

Referring to Eqs (18) and (29), MSEμ^yZ5,StRS>MSETKMst if and only if ∑h=1kωh2λh[Syh2−2RSsyph+R2Sph2]>∑h=1kNh2N2λhSyh2+Y¯h2Sxh2X¯Uh2−2Y¯hSyxhX¯Uh+λh−SyxhPUh+SyphX¯Uh+Y¯hX¯Uh−SxphX¯Uh+Sxh2PUh2Sph2X¯Uh2+Sxh2PUh2+2SxphX¯Uh.PUh. (39)

A simulation study

We conducted a comprehensive simulation study to assess the proposed estimator’s effectiveness. This involved comparing the performance of our estimator with that of several alternative estimators under StRS conditions. The percent relative efficiency (PRE) served as the criterion for evaluating estimator performance. We followed the steps outlined below to compute the PREs for our proposed estimator under StRS.

A simulated population comprising 1500 observations is established for the study variable Y. X serves as an auxiliary variable, while P represents an auxiliary attribute. Random selections are made from Normal and Uniform (symmetric), Gamma, and Weibull (asymmetric) distributions across 10,000 samples of specified sizes. Table 3 outlines the methodology for utilizing the Bernoulli distribution with specific parameter values to generate the auxiliary attribute. Moreover, population size is considered as N = N1 + N2 = 800 + 700 = 1500.

The given proportional allocation formula is used to determine the sample size from the stratum.

nh/n=Nh/Nornh⇌nNNh=nWh

Using StRS, the population was divided into two strata: 800 and 700. Further, 10,000 random samples of size 20 are drawn by taking 12 units from stratum one and 8 units from stratum two. Next, 10,000 random samples of size 50 are drawn by taking 30 units from stratum one and 20 units from stratum two. Similarly, using the proportional allocation scheme, the next 10,000 random samples of size 80 are obtained by selecting 50 units from stratum one and 30 units from stratum two. Next, 10,000 random samples of size 200 are drawn by taking 120 units from stratum one and 80 units from stratum two. Further next, 10,000 random samples of size 400 are drawn by taking 280 units from stratum one and 120 units from stratum two. The results are shown in Table 4.

The PREs of estimators are calculated by using the following expression:

PRE=MSEμ^y,StRSMSEμ^y,i×100 (40)

where a^y,StRS=∑i=1kNhNy¯hx¯hX¯ and “i” stands for the estimator, whose performance needs to be compared. To check the efficiency of proposed estimator the expression of MSE given in Eq (30) has been utilized.

10.1371/journal.pone.0307607.t003 Table 3 Simulated auxiliary variables and attributes.

	Distribution Parameters	Normal	Uniform	Gamma	Weibull	
τ	Y	X	τ	Y	X	τ	Y	X	τ	Y	X	
Stratum-I N1h = 800	n	1			1			1			1			
P	0.5			0.3			0.5			0.7			
Shape					0	0.1		2	2		2	2	
Scale		600	1200		1	1.1		1	1		1	1	
Location		550	850										
Stratum-II N2h = 700	n	1			1			1			1			
P	0.5			0.7			0.5			0.3			
Shape					0	0.1		2.5	2.5		4	4	
Scale		400	1000		1	1.1		1.5	1.5		1	1	
Location		500	900										

10.1371/journal.pone.0307607.t004 Table 4 Two homogeneous subgroups of the population.

Prop. Allocation	Stratum-I	Stratum-II	
n1h = 12,30,50,120,280	n2h = 8,20,30,80,120	
N h	800	700	
W h	800/1500	700/1500	
nW h	20*800/1500, 50*800/1500, 80*800/1500, 120*800/1500, 280*800/1500	8*700/1500, 20*700/1500, 30*700/1500, 80*700/1500, 120*700/1500	

Tables 5 and 6 present the PREs of the proposed estimator in comparison to those of competing estimators. We also assessed the impact of sample size on MSEs and employed sample sizes of 20, 50,80, 200, and 400. Under Normal distribution, for n = 20,50,80,200 and 400,the proposed generalized mixture estimator TKMst the highest PRE’s are 102.98, 103.35, 104.46, 106.21, and 107.31 respectively. Similarly, under the Uniform distribution, for n = 20,50,80,200 and 400, the proposed generalized mixture estimator TKMst the highest PREs were reported as 102.92, 103.56,103.91, 105.56, and 106.23 respectively. Moreover, similar results are evident in Table 6, particularly under the Gamma and Weibull distributions, the proposed generalized mixture estimator TKMst demonstrates significantly superior PRE values when compared to the other estimators under consideration. Additionally, it was noted that the percent relative efficiency (PRE) values demonstrated an increase as sample sizes increased, as illustrated in Fig 1.

10.1371/journal.pone.0307607.g001 Fig 1 A comparison between the proposed estimator and competitive estimators under the following distributions.

(A) normal, (B) uniform, (C) Gamma, and (D) Weibull.

10.1371/journal.pone.0307607.t005 Table 5 Proposed generalized mixture estimators’ percent relative efficiencies (PREs) and comparative estimators’ PREs for Normal and Uniform distributions.

Normal	Uniform	
Estimators	n = 20	n = 50	n = 80	n = 200	n = 400	n = 20	n = 50	n = 80	n = 200	n = 400	
μ^y,StRS	100	100	100	100	100	100	100	100	100	100	
μ^yKC1,StRS	100.64	100.69	100.77	101.23	101.98	100.50	100.55	100.64	101.11	101.65	
μ^yKC2,StRS	100.11	100.31	100.37	100.76	101.23	100.23	100.31	100.44	100.82	101.34	
μ^yKC3,StRS	100.40	100.49	100.52	100.89	101.34	100.17	100.27	100.39	100.77	101.11	
μ^yKC4,StRS	101.37	101.48	101.55	101.99	102.01	100.62	100.75	100.77	101.31	101.99	
μ^yZ1,StRS	100.13	100.20	100.36	101.21	101.56	100.69	100.76	100.80	101.24	102.01	
μ^yZ2,StRS	100.56	100.65	100.68	100.87	101.56	100.93	101.13	101.66	102.14	102.76	
μ^yZ3,StRS	101.45	101.53	101.60	101.97	102.45	100.80	100.91	100.98	101.43	101.88	
μ^yZ4,StRS	101.54	101.63	101.69	102.21	102.65	101.79	101.96	102.00	102.22	102.56	
μ^yZ5,StRS	100.49	100.59	100.70	101.02	101.45	100.34	100.56	100.62	101.04	101.76	
TKMst	102.98	103.35	104.46	106.21	107.31	102.92	103.56	103.91	105.56	106.23	
Ts1	100.45	100.76	100.89	101.35	101.99	100.32	100.43	101.64	101.98	102.16	
Ts2	100.65	100.68	100.76	101.02	102.21	100.44	100.76	100.90	101.21	101.98	
Ts3	100.49	100.70	100.76	101.12	101.96	100.55	100.73	100.79	101.13	101.87	
Ts4	100.35	100.44	100.65	101.02	101.56	100.48	100.51	100.73	101.22	101.86	
Ts5	100.44	100.63	100.75	101.22	101.86	100.34	100.62	100.82	101.09	101.78	
Ts6	100.65	100.79	100.81	101.31	101.97	100.52	100.77	100.96	101.44	101.89	
Ts7	100.33	100.49	100.57	100.97	101.44	100.21	100.53	100.63	101.17	101.95	
Ts8	100.39	100.56	100.68	101.04	101.68	100.46	100.69	100.95	101.54	102.01	
Ts9	100.79	100.85	101.29	101.87	102.68	100.81	100.98	101.35	101.95	102.65	
Ts10	100.65	100.75	100.77	101.11	101.77	100.33	100.65	100.81	101.24	101.84	
Ts11	100.53	100.69	100.74	101.24	101.95	100.41	100.57	100.83	101.31	101.92	
Ts12	100.61	100.77	100.95	101.35	102.34	100.57	100.70	100.99	101.56	102.21	
Ts13	100.34	100.79	100.88	101.54	102.01	100.38	100.66	100.78	101.33	101.83	
Ts14	100.67	100.94	101.54	101.89	102.55	100.71	100.93	101.45	101.99	102.67	
Ts15	100.44	100.85	100.86	101.21	101.88	100.48	100.90	101.00	101.76	102.65	
Ts16	100.42	100.68	100.81	101.08	101.79	100.53	100.72	100.90	101.27	102.43	
Ts17	100.58	100.67	100.83	101.23	101.98	100.46	100.77	100.99	101.67	102.32	

10.1371/journal.pone.0307607.t006 Table 6 Proposed generalized mixture estimators’ percent relative efficiencies (PREs) and comparative estimators’ PREs for Gamma and Weibull distributions.

Normal	Uniform	
Estimators	n = 20	n = 50	n = 80	n = 200	n = 400	n = 20	n = 50	n = 80	n = 200	n = 400	
μ^y,StRS	100	100	100	100	100	100	100	100	100	100	
μ^yKC1,StRS	100.65	100.70	100.79	101.21	101.87	100.42	100.49	100.52	101.15	101.84	
μ^yKC2,StRS	100.32	100.39	100.42	101.03	101.81	100.38	100.41	100.45	100.97	101.43	
μ^yKC3,StRS	100.58	100.70	100.71	101.11	101.64	100.92	101.00	101.20	101.98	102.31	
μ^yKC4,StRS	100.63	100.75	100.81	101.07	101.77	100.95	100.99	101.18	101.88	102.65	
μ^yZ1,StRS	100.71	100.79	100.87	101.32	101.83	100.88	100.90	100.93	101.65	102.03	
μ^yZ2,StRS	100.19	100.35	100.44	100.97	101.32	100.67	100.71	100.77	101.34	102.06	
μ^yZ3,StRS	100.67	100.80	100.88	101.18	101.91	100.65	100.68	100.72	101.18	101.84	
μ^yZ4,StRS	100.77	101.79	101.82	102.11	102.76	101.36	101.47	101.51	102.12	102.67	
μ^yZ5,StRS	100.78	100.84	100.89	101.15	101.46	100.22	100.34	100.55	101.06	101.76	
TKMst	101.67	101.99	102.31	104.44	105.23	101.54	101.59	101.92	103.76	104.87	
Ts1	100.61	100.77	100.95	101.23	101.91	100.55	100.78	100.98	101.54	102.38	
Ts2	100.34	100.79	100.88	101.33	101.90	100.43	100.71	100.92	101.33	102.01	
Ts3	100.67	100.74	100.86	101.17	101.87	100.57	100.80	100.98	101.54	102.41	
Ts4	100.44	100.85	100.86	101.34	102.00	100.38	100.77	100.99	101.65	102.18	
Ts5	100.42	100.68	100.81	101.29	101.93	100.41	100.63	100.87	101.47	102.43	
Ts6	100.58	100.67	100.83	101.41	102.01	100.54	100.73	100.93	101.29	102.30	
Ts7	100.45	100.76	100.89	101.27	101.92	100.51	100.67	100.91	101.36	102.22	
Ts8	100.65	100.68	100.76	100.99	101.59	100.58	100.81	100.90	101.56	102.19	
Ts9	101.18	101.32	101.42	101.89	102.74	101.09	101.27	101.39	102.06	102.87	
Ts10	100.35	100.44	100.65	101.55	102.04	100.37	100.53	100.82	101.66	102.15	
Ts11	100.44	100.63	100.75	101.34	101.99	100.39	100.69	100.83	101.27	101.99	
Ts12	100.65	100.79	100.81	101.45	102.65	100.68	100.82	100.96	101.67	102.43	
Ts13	100.33	100.49	100.57	101.43	102.32	100.23	100.44	100.81	101.53	102.44	
Ts14	100.59	100.81	101.23	101.98	102.65	100.61	100.88	101.41	101.98	102.65	
Ts15	100.49	100.70	100.76	101.44	102.03	100.32	100.53	100.88	101.73	102.34	
Ts16	100.65	100.75	100.77	101.31	102.33	100.43	100.68	100.74	101.54	102.11	
Ts17	100.53	100.69	100.74	101.24	101.96	100.42	100.72	100.87	101.38	102.28	

Exploring the Best-Fitted distribution

This section explores the most appropriate probability distribution for the proposed generalized mixture estimator across various sample sizes using EasyFit [30]. EasyFit employs two methods, namely the Kolmogorov-Smirnov and Anderson-Darling tests, for this purpose. The auxiliary variables are generated from a normal distribution, while the auxiliary attribute is generated from a binomial distribution using R language (version 4.2.2), with parameters specified in Table 7. We select one thousand samples of sizes n = 20, n = 50, and n = 80 and then construct the sampling distribution of the proposed estimator.

10.1371/journal.pone.0307607.t007 Table 7 Setting up additional variables and characteristics for the simulation study.

Variables and Attributes	N	Stratum-I	Stratum-II	
N1h = 800	N2h = 700	
	Parameters	
μ	σ	ρ xy	n	p	μ	σ	ρ xy	n	p	
Y	1500	550	600	0.99			500	400	0.88			
X	1500	850	1200	0.99			900	1000	0.88			
P	1500				2	0.5				2	0.5	

As depicted in Fig 2, the proposed generalized mixture estimator conforms to the Normal distribution, with the scale parameter and location parameter computed as 3.13 and 5.30, respectively.

10.1371/journal.pone.0307607.g002 Fig 2 Probability distribution of generalized proposed mixture estimator.

Exploring the distribution of the proposed estimator for different sample sizes

This section examines how the suggested generalized mixture estimator’s distribution behaves for different sample sizes.

Table 8 displays the probability distribution of the estimator for n = 20, 50, and 80. When the initial data is derived from the Normal distribution for a sample size of n = 20, the proposed estimator follows the Cauchy distribution, with the cutoff point being n = 35, for which the proposed estimator’s probability distribution is converged to be Normal. Hence, the proposed estimator follows the normal distribution for n = 50 and 80.

10.1371/journal.pone.0307607.t008 Table 8 Probability distribution of estimator for different sample sizes.

Distribution of the Variable	Sample Sizes	Distribution of Estimators	
Normal	20	Cauchy (n < 35)	Normal (n ≥ 35)	
50	Normal	
80	Normal	

Table 9 illustrates the probability distributions of the proposed estimator against each value of n. The Kolmogorov–Smirnov results indicate that the p-values exceed 5%, supporting the hypotheses (H0) that the data adhere to the specified distribution.

10.1371/journal.pone.0307607.t009 Table 9 Probability distribution of estimator for each sample size.

n	p-value (KS test)	Distribution of Estimator	n	p-value (KS test)	Distribution of Estimator	n	p-value (KS test)	Distribution of Estimator	
20	0.691	Cauchy	40	0.496	Normal	61	0.569	Normal	
21	0.662	Cauchy	42	0.661	Normal	62	0.211	Normal	
22	0.667	Cauchy	43	0.372	Normal	63	0.162	Normal	
23	0.692	Cauchy	44	0.451	Normal	64	0.544	Normal	
24	0.656	Cauchy	45	0.714	Normal	65	0.692	Normal	
25	0.545	Cauchy	46	0.685	Normal	66	0.734	Normal	
26	0.589	Cauchy	47	0.699	Normal	67	0.181	Normal	
27	0.576	Cauchy	48	0.127	Normal	68	0.812	Normal	
28	0.566	Cauchy	49	0.754	Normal	69	0.732	Normal	
29	0.432	Cauchy	50	0.671	Normal	70	0.541	Normal	
30	0.591	Cauchy	51	0.651	Normal	71	0.417	Normal	
31	0.612	Cauchy	52	0.161	Normal	72	0.568	Normal	
32	0.655	Cauchy	53	0.362	Normal	73	0.485	Normal	
33	0.212	Cauchy	54	0.324	Normal	74	0.721	Normal	
34	0.426	Cauchy	55	0.712	Normal	75	0.547	Normal	
35	0.525	Normal	56	0.621	Normal	76	0.167	Normal	
36	0.769	Normal	57	0.676	Normal	77	0.651	Normal	
37	0.616	Normal	58	0.269	Normal	78	0.611	Normal	
38	0.321	Normal	59	0.665	Normal	79	0.232	Normal	
39	0.263	Normal	60	0.455	Normal	80	0.423	Normal	

Illustrative examples

In this section, we present two real-life examples to illustrate the practical application of the proposed estimator. The distribution of study and auxiliary variables in each of the two data sets is discovered using EasyFit version 5.5 professional. The details of each of the data sets are given below.

Data-I: Tumor data

Data-I has been taken from Andersen, Borgan [31]. We are interested in estimating the average thickness of the tumor by including auxiliary information. The data consists of 205 entities. StRS has been used, and the given variables and attributes have been considered. Y: Thickness of tumor (mm), X: Age of patient (at operation time), and P: Gender (0 = Male, 1 = Female) are used as auxiliary variable and attribute. The variable ’whether a patient was ulcerated or not’ has been used for stratification purposes. Those patients who are not ulcerated are placed in stratum 1, while stratum 2 consists of the remaining patients. In Table 10, the parameters of the data have been presented.

10.1371/journal.pone.0307607.t010 Table 10 Descriptive details of study variable Y, auxiliary variable X, and attribute P.

Variables	Strata	Size	Mean	Variance	Minimum	Maximum	
Y	1	115	1.8113	4.778	0.1	14.66	
2	90	4.3363	10.3382	0.16	17.42	
X		205	52.463	277.95	4	95	
P		205	0.385	0.238	0	1	

The population of size 205 is split into two strata, with sizes 115 and 90, respectively, as Table 10 illustrates. The proportional allocation scheme has been utilized to select random samples of sizes 12, 30, 50, and 8, 20, 30, and PREs are computed. In stratum 1, the average thickness of the tumor is 1.8113, while in stratum 2, its value is 4.3433. The variations in the tumor thickness in strata 1 and 2 are 4.7364 and 10.4244, respectively. The average age is 52.463, and the average attribute gender is 0.385. The variation in age is 277.95, and in gender, it is 0.238.

Data-II: Wages data

In this example, the data set obtained from Bierens and Ginther [32] is about wages of employees in the United States of America (USA) is considered, where wage (in dollars per week) is the study variable Y, number of years of education is the auxiliary variable X, and living in the standard metropolitan statistical area (SMSA) (0 = Not lives, 1 = if lives) is used as the auxiliary attribute P. In Table 11, the parameters of the data have been presented.

10.1371/journal.pone.0307607.t011 Table 11 Descriptive details of study variable Y, auxiliary variable X, and attribute P.

Variables	Strata	Size	Mean	Variance	Minimum	Maximum	
Y	1	25631	640.2	197379.9	50.4	18777.2	
2	2524	233.73	139839.9	50.5	9259.26	
X		28155	13.067	8.408	0	18	
P		28155	0.7435	0.1907	0	1	

Table 11 shows that the original data set contained 28155 observations. The observations are divided into two strata of sizes 25631 and 2524, respectively. The proportional allocation scheme has been utilized to select random samples of sizes 12, 30, 50, and 8, 20, 30, and PREs is computed. In stratum 1, the average wage is 640.2, while in stratum 2, the average wage is 233.73. In strata 1, the variation in wage is 197379.9, and in strata 2, its value is 139839.9. The average number of years of education is 13.067, and the average SMSA is 0.7435. Finally, the variation in the number of years of education is 8.408, and in SMSA, the variation is 0.1907.

Tables 12 and 13 shows the necessary calculations (estimated means and PRE’s) for proposed and comparative estimators with respect to StRS for those two real-life datasets mentioned. We used 20, 50, and 80 sample sizes for each data set. In dataset I, TKMst with n = 20, has the highest PRE, coming in at 105.42. In a similar vein, for n = 50 and 80, TKMst additionally displays 105.98 and 105.99 as dominant PRE values, respectively. Similarly, the suggested generalized mixture estimator exhibits the highest PREs in dataset II. This implies that when compared to comparative estimators, the suggested estimator performs remarkably well and efficiently. Furthermore, it is clear that the PREs rise in tandem with larger sample sizes. In Fig 3, the suggested estimator’s performance is further illustrated graphically.

10.1371/journal.pone.0307607.g003 Fig 3 A comparison between the proposed estimator and competitive estimators for real-life datasets.

(A) Data-I and (B) Data-II.

10.1371/journal.pone.0307607.t012 Table 12 Estimated sample means and proposed generalized mixture estimators’ percent relative efficiencies (PREs) with comparative estimators’ PREs for data set I.

Estimators	Estimated Sample Means	Data I	
n = 20	n = 50	n = 80	n = 20	n = 50	n = 80	
μ^y,StRS	1.06	1.19	1.30	100	100	100	
μ^yKC1,StRS	1.61	1.64	1.69	101.35	101.84	102.12	
μ^yKC2,StRS	1.40	1.53	1.66	101.41	101.65	101.98	
μ^yKC3,StRS	1.51	1.59	1.65	101.86	101.9	101.99	
μ^yKC4,StRS	1.24	1.45	1.61	101.46	101.6	101.87	
μ^yZ1,StRS	1.11	1.40	1.59	101.37	101.62	101.89	
μ^yZ2,StRS	1.16	1.23	1.38	101.45	101.73	101.98	
μ^yZ3,StRS	1.20	1.39	1.41	101.14	101.38	101.88	
μ^yZ4,StRS	1.12	1.21	1.46	102.99	103.21	103.76	
μ^yZ5,StRS	1.08	1.30	1.43	101.82	101.91	101.98	
TKMst	1.67	1.71	1.78	105.42	105.98	105.99	

10.1371/journal.pone.0307607.t013 Table 13 Estimated sample means and proposed generalized mixture estimators’ percent relative efficiencies (PREs) with comparative estimators’ PREs for data set II.

Estimators	Estimated Sample Means	Data II	
n = 20	n = 50	n = 80	n = 20	n = 50	n = 80	
μ^y,StRS	635.43	636.44	637.19	100	100	100	
μ^yKC1,StRS	635.02	636.87	637.06	101.12	101.37	101.67	
μ^yKC2,StRS	636.32	636.52	637.13	101.25	101.41	101.74	
μ^yKC3,StRS	635.52	636.64	637.72	101.45	101.5	101.87	
μ^yKC4,StRS	635.20	636.17	637.12	101.1	101.34	101.79	
μ^yZ1,StRS	634.67	635.98	636.18	101.23	101.62	101.81	
μ^yZ2,StRS	633.42	636.70	637.03	101.32	101.41	101.52	
μ^yZ3,StRS	634.67	635.17	637.61	101.14	101.42	101.86	
μ^yZ4,StRS	633.26	635.11	637.87	101.87	101.89	101.92	
μ^yZ5,StRS	632.56	633.30	636.38	101.78	101.81	101.88	
TKMst	636.93	637.35	638.99	104.98	105.12	105.34	

Conclusion and recommendations

In different scenarios, using auxiliary information is a very useful strategy to improve the estimator’s efficiency. For this purpose, we have used auxiliary variables and attributes simultaneously. In this study, we introduce a generalized family of mixture estimators under Stratified Random Sampling (StRS), inspired by the work of Zaman [25], aimed at enhancing efficiency across symmetric and asymmetric distributions for estimating finite population means. Additionally, we analyze the estimator’s behavior across various sample sizes regarding its convergence to the Normal distribution. Our findings indicate that the proposed mixture estimator adheres to the Normal distribution for sample sizes greater than or equal to 35. Furthermore, we derive the Mean Squared Error (MSE) expressions for the proposed estimator, supported by simulation results. Notably, our proposed class of estimators outperforms competing estimators in terms of Percent Relative Efficiency (PRE). Through simulation studies and real-life applications in the health and finance sectors, we demonstrate that our proposed estimator consistently delivers superior results compared to competitors across Normal, Uniform, Weibull, and Gamma distributions. Ultimately, we conclude that the efficiency of the suggested estimator holds both theoretically and in practical settings. Moreover, we suggest extending the scope of this study to include other types of estimators, such as ratio, product, power, difference, exponential, and regression estimators, under stratified random sampling.

Supporting information

S1 Appendix (DOCX)

The author “Tahir Mahmood” is thankful to the University of the West of Scotland for providing research facilities for conducting this research and providing APC charges.
==== Refs
References

1 Cochran WG. William G. Cochran. New York: Wiley; 1977.
2 Hyder M , Raza SMM , Mahmood T , Moeen M , Riaz M . On moving average based location charts under modified successive sampling. Hacettepe Journal of Mathematics and Statistics. 2024;53 (2 ):506–23.
3 Martínez-Mesa J , González-Chica DA , Duquia RP , Bonamigo RR , Bastos JL . Sampling: how to select participants in my research study? Anais brasileiros de dermatologia. 2016;91 :326–30. doi: 10.1590/abd1806-4841.20165254 27438200
4 Mahmood T. Generalized linear modelling based monitoring methods for air quality surveillance. Journal of King Saud University-Science. 2024;36 (4 ):103145.
5 Amin M , Noor A , Mahmood T , Analytics. Beta regression residuals-based control charts with different link functions: an application to the thermal power plants data. International Journal of Data Science and Analytics. 2024:1–13. doi: 10.1007/s41060-023-00501-w
6 Aslam MZ , Amin M , Mahmood T , Nauman Akram M , Simulation. Shewhart ridge profiling for the Gamma response model. Journal of Statistical Computation and Simulation. 2023:1–20. doi: 10.1080/00949655.2023.2299354
7 Neyman J. On the two different aspects of the representative method: the method of stratified sampling and the method of purposive selection. Breakthroughs in statistics: Springer; 1992. p. 123–50.
8 Almanjahie IM , Ismail M , Cheema AN . Partial stratified ranked set sampling scheme for estimation of population mean and median. Plos one. 2023;18 (2 ):e0275340. doi: 10.1371/journal.pone.0275340 36791145
9 Bushan S , Kumar A , Shahab S , Lone SA , Akhtar MT . On Efficient Estimation of the Population Mean under Stratified Ranked Set Sampling. Journal of Mathematics. 2022.
10 Ahmad S , Hussain S , Aamir M , Yasmeen U , Javid Shabbir , Ahmad Z . Dual Use of Auxiliary Information for Estimating the Finite Population Mean under the Stratified Random Sampling Scheme. Journal of Mathematics. 2021.
11 Javed M , Irfan M , Bhatti SH , Onyango R . A Simulation-Based Study for Progressive Estimation of Population Mean through Traditional and Nontraditional Measures in Stratified Random Sampling. Journal of Mathematics. 2021.
12 Hussain S , Ahmad S , Saleem M , Akhtar S . Finite population distribution function estimation with dual use of auxiliary information under simple and stratified random sampling. Plos one. 2020;15 (9 ):e0239098. doi: 10.1371/journal.pone.0239098 32986764
13 Cochran W. The estimation of the yields of cereal experiments by sampling for the ratio of grain to total produce. The journal of agricultural science. 1940;30 (2 ):262–75.
14 Graunt J . Natural and political observations mentioned in a following index, and made upon the bills of mortality. Mathematical Demography: Springer; 1977. p. 11–20.
15 Robson D. Applications of multivariate polykays to the theory of unbiased ratio-type estimation. Journal of the American Statistical Association. 1957;52 (280 ):511–22.
16 Kadilar C , Cingi H . Ratio estimators in stratified random sampling. Biometrical Journal: Journal of Mathematical Methods in Biosciences. 2003;45 (2 ):218–25.
17 Samiuddin M , Hanif M , editors. Estimation in two phase sampling with complete and incomplete information. Proc 8th Islamic Countries Conference on Statistical Sciences; 2006.
18 Ahmad Z , Hanif M , Ahmad M . Generalized Regression-cum-ratio estimators for two-phase sampling using multi-auxiliary variables. Pakistan Journal of Statistics. 2009;25 (2 ):93–106.
19 Ahmad Z , Hanif M , Ahmad M . Generalized Multi-Phase Multivariate Ratio Estimators for Partial Information Case Using Multi-Auxiliary Vatiables. Communications for Statistical Applications and Methods. 2010;17 (5 ):625–37.
20 Moeen M , Shahbaz MQ , Hanif M . Mixture Ratio and Regression Estimators Using Multi-Auxiliary Variables and Attributes in Single Phase Sampling. World Applied Sciences Journal. 2012;18 (11 ):1518–26.
21 Malik S , Singh R . An improved estimator using two auxiliary attributes. Applied mathematics and computation. 2013;219 (23 ):10983–6.
22 Verma HK , Sharma P , Singh R . Some families of estimators using two auxiliary variables in stratified random sampling. Revista Investigacion Operacional. 2015;36 (2 ):140–50.
23 Zaman T. An efficient exponential estimator of the mean under stratified random sampling. Mathematical Population Studies. 2021;28 (2 ):104–21.
24 Singh HP , Ragen G , Tarray TA . Study of Yadav and Kadilar’s improved exponential type ratio estimator of population variance in two-phase sampling. Hacettepe Journal of Mathematics and Statistics. 2015;44 (5 ):1257–70.
25 Zaman T. Efficient estimators of population mean using auxiliary attribute in stratified random sampling. Advances and Applications in Statistics. 2019;56 (2 ):153–71.
26 Iqbal K , Moeen M , Ali A , Iqbal A . Mixture regression cum ratio estimators of population mean under stratified random sampling. Journal of Statistical Computation and Simulation. 2020;90 (5 ):854–68.
27 Yadav SK , Zaman T . Use of some conventional and non-conventional parameters for improving the efficiency of ratio-type estimators. Journal of statistics and Management systems. 2021;24 (5 ):1077–100.
28 Zaman T , Kadilar C . On estimating the population mean using auxiliary character in stratified random sampling. Journal of Statistics and Management Systems. 2020;23 (8 ):1415–26.
29 Lawson N , Thongsak N . A Class of Population Mean Estimators in Stratified Random Sampling: A Case Study on Fine Particulate Matter in the North of Thailand. WSEAS Transactions on Mathematics. 2024;23 :160–6.
30 EasyFit. MathWave; 2014. 5.5:[Available from: http://www.mathwave.com/products/easyfit.html.
31 Andersen PK , Borgan O , Gill RD , Keiding N . Statistical models based on counting processes: Springer Science & Business Media; 2012.
32 Bierens HJ , Ginther DK . Integrated conditional moment testing of quantile regression models. Empirical Economics. 2001;26 (1 ):307–24.
