
==== Front
Heliyon
Heliyon
Heliyon
2405-8440
Elsevier

S2405-8440(24)12795-8
10.1016/j.heliyon.2024.e36764
e36764
Research Article
A new two-parameter over-dispersed discrete distribution with mathematical properties, estimation, regression model and applications
Ahmadini Abdullah Ali H. a
Ahsan-ul-Haq Muhammad ahsanshani36@gmail.com
b⁎
Hussain Muhammad Nasir Saddam c
a Department of Mathematics, College of Science, Jazan University, Jazan, Saudi Arabia
b College of Statistical Sciences, University of the Punjab, Lahore, Pakistan
c Department of Statistics, Govt. Murray Graduate College Sialkot, Pakistan
⁎ Corresponding author. ahsanshani36@gmail.com
23 8 2024
15 9 2024
23 8 2024
10 17 e367643 2 2024
20 8 2024
21 8 2024
© 2024 The Authors. Published by Elsevier Ltd.
2024

https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
This paper focuses on the derivation of a new two-parameter discrete probability distribution. The new model is derived by mixing Poisson and Loai distributions and is named “Poisson Loai Distribution”. The paper explores various mathematical properties of the new model, introducing a count-regression model based on this distribution. The parameters of the model are estimated using the maximum likelihood estimation method. A comprehensive simulation study is utilized to assess the behavior of derived estimators. The importance of the proposed distribution is confirmed through the analysis of three real datasets. It is found that the suggested distribution has the greatest match when compared to all rival distributions, and it may be a viable alternative for assessing dispersed count data.

Keywords

Count data analysis
Dispersion
Moments
Estimation
Regression model
Data analysis
==== Body
pmc1 Introduction

The Poisson distribution stands as a prominent statistical model, particularly in the realm of modeling count data. Its one-parameter structure offers a straightforward representation of phenomena where events occur randomly and independently over a fixed interval. Notably, the Poisson distribution exhibits a distinct characteristic of the mean and variance being equal, a quality that, while often advantageous, becomes a limitation when faced with count data that deviates from this equality, presenting scenarios of over-dispersion or under-dispersion. Despite this inherent constraint, the Poisson distribution finds widespread application in diverse fields of investigation, including health sciences, support sciences, environmental sciences, actuarial studies, biology, business administration, and economics. Recognizing the need for more flexible models capable of accommodating situations where the equality of mean and variance may not hold, researchers have directed their attention toward the exploration of mixed-Poisson distributions.

Thus, numerous researchers have investigated the formation of mixed-Poisson distributions, such as Bhati et al. [1] introduced the Poisson-transmuted exponential model, and subsequently applied it to datasets related to healthcare. The Poisson-Ailamujia distribution was developed by Ref. [2] and used to analyze the biological datasets. The Poisson pseudo-Lindley distribution introduced by Ref. [3] and applied to the data sets of patients infected with ebola virus, further Poisson quasi-Xgamma distribution with INAR(1) constructed by Ref. [4], and consider it to analyze the earthquake data of turkey, Hassan et al. [5] suggested the Poisson Ishita distribution and used the new model on epileptic seizure counts dataset, Ahsan-ul-Haq et al. [6] originated Poisson XLindley distribution, they use different estimation techniques for parameter estimation of the new compound model and applied it on two different real-life datasets in the field of health and biology. Some more examples of Poisson mixture models are; Poisson Mirra [7], Binomial Poisson Lindley [8], Poisson transmuted Janardan [9], Poisson moment exponential [10], Poisson Lindley difference [11], Poisson entropy-based weighted exponential [12], Poisson Ramos Louzada [13], Bernoulli Poisson Lindley [14], Poisson new XLindley [15], Poisson Xgamma [16], Poisson Quasi-XLindley [17], and Poisson Quasi-Shanker [18]. These new compound Poisson probability distributions provide a sophisticated framework that goes beyond the limitations of the conventional Poisson model. In the realm of discrete models documented in current literature, there is always a need for new models to be developed. In light of this, we would like to suggest a flexible probability distribution that, in our opinion, would provide a better match for over-dispersed count data.

In this study, we will propose a new discrete probability model using a Poisson mixture with Loai distribution. The Loai distribution is a novel flexible two-parameter lifetime probability distribution proposed by Ref. [19]. It is obtained by mixing gamma distribution and Lindley distribution with parameters α=3, and θ respectively, with mixtire proportions (α1+α) and (11+α). The Loai distribution can be defined as:

If a random variable X has two parameters (α and θ) and is distributed according to the Loai distribution, the probability density function of the Loai distribution may be expressed as(1) f(x,α,θ)=θ2(1+α)(αθx22+1+x1+θ)e−θx;x,α,θ>0

This research endeavors to formulate a new distinctive two-parameter discrete distribution. The study has the following key objectives.• The main objective of this novel study was to propound a flexible two-parameter extension of Poisson distribution. This innovative model unites elements of the Poisson distribution with the unique features of the Loai distribution. This new count data model is termed the Poisson Loai (PLo) distribution.

• Derived its various key statistical properties including probability generating function, moment generating function, dispersion index, moments, Shannon, and Rényi entropies.

• Estimate the parameters of PLo distribution using the maximum likelihood estimation (MLE) approach. The behavior of derived MLEs is assessed via a comprehensive simulation study based on different sample sizes and different choices of parameters.

• Three datasets from different disciplines are utilized to show the applicability of the new distribution.

• Another key objective of the present study is to introduce a new count regression to analyze over-dispersed count datasets.

The structure of this document is as follows: Section 2 goes into depth about how the new distribution developed. In Section 3, we investigate and analyze some mathematical characteristics of the Poisson Loai (PLo) distribution. Section 4 describes a unique count regression model based on the PLo distribution. Section 5 provides parameter estimates as well as simulation investigation outcomes. Section 6 describes how to apply the new distribution and regression model to real-world data. Finally, Section 7 includes final observations.

2 Poisson Loai Distribution

Suppose that a random variable Z has a Poisson distribution and the parameter of the Poisson distribution follows the Loai distribution. Thus, the Poisson Loai distribution can be obtained by mixing the Poisson distribution model with the Loai distribution asp(Z=z)=∫0∞p(z|λ)g(λ,α,θ)dλ,

where p(z|λ) is the density of Poisson distribution and g(λ,α,θ) is the density of Loai distribution.

Preposition 1: A discrete random variable Z is distributed as the Poisson Loai distribution with parameters α and θ, and the pmf of PLo is given as(2) p(Z=z)=θ2(αθz2+2(2+θ+αθ)+z(2+3αθ))2(1+α)(1+θ)z+3,

for z=0,1,2… and α>−1,θ>0..Proof Let Z|λ∼Pois(λ) and λ∼Loai(α,θ), the pmf of Z is derived as

p(Z=z)=∫0∞e−λλzz!θ2(1+α)(αθλ2e−θλ2+(1+λ)(1+θ)e−θλ)dλ,

=θ2z!(1+α)(αθ2∫0∞λz+2e−λ(1+θ)dλ+1(1+θ)∫0∞λz+1e−λ(1+θ)dλ+1(1+θ)∫0∞λze−λ(1+θ)dλ),

=θ2z!(1+α)(αθΓ(z+3)2(1+θ)z+3+Γ(z+2)(1+θ)z+3+Γ(z+1)(1+θ)z+2),

=θ2(z2αθ+2(2+θ+αθ)+z(2+3αθ))2(1+α)(1+θ)z+3.

which is the PMF of PLo distribution.

The PMF graph in Fig. 1 depicts different shapes of probability mass functions for various parameter choices in the PLo distribution. The PMF plots show decreasing and unimodal behavior.Fig. 1 Behavior of PMF of PLo distribution for different values of parameters.

Fig. 1

The cumulative distribution function of the PLo distribution is expressed asF(z)=P(Z≤z)

F(z)=∑k=0zθ2(αθk2+2(2+θ+αθ)+k(2+3αθ))2(1+α)(1+θ)k+3,

(3) F(z)=2(1+θ)3+z−2θ(z+θ+3)−2α+2α(1+θ)3+z−(z+3)(2αθ+(2+z)αθ2)−22(1+α)(1+θ)z+3.

The survival function isS(z)=1−F(z)=(2+α(2+(3+z)θ(2+(2+z)θ))+2θ(3+z+θ))2(1+α)(1+θ)z+3.

The hazard function of PLo distribution is(4) h(z)=θ2(z2αθ+2(2+θ+αθ)+z(2+3αθ))2+2θ(3+z+θ)+α(2+(3+z)θ(2+(2+z)θ)).

The reverse hazard rate of PLo distribution is given as(5) h*(z)=θ2(z2αθ+2(2+θ+αθ)+z(2+3αθ))2(1+θ)3+z−2θ(z+θ+3)−2α+2α(1+θ)3+z−(z+3)(2αθ+(2+z)αθ2)−2.

The Mills ratio is as follows(6) MR=(2+α(2+(3+z)θ(2+(2+z)θ))+2θ(3+z+θ))θ2(z2αθ+2(2+θ+αθ)+z(2+3αθ))

The hazard function plots for different combinations of parameters are presented in Fig. 2. It is observed that the hazard function increases on random variable Z.Fig. 2 Behavior of Hazard function of PLo distribution for different values of parameters.

Fig. 2

3 Mathematical characteristics

In this section, some important mathematical properties are derived using equation (2) and discussed.

3.1 Factorial moments

The factorial moments function of the random variable Z can be calculated as follows:μr;′=E[E(X(r)|δ)]=∫0∞(∑0∞z(r)e−δδzz!)θ2(1+α)(αθx22+1+x1+θ)e−θxdδ,

where; z(r)=z(z−1)(z−2)...(z−r+1) and then:μ(r)′=∫0∞δr(∑y=r∞e−λλ(z−r)(z−r)!)θ2(1+α)(αθδ22+1+δ1+θ)e−θδdδ,

Using z+r instead of r,μ(r)′=θ2(1+α)∫0∞δr(∑y=0∞e−λλrz!)(αθδ22+1+δ1+θ)e−θδdδ,

After simplification(7) =θ2(1+α)(αθ2θr+3Γ(r+3)+11+θ(Γ(r+1)θr+1+Γ(r+2)θr+2)).

3.2 Probability generating function

The probability-generating function of the random variable Z can be articulated as follows:G(t)=∑z=0∞tzp(Z=z)=∑z=0∞tzθ2(z2αθ+2(2+θ+αθ)+z(2+3αθ))2(1+α)(1+θ)z+3,

=θ22(1+α)(αθ∑z=0∞tzz2(1+θ)z+3+2(2+θ+αθ)∑z=0∞tz(1+θ)z+3+(2+3αθ)∑z=0∞ztz(1+θ)z+3),

(8) =tαθ3(1+t+θ)+2θ2(2+θ+αθ)(1−t+θ)2+tθ2(2+3αθ)(1−t+θ)2(1+α)(1+θ)2(1−t+θ)3.

3.3 Moment generating function

Correspondingly the moment-generating function of Z is formulated as follows:M(t)=∑z=0∞etzp(Z=z)=∑z=0∞etzθ2(z2αθ+2(2+θ+αθ)+z(2+3αθ))2(1+α)(1+θ)z+3,

=θ22(1+α)(αθ∑z=0∞etzz2(1+θ)z+3+2(2+θ+αθ)∑z=0∞etz(1+θ)z+3+(2+3αθ)∑z=0∞zetz(1+θ)z+3),

(9) =etαθ3(1+et+θ)+2θ2(2+θ+αθ)(1−et+θ)2+etθ2(2+3αθ)(1−et+θ)2(1+α)(1+θ)2(1−et+θ)3.

The moments, variance, dispersion index, coefficient of skewness, and kurtosis are given as follows asμ1′=2+3α+θ+3αθθ(1+α)(1+θ),

μ2′=6+12α+4θ+15αθ+θ2+3αθ2(1+α)θ2(1+θ),

μ3′=24+60α+24θ+96αθ+8θ2+39αθ2+θ3+3αθ3(1+α)θ3(1+θ),

μ4′=120+360α+168θ+720αθ+78θ2+444αθ2+16θ3+87αθ3+θ4+3αθ4(1+α)θ4(1+θ),

The variance and dispersion index (DI) are formulated asμ2=2+6α+3α2+6θ+19αθ+9α2θ+4θ2+17αθ2+9α2θ2+θ3+4αθ3+3α2θ3((1+α)θ(1+θ))2,

andDI(z)=μ2μ1′=2+6α+3α2+6θ+19αθ+9α2θ+4θ2+17αθ2+9α2θ2+θ3+4αθ3+3α2θ3(2+θ+3α(1+θ))(1+α)θ(1+θ).

Furthermore, the coefficient of Skewness (SK) and Kurtosis (K) can be obtained using the following measures:SK=μ3′−3(μ1′)(μ2′)+2(μ1′)3(μ2′−(μ1′)2)32andK=μ4′−4μ3′μ1′+6μ2′(μ1′)2−3(μ1′)3(μ2′−(μ1′)2)2.

Some numerical values of the first four moments about the origin, variance (Var), DI, skewness, and kurtosis are presented in Table 1. The mean and variance increase with an increase in parameter α, while the DI, coefficient of skewness, and kurtosis measures show a decreasing pattern. On the other side, there is an opposite pattern with an increase in parameter θ.Table 1 The numerical values of various descriptive metrics depend on particular parameter selections.

Table 1Para.	Descriptive Measures	
θ	α	E(Z)	E(Z2)	E(Z3)	E(Z4)	Var(Z)	DI(Z)	SK(Z)	K(Z)	
0.5	0.5	4.222	32.667	345.56	4598.0	14.842	3.515	1.440	5.914	
1.0	4.667	38.000	416.67	5694.0	16.219	3.475	1.346	5.544	
1.5	4.933	41.200	459.33	6351.6	16.866	3.419	1.295	5.368	
2.0	5.111	43.333	487.78	6790.0	17.211	3.367	1.266	5.275	
2.5	5.238	44.857	508.09	7103.1	17.420	3.326	1.247	5.218	
1.0	0.5	2.000	8.667	52.000	396.67	4.667	2.334	1.587	6.459	
1.0	2.250	10.250	63.750	499.25	5.188	2.306	1.468	5.944	
1.5	2.400	11.200	70.800	560.80	5.440	2.267	1.404	5.699	
2.0	2.500	11.833	75.500	601.83	5.583	2.233	1.365	5.563	
2.5	2.571	12.286	78.857	631.14	5.676	2.208	1.337	5.474	
1.5	0.5	1.289	4.133	18.356	104.32	2.471	1.917	1.713	6.974	
1.0	1.467	4.933	22.711	132.52	2.781	1.896	1.577	6.343	
1.5	1.573	5.413	25.324	149.44	2.939	1.868	1.502	6.032	
2.0	1.644	5.733	27.067	160.72	3.030	1.843	1.456	5.857	
2.5	1.695	5.962	28.311	168.77	3.089	1.822	1.425	5.747	
2.0	0.5	0.944	2.500	9.111	42.667	1.609	1.704	1.820	7.436	
1.0	1.083	3.000	11.333	54.500	1.827	1.687	1.671	6.707	
1.5	1.167	3.300	12.667	61.600	1.938	1.661	1.591	6.355	
2.0	1.222	3.500	13.556	66.333	2.007	1.642	1.539	6.144	
2.5	1.262	3.643	14.190	69.714	2.050	1.625	1.505	6.015	
2.5	0.5	0.743	1.718	5.424	22.052	1.166	1.569	1.918	7.877	
1.0	0.857	2.069	6.768	28.263	1.335	1.557	1.756	7.053	
1.5	0.926	2.279	7.574	31.989	1.422	1.535	1.670	6.658	
2.0	0.971	2.419	8.112	34.474	1.476	1.520	1.615	6.418	
2.5	1.004	2.519	8.496	36.248	1.511	1.505	1.579	6.270	
3.0	0.5	0.611	1.278	3.611	13.154	0.905	1.481	2.004	8.275	
1.0	0.708	1.542	4.514	16.894	1.041	1.470	1.835	7.381	
1.5	0.767	1.700	5.056	19.137	1.112	1.449	1.746	6.948	
2.0	0.806	1.806	5.417	20.633	1.156	1.435	1.687	6.687	
2.5	0.833	1.881	5.675	21.701	1.187	1.425	1.647	6.513	

3.4 Shannon and Rényi entropy

Entropy was first introduced by Claude Shannon in 1948, as an uncertainty metric for a stochastic random variable. Entropies are widely used in the theory of information and communication, ecology, statistics, etc. The most prominent entropy metrics are Rényi entropy and Shannon entropy, both of which are widely available in the literature. Researchers also developed goodness of fit tests based on the entropies and divergence measures for the introduced distributions, for more detail see the following references [20,21].

Shannon entropy of discrete random variable Z isH(Z)=−∑z=0∞p(Z=z)logp(Z=z).

So, we may obtain the Shannon entropy for the PLo distribution as,H(Z)=−log(θ22(1+α))+∑z=0∞p(z)log(αθz2+2(2+θ+αθ)+z(2+3αθ)(1+θ)z+3).

The Rényi entropy is now determined byHγ(Z)=11−γlog∑z=0∞(p(Z=z))z,forγ>0andγ≠1.

Using equation (2) in the context of the PLo distribution, we obtain∑z=0∞(p(Z=z))z=(θ22(1+α))z∑z=0∞(αθz2+2(2+θ+αθ)+z(2+3αθ)(1+θ)z+3)z.

Thus, the Rényi entropy for a discrete PLo distribution is simplified to the following formula:Hγ(Z)=11−γlog{(θ22(1+α))z∑z=0∞(αθz2+2(2+θ+αθ)+z(2+3αθ)(1+θ)z+3)z}.

Some numerical values of Shannon entropy are given in Table 2. The Shannon entropy shows decreasing behavior with an increase in parameter θ while parameter α is fixed. On the other, there is an increasing trend for fixed θ.Table 2 H(Z) values for some choices of parameters.

Table 2	α=0.1	α=0.5	α=1.0	α=1.5	α=2.0	α=2.5	α=3.0	α=3.5	α=4.0	
θ=0.1	3.95440	4.06170	4.11819	4.14559	4.16206	4.17454	4.18582	4.19695	4.20826	
θ=0.5	2.38727	2.52382	2.59936	2.63757	2.65997	2.67442	2.68437	2.69157	2.69699	
θ=1.0	1.74611	1.89989	1.98725	2.03274	2.06015	2.07827	2.09105	2.10050	2.10775	
θ=1.5	1.40589	1.56374	1.65486	1.70311	1.73262	1.75240	1.76651	1.77705	1.78521	
θ=2.0	1.18821	1.34453	1.43577	1.48463	1.51482	1.53521	1.54986	1.56088	1.56945	
θ=2.5	1.03510	1.18753	1.27725	1.32570	1.35583	1.37630	1.39108	1.40224	1.41096	
θ=3.0	0.92077	1.06836	1.15583	1.20337	1.23308	1.25336	1.26805	1.27917	1.28788	
θ=3.5	0.83174	0.97422	1.05914	1.10552	1.13464	1.15457	1.16904	1.18003	1.18865	
θ=4.0	0.76020	0.89761	0.97990	1.02504	1.05347	1.07297	1.08718	1.09797	1.10645	

4 The PLo regression model

As, the previous section of this study reveals the PLo distribution can be utilized to analyze the over-dispersed data sets, which is an important aspect since most data in real life is over-dispersed. In this section, a new count regression based on PLo distribution is proposed to evaluate over-dispersed data sets.

Suppose that Z follows PLo distribution. To begin with, let's consider the following reparameterization:α=2+θ−θμ−θ2μ(1+θ)(θμ−3).

By using this approach, we can obtain the PMF of the PLo distribution in terms of mean E(Z)=μ>0 and θ>0. Then the corresponding PMF is obtained as(10) P(Z)=θ2(z2(2+θ−θμ−θ2μ(1+θ)(θμ−3))θ+2(2+θ+θ(2+θ−θμ−θ2μ(1+θ)(θμ−3)))+z(2+3θ(2+θ−θμ−θ2μ(1+θ)(θμ−3))))2(1+(2+θ−θμ−θ2μ(1+θ)(θμ−3)))(1+θ)z+3.

where z=0,1,2,… Hence, we may define the PMF in equation (2) as PLo(μ,θ) i.e. Z∼PLo(μ,θ). Let's configure that Z∼PLo(μ,θ), and the covariates are in relation with the ith mean by the log-link function as follows:(11) μi=E(Zi)=exp(xiT,β),i=1,2,3,…,n

where xiT=(1,xi1,xi2,…,xip) is the vector representing covariates and β=(β0,β1,…,βp)T is the regression coefficients unknown vector.

5 Parameter estimation and simulation

In this section, we utilized the maximum likelihood estimation (MLE) approach to estimate the PLo distribution parameters. The MLEs are consistent, efficient, and asymptotically unbiased. Due to these characteristics, MLE is commonly favored among researchers. Many authors have utilized this approach to find estimates of Poisson compound models se references [5,17,22].

Therefore, in this segment, we delve into the pursuit of MLE for the parameters of the PLo distribution. Let's denote a random sample of size n from the PLo distribution as Z1,Z2,…,Zn, with corresponding observed values z1,z2,…,zn. The MLEs for the PLo distribution parameters are derived through the direct maximization of the log-likelihood l(z) function given as follows.(12) l(z)=2nlog(θ)−nlog(2)−nlog(1+α)+∑i=1n(zi+3)log(1+θ)+∑i=1nlog(zi2αθ+2(2+θ+αθ)+zi(2+3αθ)).

Now, The ML estimates of α and θ can be obtained by partially differentiating equation (12) and the resulting normal equations are follows as:(13) ∂l(z)∂α=−n(1+α)+∑i=1nzi2θ+2θ+3ziθ(zi2αθ+2(2+θ+αθ)+zi(2+3αθ)).

and(14) ∂l(z)∂θ=2nθ+1(1+θ)∑i=1n(zi+3)+∑i=1n(zi2α+2+2α+3ziα)(zi2αθ+2(2+θ+αθ)+zi(2+3αθ)).

The above nonlinear equations can be solved numerically by using the Newton-Raphson method. For this purpose, we will utilize the optim function using R-software.

A detailed simulation study is performed to demonstrate the performance of the derived estimators. To evaluate the long-term accuracy of the maximum likelihood (ML) estimators for the parameters of the PLo distribution. We generated samples of sizes n = 25, 50, 100, 200, and 300 from the PLo distribution under different combinations of parameters. This process was repeated 10,000 times for each combination of parameter values and sample sizes. Subsequently, metrics such as the average bias (AB), mean relative error (MRE) and mean squared error (MSE) were computed for each parameter estimate across all iterations for the respective sample sizes. The results are presented in Table 3. It was observed that AB, MRE, and MSE decreased as the sample size increased across the various parameter combinations.Table 3 Simulation results of PLo distribution for some choices of parameters.

Table 3Para.	n	θˆ	αˆ	
AE	AB	MRE	MSE	AE	AB	MRE	MSE	
θ=0.1
α=0.5	25	0.10640	0.00640	0.06400	0.00050	1.34240	0.84240	1.68480	2.33680	
50	0.10360	0.00360	0.03560	0.00030	1.12920	0.62920	1.25840	1.73730	
100	0.10170	0.00170	0.01740	0.00020	0.92900	0.42900	0.85790	1.16770	
200	0.10040	0.00040	0.00390	0.00010	0.75080	0.25080	0.50160	0.63400	
300	0.09980	0.00020	0.00160	0.00010	0.65440	0.15440	0.30880	0.38730	
θ=0.5
α=0.5	25	0.54290	0.04290	0.08580	0.02140	1.32500	0.82500	1.65000	2.27190	
50	0.52460	0.02460	0.04920	0.01370	1.13960	0.63960	1.27920	1.74420	
100	0.51270	0.01270	0.02540	0.00910	0.91650	0.41650	0.83310	1.10870	
200	0.50130	0.00130	0.00260	0.00590	0.72500	0.22500	0.45000	0.60240	
300	0.49880	0.00120	0.00250	0.00460	0.63910	0.13910	0.27830	0.36860	
θ=1.0
α=0.5	25	1.13160	0.13160	0.13160	0.14030	1.41290	0.91290	1.82580	2.48550	
50	1.07430	0.07430	0.07430	0.07770	1.19150	0.69150	1.38300	1.89150	
100	1.04900	0.04900	0.04900	0.05320	1.02150	0.52150	1.04310	1.35030	
200	1.01340	0.01340	0.01340	0.03360	0.77320	0.27320	0.54640	0.70700	
300	1.00690	0.00690	0.00690	0.02700	0.69440	0.19440	0.38880	0.48940	
θ=0.1
α=1.0	25	0.10120	0.00120	0.01180	0.00040	1.61680	0.61680	0.61680	2.01650	
50	0.09910	0.00090	0.00920	0.00030	1.42980	0.42980	0.42980	1.61190	
100	0.09910	0.00090	0.00890	0.00020	1.35320	0.35320	0.35320	1.29430	
200	0.09930	0.00070	0.00690	0.00010	1.26420	0.26420	0.26420	0.91910	
300	0.09940	0.00060	0.00590	0.00010	1.19140	0.19140	0.19140	0.66360	
θ=0.5
α=1.0	25	0.50780	0.00780	0.01550	0.01530	1.59830	0.59830	0.59830	1.99880	
50	0.50050	0.00050	0.00110	0.01030	1.47060	0.47060	0.47060	1.66360	
100	0.49590	0.00410	0.00820	0.00660	1.35140	0.35140	0.35140	1.25180	
200	0.49560	0.00440	0.00880	0.00420	1.23590	0.23590	0.23590	0.83180	
300	0.49760	0.00240	0.00480	0.00300	1.17870	0.17870	0.17870	0.61710	
θ=1.0
α=1.0	25	1.04460	0.04460	0.04460	0.09030	1.65160	0.65160	0.65160	2.04800	
50	1.01100	0.01100	0.01100	0.05590	1.52450	0.52450	0.52450	1.75460	
100	0.99260	0.00740	0.00740	0.03760	1.38900	0.38900	0.38900	1.39030	
200	0.99100	0.00900	0.00900	0.02480	1.25880	0.25880	0.25880	0.96010	
300	0.99260	0.00740	0.00740	0.01810	1.19830	0.19830	0.19830	0.71660	
θ=0.1
α=1.5	25	0.09890	0.00110	0.01150	0.00030	1.82700	0.32700	0.21800	1.68270	
50	0.09820	0.00180	0.01800	0.00020	1.75270	0.25270	0.16840	1.47630	
100	0.09840	0.00160	0.01610	0.00010	1.72430	0.22430	0.14950	1.23500	
200	0.09900	0.00100	0.01050	0.00010	1.68390	0.18390	0.12260	0.94440	
300	0.09940	0.00060	0.00580	0.00010	1.67130	0.17130	0.11420	0.74960	
θ=0.5
α=1.5	25	0.49500	0.00500	0.01000	0.01320	1.82230	0.32230	0.21490	1.66780	
50	0.48850	0.01150	0.02300	0.00850	1.71730	0.21730	0.14490	1.44520	
100	0.49000	0.01000	0.02000	0.00570	1.70780	0.20780	0.13850	1.22230	
200	0.49510	0.00490	0.00990	0.00290	1.69520	0.19520	0.13010	0.90010	
300	0.49700	0.00300	0.00590	0.00200	1.66040	0.16040	0.10690	0.70670	
θ=1.0
α=1.5	25	0.99880	0.00120	0.00120	0.07190	1.82070	0.32070	0.21380	1.67300	
50	0.97540	0.02460	0.02460	0.04540	1.75650	0.25650	0.17100	1.51940	
100	0.97740	0.02260	0.02260	0.02870	1.70520	0.20520	0.13680	1.25970	
200	0.98720	0.01280	0.01280	0.01770	1.68980	0.18980	0.12650	0.98400	
300	0.98950	0.01050	0.01050	0.01160	1.66320	0.16320	0.10880	0.79570	
θ=1.5
α=1.5	25	1.50520	0.00520	0.00350	0.20150	1.83360	0.33360	0.22240	1.70640	
50	1.46930	0.03070	0.02040	0.12530	1.75660	0.25660	0.17110	1.56020	
100	1.46150	0.03850	0.02570	0.07860	1.73690	0.23690	0.15790	1.36030	
200	1.47270	0.02730	0.01820	0.04990	1.71860	0.21860	0.14570	1.09000	
300	1.48780	0.01220	0.00810	0.03520	1.71320	0.21320	0.14210	0.91180	

6 Application

In this section, we delve into the examination of three applications, each leveraging unique real-world datasets. We aim to evaluate the practicality and adaptability of the proposed model when confronted with real phenomena. Additionally, we seek to determine whether the suggested model stands as a superior choice compared to other competitive models such as New Three Parameter Poisson Lindley (NTPPL), Poisson Transmuted Moment Exponential (PTMEx) [23], Poisson transmuted record type exponential (PTRTE) [24], Poisson Moment Exponential (PMEx) [10], Poisson and Poisson Ramos Lozada (PRL) [13], discrete inverted Topp-Leone (DITL) [25] distributions in addressing these real-world scenarios. For regression modeling, we compared the PLo regression model with Poisson Quasi-Lindley (PQL) [26] and Poisson regression models.

To evaluate the goodness of fit of the PLo distribution, various metrics were employed, including negative log-likelihood (−l), AIC (Akaike Information Criterion), BIC (Bayesian Information Criterion), and the Chi-Square Statistic. The test statistics are.AIC=2×numberofparameters−2×log−likelihood

BIC=log(numberofobservations)×numberofparameters−2×log−likelihood

andχ2=∑(observed−expected)2expected.

6.1 Data-I (European corn borers)

The first dataset, comprising 120 observations, corresponds to a biological study involving European corn borers. The recorded observations can be found in Table 4. The experiment, conducted randomly across 8 hills in 15 replications, involved counting the number of borers per corn hill. The provided table details the estimates, along with comparative metrics for the PLo distribution and other competing distributions. The nonparametric plots for the first dataset are given in Fig. 3.Table 4 The MLEs, AIC, BIC, and Chi-square values of all the fitted models for the first dataset.

Table 4Count	Freq.	Expected	
PLo	NTPPL	PTMEx	PTRTE	PMEx	DITL	Poisson	PRL	
0	43	44.0707	44.6144	40.4161	42.1403	39.5580	52.1885	27.2265	18.0502	
1	35	30.3743	30.4606	33.6581	28.7761	33.6914	30.4244	40.3851	17.4679	
2	17	19.3091	19.0667	21.0523	18.0131	21.5212	14.1124	29.9516	15.7629	
3	11	11.6373	11.3366	11.8187	10.7101	12.2197	7.4663	14.8091	13.6192	
4	5	6.7372	6.5149	6.3104	6.1547	6.5047	4.3900	5.49158	11.4229	
5	4	3.7709	3.6544	3.2877	3.4522	3.3239	2.7906	1.62913	9.3766	
6	1	2.0457	2.0132	1.6924	1.9016	1.6514	1.8811	0.40275	7.5718	
7	2	1.0749	1.0935	0.8661	1.0329	0.8037	1.3271	0.08534	6.0364	
8	2	0.9798	1.2454	0.8979	7.8189	0.7259	5.4195	0.01888	20.6917	
Total	n=120									
MLE	θˆ	0.8325	1.0584	0.4648	1.0586	0.7416	1.9840	1.4833	2.5243	
αˆ	−0.1764	0.6111	0.8943	0.5703	–		–	–	
βˆ	–	0.8576	–	–	–		–	–	
GOF Measures	−l	200.414	200.415	200.821	200.415	201.221	205.152	219.188	244.501	
AIC	404.827	406.830	405.641	404.830	404.442	412.303	440.376	491.002	
BIC	410.402	415.193	411.216	410.405	407.230	415.091	443.163	493.790	
χ2	1.6513	1.4444	2.0822	6.0345	2.7268	6.9766	38.536	84.667	
Degree of Freedom	6	5	6	6	7	7	7	7	
P-value	0.6478	0.4856	0.5555	0.1965	0.6045	0.1371	<0.0001	0.0000	

Fig. 3 Boxplot, Violin, q-q, and histogram for the first dataset.

Fig. 3

Fig. 3 demonstrates the Boxplot, Violin, Q-Q, and histogram for the first dataset. The four plots elucidate that the European corn borers data set is right-skewed. The box plot shows that most values in the data set are concentrated at the lower end with some outliers. The violin plot illustrates a longer peak at lower values and a narrowing tail. The Q-Q plot confirms the deviations from normality, confirming the skewness. These findings are further supported by the histogram, which shows a frequency distribution with most data points at lower values and fewer at higher values. When combined, these plots show a distribution that is right-skewed, concentrated in lower values, and contains a few outliers.

The MLEs of all fitted models and their respected goodness-of-fit measures for the first dataset are reported in Table 4. The observed and expected frequency plots are given in Fig. 4.Fig. 4 The fitted probability mass functions of all the fitted models for the first dataset.

Fig. 4

Table 4 compares the observed counts of European corn borers with expected counts under the different probability distribution models mentioned above. Different goodness-of-fit measures are provided, such as −l, AIC, BIC, chi-square (χ2) and P-values. The AIC and BIC values are used to compare model fit, where lower values indicate a better fit. Among the models, PLo has the lowest AIC (404.827) and BIC (410.402), suggesting it is the most suitable model. Also, the PRL distribution, with the highest χ2 (84.667) and a P-value of 0.000, shows a poor fit. PLo distribution, with a minimum χ2 (1.6513) and highest P-value (0.6478), which best aligns with the observed data, indicating that the PLo distribution best fits among the competitor models. Fig. 4 demonstrates the graph of observed data (pink lines) with fitted distributions (black lines) for all the competitive models. It can be observed that the NTPPL, PTMEx, and PMEx models fit the observed data well, but the PLo distribution fits the observed data best among all. By combining both statistical and graphical analyses, it can be concluded that the PLo distribution is the best fit for the observed data which satisfies the null hypothesis.

6.2 Data-II (counts of seizures)

The dataset in Table 5 pertains to a patient grappling with intractable epilepsy and undergoing continuous treatment with a consistent anticonvulsant drug. The individual was observed over 351 days, revealing an average occurrence of 1.54 seizures per day. The presented table outlines the estimations and comparative measures for the PLo distribution in contrast to other competing distributions. The boxplot, Volin, q-q plot, and histogram for the second dataset are given in Fig. 5. All the plots reveal that the second data is right skewed with few outliers and most values in the data set are concentrated at the lower end. The MLEs of all fitted models and their respected goodness-of-fit measures for the second dataset are presented in Table 5. The observed and expected frequency plots are given in Fig. 6.Table 5 The MLEs, AIC, BIC, and Chi-square values of all the fitted models for the second dataset.

Table 5Count	Freq.	Expected	
PLo	NTPPL	PTMEx	PTRTE	PMEx	DITL	Poisson	PRL	
0	126	121.512	121.878	113.288	121.861	111.771	148.063	74.9324	52.2709	
1	80	89.7280	90.9476	97.0969	90.9486	97.3971	88.3486	115.711	51.2296	
2	59	59.3401	58.7444	62.7377	58.7491	63.6535	41.8858	89.3402	46.4659	
3	42	36.1452	35.2131	36.3287	35.2173	36.9783	22.5509	45.9864	40.2064	
4	24	20.7307	20.1633	19.8976	20.1663	20.1392	13.4520	17.7530	33.7052	
5	8	11.3660	11.1937	10.5495	11.1956	10.5296	8.6560	5.4828	27.6186	
6	5	6.0191	6.0769	5.4768	6.0780	5.3523	5.8972	1.4111	22.2457	
7	4	3.1014	3.2438	2.8015	3.2444	2.6652	4.1995	0.3113	17.6793	
8	3	3.0569	3.5388	2.8228	3.5396	2.5135	17.9468	0.0722	59.5783	
Total	n=351									
MLE	θˆ	1.4287	1.1157	0.2030	1.1157	0.7721	1.9044	1.5442	2.4780	
αˆ	0.9980	0.5731	0.8353	0.7228	–		–	–	
βˆ	–	1.6666	–	–	–		–	–	
GOF Measures	−l	593.955	594.482	595.727	594.482	595.863	620.056	636.046	716.218	
AIC	1191.91	1194.96	1195.45	1192.96	1193.73	1242.11	1274.09	1434.44	
BIC	1199.63	1206.55	1203.18	1200.69	1197.59	1245.97	1277.95	1438.30	
χ2	3.9707	4.6055	7.3848	4.6050	7.9525	48.7311	117.861	218.028	
Degree of Freedom	6	5	6	6	7	7	7	7	
P-value	0.5536	0.33022	0.1935	0.4659	0.2416	0.0000	0.0000	0.0000	

Fig. 5 Boxplot, Violin, q-q, and histogram for the second dataset.

Fig. 5

Fig. 6 The fitted probability mass functions of all the fitted models for the second dataset.

Fig. 6

The findings in Table 5 distinctly demonstrate that the PLo distribution outshines other competitive models by displaying the most favorable combination of the lowest statistical values such as AIC (1191.91), BIC (1199.63), and the highest P-value (0.5536). This pattern strongly indicates that the PLo distribution stands as the most suitable model among others. Fig. 6 shows the plots of observed and expected counts for the data set of Counts of seizures. It reveals that the NTPPL, PTRTE, and PMEx models correspond well with the actual data, but the PLo distribution corresponds best among all models. Combining statistical and graphical analysis yields the conclusion that the PLo distribution is the best match for the observed data, hence satisfying the null hypothesis.

6.3 Data-III (number of doctor visits)

Data III regarding the number of doctor visits originated from the Medicaid Consumer Survey which is funded by the Health Care Financing Administration took place in 1986. Gurmu [27] provides comprehensive and detailed information about the above-said data set. One can find the data set easily in the Ecdat package of R software. The goal of the study is to look at how the number of doctor visits affects the number of children in the household (x1i) and health status (x2i) of the family. The regression model given below is fitted using Poisson, Poisson Quasi Lindley (PQL), and PLo distributions.ln(μi)=β0+β1x1i+β2x2i.

Table 6 demonstrates the parameter estimates and their respective standard errors for the fitted models. It can be observed that AIC and BIC statistics for PLo distribution have the lowest values compared with PQL and Poisson distribution. So, it is inferred that the PLo distribution regression model fits better as compared to the other fitted models suggesting greater modeling accuracy. It can also be concluded that both the covariates i.e. the number of children and the health status have a significant effect on the number of doctor visits.Table 6 The findings of fitted count regression models of AZPRO data.

Table 6Para.	PLo	PQL	Poisson	
MLEs (SE)	p-Value	MLEs (SE)	p-Value	MLEs (SE)	p-Value	
β0	0.0800 (0.8909)	<0.0001	0.0413 (0.0994)	0.6800	0.7537 (0.0747)	<0.0001	
β1	−0.0529 (0.0715)	<0.0001	−0.1729 (0.0407)	<0.0001	−0.1781 (0.0318)	<0.0001	
β2	0.0903 (0.2158)	<0.0001	0.3052 (0.0322)	0.0001	0.2798 (0.0180)	<0.0001	
α	0.2602 (1.15222)	–	0.0001 (0.1612)	–	–	–	
−l	841.89	845.02	1097.7	
AIC	1689.8	1696.0	2201.4	
BIC	1702.3	1708.6	2213.9	

7 Conclusion

To model the count datasets, discrete distributions play a pivotal role. It has been seen that count datasets are generally over-dispersed. This study introduced a new two-parameter count distribution namely, the Poisson Loai distribution to deal with over-dispersed observation. The new model is derived by compounding discrete Poisson and two-parameter continuous Loi distributions. The new distribution is flexible due to its density function shapes. PLo distribution is showing exponentially decreasing and unimodal behavior based on parameter values. The hazard function (failure rate) of the PLo distribution increases with an increase in parameter values. Different statistical and reliability measures such as generating functions, moments, skewness, kurtosis, dispersion index, survival function, hazard rate, and entropies were mathematically, and numerically expressed. The mean and variance increase with an increase in parameter α, while the DI, coefficient of skewness, and kurtosis measures show a decreasing pattern. On the other side, there is an opposite pattern with an increase in parameter θ. The above-mentioned measures suggest that the novel model is well suited for over-dispersed, positively skewed, and increasing hazard rate datasets. The method of maximum likelihood estimation was used to estimate the PLo parameters. To evaluate the consistency of the parameters a comprehensive simulation study was also conducted. In addition, for count data sets a novel regression model based on the PLo distribution is introduced and compared with its competitive regression models based on a real dataset. Collectively, three real-world datasets were considered to illustrate the use of the novel model, the first regarding European corn borers, the second regarding counts of seizures, and the third regarding the doctor visits data. As a result, the new distribution outperforms all other studied distributions for fitting data sets, making it a valuable contribution to count data modeling.

8 Future work

Here are some potential topics for further study on the new two-parameter distribution:• The proposed distribution can be further investigated in a variety of dimensions. A particularly attractive path is to extend its neutrosophic extension for assessing datasets with uncertainty [[28], [29], [30], [31], [32]].

• The model parameters can be estimated using different estimation methods, for example, methods of moments, Anderson Darling, Cramer von Misses, Ordinary least squares, and weighted least squares.

• Bayesian estimation may be used to estimate model parameters using different loss functions and approximation strategies.

Data availability statement

Data included in article/supplementary material/referenced in article.

CRediT authorship contribution statement

Abdullah Ali H. Ahmadini: Supervision, Funding acquisition. Muhammad Ahsan-ul-Haq: Writing – review & editing, Writing – original draft, Methodology, Conceptualization. Muhammad Nasir Saddam Hussain: Writing – original draft.

Declaration of competing interest

The Authors have no conflict of interest.
==== Refs
References

1 Bhati D. Kumawat P. Gómez–Déniz E. A new count model generated from mixed Poisson transmuted exponential family with an application to health care data, Commun. Stat Methods 46 2017 11060 11076
2 Aljohani H.M. Akdoğan Y. Cordeiro G.M. Afify A.Z. The uniform Poisson–Ailamujia distribution: actuarial measures and applications in biological science Symmetry 13 2021 1258
3 Zeghdoudi H. Nedjar S. On Poisson pseudo Lindley distribution: properties and applications J. Probab. Stat. Sci 15 2017 19 28
4 Altun E. Bhati D. Khan N.M. A new approach to model the counts of earthquakes: INARPQX(1) process SN Appl. Sci. 3 2021 10.1007/s42452-020-04109-8
5 Hassan H. Dar S.A. Ahmad P.B. Poisson Ishita distribution: a new compounding probability model IOSR J. Eng. 9 2019 38 46
6 Ahsan-ul-Haq M. Al-Bossly A. El-Morshedy M. Eliwa M.S. Poisson XLindley distribution for count data: statistical and reliability properties with estimation techniques and inference Comput. Intell. Neurosci. 2022 2022 10.1155/2022/6503670
7 Maya R. Irshad M.R. Chesneau C. Nitin S.L. Shibu D.S. On discrete Poisson–Mirra distribution: regression, INAR (1) process and applications Axioms 11 2022 193
8 Chesneau C. The binomial-discrete Poisson-lindley model: modeling and applications to count regression Commun. Math. Res. 38 2022 28 51 10.4208/cmr.2021-0045
9 Bodhisuwan W. Aryuyuen S. The Poisson-transmuted janardan distribution for modelling count data Trends Sci 19 2022 10.48048/tis.2022.2898
10 Ahsan-ul-Haq M. On Poisson moment exponential distribution with applications Ann. Data Sci. 2022 10.1007/s40745-022-00400-0
11 Chesneau C. Bakouch H.S. Tomy L. Veena G. The Poisson-Lindley difference model with application to discrete stock price change Int. J. Model. Simul. 00 2022 1 10 10.1080/02286203.2022.2086422
12 Alomair A. Ahsan-ul-Haq M. A new extension of Poisson distribution for asymmetric count data: theory, classical and Bayesian estimation with application to lifetime data PeerJ Comput. Sci. 9 2023 e1748 10.7717/peerj-cs.1748
13 Alkhairy I. Classical and Bayesian inference for the discrete Poisson Ramos-Louzada distribution with application to COVID-19 data Math. Biosci. Eng. 20 2023 14061 14080 10.3934/mbe.2023628 37679125
14 Aswathy S. Qarmalah N. A Flexible Dispersed Count Model Based on Bernoulli Poisson–Lindley Convolution and its Regression Model 2023 1 20
15 Seghier F.Z. Ahsan-ul-Haq M. Zeghdoudi H. Hashmi S. A new generalization of Poisson distribution for over-dispersed, count data: mathematical properties, regression model and applications Lobachevskii J. Math. 44 2023 3850 3859 10.1134/S1995080223090378
16 Altun E. Cordeiro G.M. Ristić M.M. An one-parameter compounding discrete distribution J. Appl. Stat. 49 2022 1935 1956 10.1080/02664763.2021.1884846 35757587
17 Alghamdi F.M. Ahsan-ul-Haq M. Hussain M.N.S. Hussam E. Almetwally E.M. Aljohani H.M. Mustafa M.S. Alshawarbeh E. Yusuf M. Discrete Poisson Quasi-XLindley distribution with mathematical properties, regression model, and data analysis J. Radiat. Res. Appl. Sci. 17 2024 100874
18 Zaagan A.A. Mahnashi A.M. Analysis of leukemia and forest fires data using new Poisson Quasi-Shanker distribution Alexandria Eng. J. 104 2024 701 709 10.1016/j.aej.2024.08.012
19 Alzoubi L. Gharaibeh M.M. Khazaleh A. Benrabia M. Loai Distribution : Properties , Parameters Estimation and Application to Covid-19 Real Data 2022
20 Mahdizadeh M. Zamanzade E. Goodness of fit tests for Rayleigh distribution based on Phi-divergence Rev. Colomb. Estadística 40 2017 279 290
21 Zamanzade E. Mahdizadeh M. Entropy estimation from ranked set samples with application to test of fit Rev. Colomb. Estadística 40 2017 223 241
22 Maya R. Chesneau C. Krishna A. Irshad M.R. Poisson extended exponential distribution with associated INAR (1) process and applications Stats 5 2022 755 772
23 Alrumayh A. Khogeer H.A. A new two-parameter discrete distribution for overdispersed and asymmetric data: its properties, estimation, regression model, and applications Symmetry (Basel). 15 2023 1289 10.3390/sym15061289
24 Erbayram T. Akdoğan Y. A new discrete model generated from mixed Poisson transmuted record type exponential distribution Ric. Di Mat 2023 10.1007/s11587-022-00755-9
25 Eldeeb A.S. Ahsan-ul-Haq M. Babar A. A discrete analog of inverted topp-leone distribution: properties, estimation and applications Int. J. Anal. Appl. 19 2021 695 708
26 Grine R. Zeghdoudi H. On Poisson quasi-lindley distribution and its applications J. Mod. Appl. Stat. Methods 16 2017 403 417 10.22237/jmasm/1509495660
27 Gurmu S. Semi‐parametric estimation of hurdle regression models with an application to Medicaid utilization J. Appl. Econom. 12 1997 225 242
28 Albassam M. Ahsan-Ul-haq M. Aslam M. Weibull distribution under indeterminacy with applications AIMS Math 8 2023 10745 10757 10.3934/math.2023545
29 Aslam M. A new goodness of fit test in the presence of uncertain parameters Complex Intell. Syst. 7 2021 359 365 10.1007/s40747-020-00214-8
30 Ahsan-ul-Haq M. Neutrosophic kumaraswamy distribution with engineering application Neutrosophic Sets Syst 49 2022 269 276 https://digitalrepository.unm.edu/nss_journal/vol49/iss1/17
31 Ahsan-ul-Haq M. A new cramèr-von Mises goodness-of-fit test under uncertainty Neutrosophic Sets Syst. 49 2022 262 268 https://digitalrepository.unm.edu/nss_journal/vol49/iss1/16
32 Zeina M.B. Hatip A. Neutrosophic random variables Neutrosophic Sets Syst. 39 2021 4
