==== Front J Transl Med J Transl Med Journal of Translational Medicine 1479-5876 BioMed Central London 2628 10.1186/s12967-020-02628-x Research A continuous data driven translational model to evaluate effectiveness of population-level health interventions: case study, smoking ban in public places on hospital admissions for acute coronary events https://orcid.org/0000-0001-6169-3654Bonakdari Hossein hossein.bonakdari@fsaa.ulaval.ca 12 https://orcid.org/0000-0001-9930-6453Pelletier Jean-Pierre dr@jppelletier.ca 1 http://orcid.org/0000-0003-2618-383XMartel-Pelletier Johanne jm@martelpelletier.ca 1 1 grid.410559.c0000 0001 0743 2111Osteoarthritis Research Unit, University of Montreal Hospital Research Centre (CRCHUM), 900 Saint-Denis Street, R11.412, Montreal, QC H2X 0A9 Canada 2 grid.23856.3a0000 0004 1936 8390Department of Soil and Agri-Food Engineering, Laval University, 2425 rue de l’Agriculture, Québec, QC G1V 0A6 Canada 9 12 2020 9 12 2020 2020 18 46629 7 2020 20 11 2020 © The Author(s) 2020Open AccessThis article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/. The Creative Commons Public Domain Dedication waiver (http://creativecommons.org/publicdomain/zero/1.0/) applies to the data made available in this article, unless otherwise stated in a credit line to the data.Background An important task in developing accurate public health intervention evaluation methods based on historical interrupted time series (ITS) records is to determine the exact lag time between pre- and post-intervention. We propose a novel continuous transitional data-driven hybrid methodology using a non-linear approach based on a combination of stochastic and artificial intelligence methods that facilitate the evaluation of ITS data without knowledge of lag time. Understanding the influence of implemented intervention on outcome(s) is imperative for decision makers in order to manage health systems accurately and in a timely manner. Methods To validate a developed hybrid model, we used, as an example, a published dataset based on a real health problem on the effects of the Italian smoking ban in public spaces on hospital admissions for acute coronary events. We employed a continuous methodology based on data preprocessing to identify linear and nonlinear components in which autoregressive moving average and generalized structure group method of data handling were combined to model stochastic and nonlinear components of ITS. We analyzed the rate of admission for acute coronary events from January 2002 to November 2006 using this new data-driven hybrid methodology that allowed for long-term outcome prediction. Results Our results showed the Pearson correlation coefficient of the proposed combined transitional data-driven model exhibited an average of 17.74% enhancement from the single stochastic model and 2.05% from the nonlinear model. In addition, data demonstrated that the developed model improved the mean absolute percentage error and correlation coefficient values for which 2.77% and 0.89 were found compared to 4.02% and 0.76, respectively. Importantly, this model does not use any predefined lag time between pre- and post-intervention. Conclusions Most of the previous studies employed the linear regression and considered a lag time to interpret the impact of intervention on public health outcome. The proposed hybrid methodology improved ITS prediction from conventional methods and could be used as a reliable alternative in public health intervention evaluation. Keywords Transitional modelData processingComputer simulationHybrid modelInterrupted time seriesLag timeNonlinearPublic healthOsteoarthritis Research Unit, CRCHUMChair in Osteoarthritis, University of Montrealissue-copyright-statement© The Author(s) 2020 ==== Body Introduction Due to advances in technology and improvements in recording reliable data and sharing methods, the time series (TS) concept has emerged in many theoretical and practical studies over the past few decades [1]. This concept allows researchers to access the outcome of any phenomenon or intervention, at any time, with minimum cost and effort, and to plan possible solutions and control measures based on the forecasted data [2]. Therefore, improving knowledge about studying TS, preprocessing, modeling and, if needed, post-processing is imperative [3]. In the domain of public health interventions, the interrupted time series (ITS) concept has been widely employed to evaluate the impact of a new intervention at a known point in time in routinely observed data [4–10]. ITS is fundamentally a sequence of outcomes over uniformly time-spaced intervals that are affected by an intervention at specific points in time or by change points. The outcome of interest shows a variation from its previous pattern due to the effect of the intervention. The applied intervention splits TS data into pre- and post-intervention periods. Based on this definition, Wagner et al. [11] proposed segmented regression analysis for evaluating intervention impacts on the outcomes of interest in ITS studies. In this approach, the choice of each segment is based on the change point, with the possible additional time lag in some cases, in order for the intervention to have an effect [12–17]. In addition, for pre- and post-intervention period segments of a TS, the level and trend values should be determined either by linear [17] or nonlinear [6] approaches. Therefore, accurate values of the change point and time lag parameters are essential in segmented regression analysis. Affecting an intervention at a change point produces different possible outcome patterns in the post-intervention period for both level and trend parameters. Figure 1 illustrates some possible impacts of an intervention on the post-intervention period. As shown in Fig. 1a–c, a change in level (or intercept) may lead to a change in level after a time lag or a temporary level change after the intervention. Other possible patterns are a change in slope (or trend) with a change in slope after a time lag, or a temporary slope change as shown in Fig. 1d–f, respectively. In some cases, a change in both of these parameters could take place as an immediate change, e.g., a change after a time lag or temporary level and slope changes (Fig. 1g–i).Fig. 1 Possible patterns in interrupted time series post-intervention period data analysis Regardless of the popularity and consensus on using segmented regression-based methods for solving ITS problems, selecting the most appropriate time lag is a challenging task with an important impact on results in this type of modeling. The reason for the delicacy of this task is that there is no specific rule to define the time lag produced between the pre- and post-intervention periods. In some cases, the outcomes of interventions have an unknown delayed response to the implemented strategies and a lag time may occur long after an intervention. However, in ITS modeling, when segmented regression approaches are used, the exact time lag after an intervention should be taken into consideration to guarantee modeling result accuracy and appropriateness. In addition, an undocumented change point seriously complicates ITS analysis. Applying a continuous nonlinear TS method is considered reliable if the ITS analysis can be released from all these fundamental concerns. Therefore, there is a necessity to introduce potential uses of linear, nonlinear or a combination of both models for solving such problems. Over the past few years, soft computing methods have been employed across domains and have established reliable tools for modeling complex systems and predicting different phenomena in healthcare [18–24]. Among soft computing techniques, the Group Method of Data Handling (GMDH) is a common self-organizing heuristic model, which can be used for simulating complicated nonlinear problems. This evolutionary procedure is performed by dividing a complex problem into some smaller and simpler problems. Based on GMDH, this study proposes a novel methodology of the continuous modeling of an ITS based on data preprocessing. An example of the novel ITS modeling uses a linear-based stochastic model, a nonlinear-based model and an integration of a stochastic and a nonlinear model (hybrid). In order to run the models, certain tests and preprocessing methods are initially applied to the TS to prepare the data for stochastic modeling. It is crucial to investigate the structure of the TS being studied prior to modeling. Therefore, the TS undergoes stationarity testing along with normality testing. After surveying the characteristics of the TS, stationarizing methods appropriate to the TS are used. Then, in case of non-normal distribution, a normal transformation is applied to the stationarized TS. For the second TS modeling approach, the dataset is modeled with an artificial intelligence (AI) method which is, in this case, the Generalized Structure Group Method of Data Handling (GS-GMDH). In the third and final step, a hybrid model that combines the linear and nonlinear results is applied. Finally, the results are compared according to various indices and methods. Therefore, using this method facilitates modeling the ITS continuously, i.e. there is no need to identify the change point and intervention lag time. Dataset description Barone-Adesi et al. [25] carried out an extensive study on the effect of a smoking ban in public places on hospital admissions for acute coronary events (ACEs). In January 2005, Italy introduced legislation that prohibits smoking in indoor public spaces, the goal of which was the reduction of health issues caused by second-hand smoke [25]. Second-hand smoke consists of smoke exhaled by smokers and from lit cigarettes and causes numerous health problems in non-smokers every year, as well as high treatment costs for both patients and the government. The ban was undertaken on 10 January 2005 to confront the growing trend of ACEs and to control this problem. Bernal et al. [6] used a subset of ACEs data from subjects in Sicily, Italy, between 2002 and 2006 among those aged 0–69 years. They analysed the ITS data by applying segmented linear regression to the standardized rate of ACEs TS associated with the implementation of a ban on smoking in all indoor public places, to calculate the change in the subsequent outcome levels and trends. Based on Barone-Adesi et al.’s [25] assumption, Bernal et al. [6] considered only a level change in ACEs occurring and there was no lag between the pre- and post-segments in the modeling procedure. Here, we used the dataset from Bernal et al. (Fig. 4, [6]) as an example to illustrate the proposed method’s performance in ITS simulation from real data regarding a health problem; it is not meant to contribute to the substantive evidence on the topic. The dataset employed comprises routine hospital admissions with 600–1100 ACEs. More information about the dataset can be found in the Barone-Adesi et al. article [25]. Methods Preprocessing Time series are data recorded continuously and based on time to institute a sequence of measures, each of which refers to a time. Thus, the ACE data collected monthly from 2002 to 2006 is a TS. Each TS consists of four terms: jump + trend + period + stochastic component. The first three terms, known as deterministic terms, are calculable and removable. The jump term represents the sudden changes that occur in TS. These changes are detectable as steps in TS plots or by numerical tests. The trend term represents the gradual upward or downward changes that take place during a long period of time; this term is denoted in TS as a linear fitted line. The third deterministic term, the period, represents the periodic alternations in TS, which are seen as sinusoidal variations. Therefore, only the remaining stochastic term is required for use in stochastic or nonlinear modeling. This term is achieved while stationarity (absence of a deterministic term) occurs. Numerous tests and methods exist for investigating and omitting deterministic terms and some are presented below. In stochastic modeling, two conditions must be met: the first is the stationarity (for details see below, Stationarizing methods); second, the distribution of the TS should be normal. Thus, in order to start stochastic-based modeling, the existence of deterministic terms must be checked, and when present, they should be removed. The Mann–Whitney (MW), Fisher, and Mann–Kendall (MK) tests are employed to check the jump, period and trend, respectively; the Kwiatkowski–Phillips–Schmidt–Shin (KPSS) test to assess the overall stationarity of the TS; and the Jarque–Bera (JB) test to check the normality of the TS. Trend A non-parametric test is used to assess the trend term in the studied TS. The MK test was developed to detect the gradual changes in TS, both seasonal and non-seasonal. The test equation is as follows [26]: 1 UMK=MK-1varMK-0.5MK>00MK=0MK+1varMK-0.5MK<0 where UMK is the standard Mann–Kendall statistic, MK is the Mann–Kendall statistic, and var(MK) is the variance of MK. MK and var(MK) are defined as: 2 MK=∑i=1N-1∑j=i+1Nsgnxj-xi 3 varMK=118NN-12N+5-∑gptgtg-12tg+5 where p is the number of identical groups, tg is the observation number in the gth group, sgn is the sign function, and N is the number of samples. The (MK) test equation for a seasonal trend is expressed as follows: 4 Sk=∑i=1Nk1∑j=i+1Nk-1sgnxki-xkj 5 SMK=∑k=1ωSk-sgnSk 6 varSMK=∑kωNkNk-12Nk+518+2∑i=1ω-1∑j=i+1ωσij 7 USMK=MKvarMK-0.5 where ω is the number of seasons in a year and σij is the covariance of the statistic test in seasons i and j. The trend in the TS is insignificant if Uα/2