
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

39232080
71762
10.1038/s41598-024-71762-z
Article
The value of error-correcting responses for cognitive assessment in games
http://orcid.org/0009-0003-2634-8056
Markovitch Benny b.m.markovitch@tue.nl

1
Evans Nathan J. 23
Birk Max V. 1
1 https://ror.org/02c2kyt77 grid.6852.9 0000 0004 0398 8763 Human Technology Interaction, Eindhoven University of Technology, 5612 Eindhoven, AZ The Netherlands
2 https://ror.org/05591te55 grid.5252.0 0000 0004 1936 973X Department of Psychology, Ludwig Maximilian University of Munich, 80799 Munich, Germany
3 https://ror.org/00rqy9422 grid.1003.2 0000 0000 9320 7537 School of Psychology, University of Queensland, St Lucia, 4067 Australia
4 9 2024
4 9 2024
2024
14 2065715 3 2024
30 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by/4.0/ Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
Traditional conflict-based cognitive assessment tools are highly behaviorally restrictive, which prevents them from capturing the dynamic nature of human cognition, such as the tendency to make error-correcting responses. The cognitive game Tunnel Runner measures interference control, response inhibition, and response-rule switching in a less restrictive manner than traditional cognitive assessment tools by giving players movement control after an initial response and encouraging error-correcting responses. Nevertheless, error-correcting responses remain unused due to a limited understanding of what they measure and how to use them. To facilitate the use of error-correcting responses to measure and understand human cognition, we developed theoretically-grounded measures of error-correcting responses in Tunnel Runner and assessed whether they reflected the same cognitive functions measured via initial responses. Furthermore, we evaluated the measurement potential of error-correcting responses. We found that initial and error-correcting responses similarly reflected players’ response inhibition and interference control, but not their response-rule switching. Furthermore, combining the two response types increased the reliability of interference control and response inhibition measurements. Lastly, error-correcting responses showed the potential to measure response inhibition on their own. Our results pave the way toward understanding and using post-decision change of mind data for cognitive measurement and other research and application contexts.

Keywords

Game-based cognitive assessment
Interference control
Response inhibition
Response-rule switching
Change of mind
Error-correction
Subject terms

Human behaviour
Diagnostic markers
The Dutch Research Council (NWO). VI.Veni.202.171issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

The behaviorally restrictive design of typical cognitive assessment tools has been criticized for failing to capture the dynamic nature of human cognition and decision-making1–4, including common behaviors such as error-corrections5–7. This is because most cognitive assessment tools are too behaviorally restrictive to allow error-correcting responses to follow mistaken initial responses in the same trial4–7. This limitation also applies to most cognitive games, which use game-like environments to assess cognitive functions, yet are typically as behaviorally restrictive as other cognitive assessment tools4,8–14. Critiques of the overly restrictive designs of cognitive assessment tools motivated the development of a less restrictive cognitive game called Tunnel Runner4. Tunnel Runner gives players continuous control over their interaction with the cognitive assessment system and provides players with natural opportunities and incentives to correct mistaken initial responses4. However, since error-correcting responses are rarely considered in cognitive assessment, the psychometric and theoretical foundations needed to use error-correcting responses are lacking, meaning that error-correcting responses remain unused even when allowed, encouraged, and measured4. Crucially, this limits the potential benefits of cognitive games, such as Tunnel Runner, which tend to evoke relatively high rates of mistaken initial responses4,13 that are ignored by reaction-time-based measurements, leading to increased data loss and reduced measurement efficiency. However, since error-correcting responses can reflect the same cognitive functions measured with initial response data5,15 and could create new cognitive measurement opportunities5–7, they might be a valuable source of cognitive data.

Historically, cognitive tasks were designed to evoke repeatable experimental effects, rather than reliably assess individual differences16. This is reflected in the reliance on conflict effects, which compare performance on regular trials against conflict trials that place additional demands on specific cognitive functions, such as response inhibition in stop-signal tasks17,18, and interference control in flanker tasks19. Conflict tasks’ psychometric reliability, defined by the ratio between individual differences and measurement error20 is limited by the tasks’ reliance on reaction time (RT) differences between trial types20,21, and their tendency to be demotivating and disengaging4 which can lead people to respond inconsistently and fail to perform to their full ability. These and other factors20 lead many conflict-based tasks to exhibit unacceptably low reliability16. However, due to the considerable interest in using conflict-based measurements in scientific research20,22–27 and health care28,29, several approaches have been proposed to improve their reliability. These proposals include using theoretically-informed response models30, evoking stronger conflict effects31, and providing better experiences via cognitive games4,8,14. These different approaches can be complemented by the use of error-correction data, which may increase the amount of response data, and could provide new opportunities to measure cognition.

Using error-correcting responses for cognitive assessment requires a theoretically-grounded understanding of the cognitive functions involved and how they may relate to the cognitive functions measured by correct initial responses. Models of cognition involving processes such as evidence accumulation in favor of specific decisions7,32, or competition between response inhibition and initiation processes17,18, have traditionally been developed to understand initial response data. However, evidence accumulation models have shown the ability to account for error-correcting responses by assuming that error-correcting responses reflect a continuation of the evidence accumulation process after initial responses5,33. These findings converge with the observation that responses in trials following erroneous responses are faster when they match what would have been error-correcting responses34,35 to support the hypothesis of cognitive continuity. According to the hypothesis of cognitive continuity (or spillover5), error-correcting responses reflect instances where the cognitive processes responsible for initial responses continue their operation after erroneous initial responses were made to later evoke error-correcting responses in the same trial. By outlining how error-correcting responses relate to the cognitive functions measured with correct first responses, cognitive continuity provides a path toward understanding and using error-correcting responses.

Cognitive continuity implies that error-correcting responses can be used for cognitive assessment alongside correct initial responses because error-correcting responses reflect a continuation of the processes measured by the initial responses5. However, the extent of cognitive continuity is contested33, and continuity between correct initial and error-correcting responses has not been directly examined in conflict-based measurements, whose use of difference scores may limit cognitive continuity as continuity between response types per each trial type does not guarantee continuity in the differences between conflict and non-conflict trials. However, if cognitive continuity holds with conflict-based measurements, then error-correcting responses could be used to supplement or even replace first-response-based measurements, as both response types would reflect the same cognitive functions. This could allow error-correcting responses to become an important part of cognitive assessment.

Although cognitive continuity may resolve theoretical challenges to the use of error-correcting responses in conflict-based measurements, psychometric challenges remain. Specifically, statistical approaches for the use of error-correcting responses in conflict-based cognitive assessment have, to our knowledge, been neither developed nor evaluated. This is a problem, because the estimation of individual differences in cognitive functions using initial responses alongside error-correcting responses can be complex. Specifically, the statistical estimation method should reflect both continuity and differences between the two response types across different trial types33. However, since this type of assessment has not been performed with conflict-based measures, it is unclear how error-correcting responses can be used for conflict-based cognitive assessment.

In this article, we address key theoretical and psychometric challenges to the use of error-correcting responses in conflict-based cognitive assessment using behavioral data collected via Tunnel Runner. Tunnel Runner includes conflict-based measures of interference control, response inhibition, and response-rule switching4, which are used to limit the influence of irrelevant stimuli on behavior22, suppress dominant but inappropriate behavioral tendencies22, and flexibly adapt to changing environments22,36, respectively. Crucially, Tunnel Runner’s continuous player control allows and encourages players to make error-correcting responses in trials where they made incorrect initial responses4; as unlike typical cognitive measurement tools5,6, Tunnel Runner’s trials do not terminate immediately after an initial response. This feature uniquely enables Tunnel Runner to naturally encourage and measure players’ error-correcting responses, and makes it a unique research paradigm to assess cognitive continuity with different conflict effects and to develop and evaluate statistical approaches to use error-correcting responses for conflict-based cognitive measurement. For these reasons, we used behavioral data collected through Tunnel Runner to assess the potential value of error-correcting responses for cognitive assessment in conflict-based tasks. Specifically, we sought to answer the following interconnected research questions: RQ1: Do error-correcting responses show cognitive continuity with initial responses in conflict-based cognitive measurements?

RQ2: Can error-correcting responses be used alongside correct initial responses to increase the reliability of conflict-based cognitive measurements?

RQ3: Can error-correcting responses be used as cognitive measurements on their own?

Material and methods

Tunnel runner

Tunnel Runner4 is an infinite runner game in which a group of five player-controlled rats run through a tunnel filled with obstacles, with different aspects of the game designed to create response conflicts to assess players’ interference control, response inhibition, or response-rule switching. We focus on describing the relevant gameplay sections and cognitive measurements, as a full description and validation of the game has been published elsewhere4.

The goal of the player is to guide the central rat to the section of the circular obstacle upcoming in the tunnel that matches the central rat’s color, requiring the player to rotate the rats’ positioning (Fig. 1). As the rats move through the tunnel, players can hold the A key to rotate the rats to the left and the L key to rotate them to the right, which also allows players to change directions and correct mistaken initial responses. The rotation is continuous rather than ballistic; only lasting while players hold down a direction key. Good performance is motivated by rewarding players with points equivalent to the number of times in a row (the streak) that the central rat passed through the correct section of the obstacle. After passing through an obstacle, all the rats are colored gray and their rotation is disabled. New colors are then assigned to the central and flanking rats, 433 and 350 milliseconds, respectively, after the new obstacle is presented. Rotation is re-enabled once the central rat’s color was assigned.

Tunnel Runner starts with 40 training trials, which are not used for assessment. These are followed by 126 regular trials, 126 mismatching flanker trials, 84 lava trials, and 120 ice trials. There are 2 breaks after the training trials, dividing the test trials into 3 blocks.

Regular trials

In 126 regular (congruent flanker) trials, shown at the top left of Fig. 1, the central rat and the four flanking rats are assigned the same color after the obstacle is presented. The rats can be rotated after the central rat’s color assignment, leaving players with 1317 ms to pass through the correct section of the obstacle before the next trial begins. The time until obstacle collision is adapted based on the outcomes of regular and mismatching flanker trials. Passing through the correct section reduces the time between the next color assignment and obstacle collision, whereas passing through the incorrect section increases the time between color assignment and obstacle collision. This adaptation targets success rates of 80%.Fig. 1 Tunnel Runner’s trial types. In regular trials (top left), all rats have matching color. In mismatching flanker trials (top right), the flanker rats are differently colored from the central rat. In ice trials (bottom left), the relation between player input and the rats’ movement is reversed. In lava trials (bottom right), a delayed stop-signal appears in the form of lava, penalizing players for moving the rats.

Mismatching flanker trials

In 126 regular mismatching flanker trials, depicted on the upper right of Fig. 1, the flanker rats are assigned the color of the section opposite to the correct section. Since flanker rats match the correct section in 50% of trials, players are asked to ignore them. These trials enable the measurement of players’ capacity for interference control4 via the differences between RTs on matching and mismatching flanker trials.

Lava trials

During 84 lava trials, which measure players’ response inhibition4 and are depicted at the bottom right of Fig.  1, the rats are surrounded by lava (a stop-signal) after color assignment. Touching the lava, which can only be prevented by not moving the rats, leads players to continuously lose points until they rotate the rats back to the starting point. Lava trials are independent of the flanker condition, such that the colors of the central and flanker rats are equally likely to match or mismatch. Creating 42 matching and 42 mismatching flanker lava trials. At first, the lava appears 300 milliseconds after the central rat is assigned a color. The stop-signal delay is then adapted based on the player’s performance, and the delay in matching flanker trials is adapted independently of the delay in mismatching flanker trials. Moving the rats after the stop-signal results in lava appearing 50 milliseconds earlier, giving players more time to stop early. Successful inhibition causes lava to appear 50 milliseconds later, giving players less time to stop early. This adaptation leads players to inhibit around 50% of responses in lava trials4, which is needed to calculate the stop-signal reaction time measure of response inhibition by the integration method17,18.

Ice trials

During 120 ice trials, depicted on the bottom left of Figure 1, the tunnel is filled with ice as soon as the rats pass the previous obstacle, reversing the key mapping so that holding the A key rotates the rats to the right, while the L key rotates them to the left. Matching and mismatching flanker trials are equally spread across ice trials, while fire trials rarely overlap with ice trials as this combination is not used for measurement. Ice trials enable the measurement of players’ capacity for response-rule switching via the differences between RT in ice and non-ice trials4.

Operationalizing error-correcting responses in tunnel runner

We aimed to define and measure error-correcting responses according to leading models of speeded binary decision-making and response inhibition. Importantly, as Tunnel Runner enables different types of error-correcting behaviors, such as response-switching in non-lava trials and late stopping in lava trials, and these different behaviors are typically modeled using different frameworks, we considered distinct theoretical perspectives for each type of error-correcting behavior. Specifically, we considered error-correcting responses that involve response-switching to result from an evidence accumulation process5,15,32,37, which is the dominant perspective on speeded binary decisions7. In contrast, we considered error-correcting responses involving late stopping in lava trials to result from a competition between two independent cognitive processes17, which is the dominant perspective on response inhibition in the stop-signal literature18.

Error-correction in non-lava trials

To operationalize error-correcting responses in non-lava trials, we assumed that the timing and nature of players’ initial and error-correcting responses depend on the dynamics of an evidence accumulation process. The evidence accumulation perspective7,32 involves the accumulation of evidence until it reaches a response boundary for one of the competing decisions. Once enough evidence has accumulated to reach a response boundary, the corresponding response is initiated.

The hypothesis of cognitive continuity implies that the evidence accumulation process does not end once a response boundary is reached; rather, the evidence accumulation process continues and could lead to a reversal of the initial response5,33. Thus, the time from target presentation to a correct first response, or to an error-correcting response, would each reflect the time required for the evidence accumulation process to reach the correct response boundary. Consequently, in non-lava trials, we measured error-correcting responses as the time between the central rat’s color assignment and the initiation of the reversal of an incorrect first response. We name this measure RT2, which is intended to provide an additional measure of the time required for the evidence accumulation process to reach the correct response boundary. A simplified evidence accumulation process is illustrated in Fig.  2.Fig. 2 A simplified evidence accumulation process representation of two response types. For the correct first response, evidence accumulates to evoke a correct first response, enabling standard RT measurement. For the error-correcting response, evidence first accumulates to evoke an initial incorrect response and then continues to accumulate to later evokes an error-correcting response. This enables the measurement of RT2, the time from target presentation until the error-correcting response.

Error-correction in lava trials

To operationalize error-correction in lava trials, we assumed that go and stop responses are determined by a competition between two independent cognitive processes, as outlined by the independent race model17,18. From this perspective, the winner of a competition between a ‘go runner’ and a ‘stop runner’ determines whether a response is initiated or inhibited. Thus, stop-signal reaction time (SSRT), which is calculated in typical stop-signal tasks, is an estimate of the time it takes the stop runner to reach the competition’s ‘finish line’ and inhibit a response17,18.

The hypothesis of cognitive continuity implies that the stop runner does not ‘quit’ once a go response is initiated and can reach the finish line in time to inhibit the ongoing response, as shown in Fig. 3. In Tunnel Runner’s lava trials, the inhibition of an ongoing response occurs when a player stops pressing the initial response button. Thus, the time from the presentation of a stop-signal until the inhibition of the ongoing go response, which we call the time-to-stop (TTS), reflects the time it takes for the stop runner to reach the finish line in trials where the go runner made it first. This means that both TTS and SSRT measurements should reflect the time it takes for the stop runner to reach the finish line. However, since players may inhibit their responses for reasons other than the stop-signal, TTS measurements require response inhibition to be followed by a corrective response (such as pressing L to move back and away from the lava), showing clear recognition of the mistaken initial response. On average, players took 191 ms between the inhibition of the ongoing response and the initiation of the corrective response. Thus, TTS is measured as the time between the stop signal and the stopping of the ongoing initial response, which was then followed by a corrective response.Fig. 3 A competition between response initiation (go) and inhibition (stop) processes. In the top part, the go process reaches the finish line before the stop process, leading to an initial response that is later stopped. In the bottom part, the stop process reaches the finish line before the stop process. SSD, stop-signal delay, is the time between the target presentation and the stop-signal presentation. TTS, time-to-stop, is the time between the onset of the stop-signal and the late stop. SSRT is the stop-signal reaction time.

Error-correcting responses differ from correct first responses

Cognitive continuity does not imply that error-correcting responses are identical to correct initial responses, since error-correcting responses can only be observed when the evidence accumulation process initially favored an incorrect first response or the go-runner initially won the competition. In other words, error-correcting responses selectively reflect trials where the cognitive processes under investigation performed worse than when a correct first response was initiated. Thus, error-correcting responses should take longer to make than correct first responses in a manner that may differ between individuals. Furthermore, several studies suggest that while initial and error-correcting responses share much in common, there are discontinuities between the two33,37,38. As such, error-correcting responses are unlikely to be directly comparable to correct first responses, and their use requires appropriate statistical adjustments.

Procedure

We used the data that we originally collected for validating Tunnel Runner4, which consisted of two online studies of Tunnel Runner. We kept the original two-studies structure to limit our researcher degree-of-freedom, and because it enabled us to independently replicate the results of our analyses. In the studies, before the informed consent form, participants completed a brief test to check whether Tunnel Runner could be displayed with 50 frames-per-second or more, ensuring good player experience and precise measurement4. Following consent, participants completed questionnaires and played Tunnel Runner.

Sample description

We conducted two studies through CloudResearch39, recruiting CloudResearch-approved Mechanical Turk users from the USA with at least 95% approval rate on at least 1,000 human intelligence tasks. These criteria should ensure a high-quality participant pool39. Participants’ median age was 38 (interquartile range: 17) in study 1, and 38 (interquartile range: 13) in study 2. Of the 117 participants in study 1, 73 identified as men, and 99 reported gaming on a weekly basis. Of the 121 participants in study 2, 81 identified as men, and 101 reported gaming on a daily basis.

As recommended to ensure high data quality in Tunnel Runner4, we excluded data at both the trial and individual levels. At the player level, we used a scoring system4 in which specific response patterns incur one or two points, with two or more points leading to the exclusion of player data. The following criteria led to a player’s exclusion: average frames-per-second lower than 35 across the study (study 1: 2, study 2: 1); first response accuracy no higher than 3 standard errors from 0.5 (11, 18); and non-response on more than 10% of non-lava trials (7, 2). Whereas at least two of the following were sufficient for exclusion: stopping rate above 0.7 or below 0.3 in lava trials (11, 8); correcting mistaken first movements in less than 30% of opportunities in non-lava trials (18, 13); correcting first movements in less than 30% of opportunities in lava trials (11, 9); failing to respond in more than 3% of non-lava trials (18, 17); and average frames-per-second below 45 (2, 2). Failure to meet these criteria led to the loss of 22 players in study 1 and 31 in study 2.

We applied additional player-level filters separately to the analyses of the ice effect, flanker effect, and SSRT and TTS measures. Since we focused on the use of error-correction data, we only considered players who had at least 3 error-correcting responses per relevant trial type. This led to a loss of 15 and 11 players for studies 1 and 2’s flanker effects, a loss of 2 and 0 players for studies 1 and 2’s ice effects, and a loss of 1 and 2 players for studies 1 and 2’s SSRT. Furthermore, we only calculated SSRT for players whose rate of stopping on lava trials was neither higher than 70% nor lower than 30%, as required by the integration method8,18, losing 2 and 0 players per studies 1 and 2’s SSRT and TTS measures.

Our analyses of ice and flanker effects on correct first response RT (RT1) excluded trials4 with RT1 lower than 300ms or higher than 1,500ms after the central rat’s color assignment or whose responses were more than 3 standard deviations from a players’ average RT1 per condition. We also excluded error-correcting responses that came more than 1 second after an initial response or a stop-signal or were further than 3 standard deviations from a player’s average error-correction time per condition. We did not apply trial-level filtering to the calculation of SSRTs18. The average number of trials analyzed per participant, measurement type, and study are described in Table 1.

Data analytic approach

All statistical tests were two-sided with a α of .05 and were accompanied by 95% confidence intervals. We fitted hierarchical regression models using R package lme440 and tested the models’ fixed effects with cluster-robust standard errors of type 241 from the ClubSandwich package42 to mitigate heteroscedasticity. We used hierarchical regressions to account for the clustering of responses at the level of an individual participant and estimated individual differences in responses to cognitive challenges via the corresponding random slopes or intercepts obtained from the regression models. Hierarchical models, particularly joint models, are highly effective in estimating individual differences and their associations in the presence of measurement error30,43–45. When testing the models’ fixed terms, we assessed the normality of the models’ residuals and random effects with QQ-plots. Hierarchical regression models are robust against non-normality46, such that only very severe non-normality would have required us to change the analyses.

Hierarchical regression models allow cognitive measurements to reflect theoretical assumptions30, making these models suitable for embodying the hypothesis of cognitive continuity when jointly modeling the time from target presentation to correct first responses (RT1), and from target presentation to an error-correcting response (RT2). We modeled and measured individual differences in ice and flanker effects with fixed effects that accounted for trial type, response type, and for their interaction. Crucially, the models’ random effects accounted for individual-level variability per trial type and response type but not for their interaction. This random effect structure embodied the assumption that individual differences in the effect of trial type (the conflict effect) are shared between response types.

Since we calculated players’ SSRTs separately per matching and mismatching flanker trials using the integration method17,18, hierarchical models could not model players’ SSRT alongside their time-to-stop (TTS) measures. Where TTS reflects the time between a stop-signal and the stopping of an ongoing incorrect response. Previous work on Tunnel Runner4 showed that its SSRT can be validly and reliably estimated as an average of two z-transformed SSRTs calculated separately in matching and mismatching flanker trials. Thus, when using players’ TTS to measure SSRT, we first z-transformed TTS and then averaged it with the two z-transformed SSRT sub-measures, thereby using TTS as a third SSRT sub-measure.

We estimated measurement reliability via McDonald’s ω47 for SSRTs, and with the even-odd split-half method48 for the other measures. McDonald’s ω is a natural fit for SSRT, since49 it is calculated as the average of several z-transformed measures: two SSRT sub-measures calculated separately per matching and mismatching flanker trials, and also TTS when applicable. Split-half reliability is a natural fit for measures based on hierarchical regression, as it is often used with cognitive measurements and would penalize the model-based estimates for overfitting.

To estimate the uncertainty around reliability estimates, we used the recommended bias-corrected and accelerated (BCa) bootstrapping49,50 with 100,000 iterations. However, this procedure has not been established for calculating differences between the reliability of nested measurements, such as SSRT calculated with or without TTS, or flanker effect calculated with or without RT2. We used simulations to assess the validity of bootstrapping in this context of nested measurements, and we found that bootstrapping incorrectly estimates the correlation between the nested measurements’ reliability , and therefore incorrectly estimates the uncertainty around their differences. For this reason, we will not report confidence intervals around reliability differences and we will not generalize these differences beyond our specific samples and cognitive measurements.

To estimate and compare the potential of measurements based on error-correcting responses and initial responses, we needed to separate the number of trials used per measurement from the measurement’s ability to assess individual differences. For measurements based on hierarchical modeling, this was achieved with the precision statistic η31, which is the ratio between the standard deviations of individual differences and the residual standard deviation estimated by a statistical model. Since we calculated the two first-response-based SSRT sub-measures via the integration method, their precision could not be estimated in a comparable way.

Statistical expectations

We aimed to assess cognitive continuity between conflict-based measures based on correct first responses and error-correcting responses, and to evaluate the measurement potential of error-correcting responses alongside their ability to supplement measures based on first responses. The assessment of cognitive continuity was driven by statistical expectations, and shaped our expectations regarding the ability of error-correcting responses to supplement initial response data. This, in turn, reflected back on our conclusions regarding cognitive continuity.

To establish cognitive continuity between conflict-based measures of different response types, we expected error-correcting responses to replicate the conflict effects seen in Tunnel Runner’s correct first responses4. These include the flanker effects of increased time until correct initial responses (RT1) and stop-signal reaction time (SSRT), and the ice effect of increased RT1. Furthermore, if conflict-based measures of error-correcting responses are cognitively continuous with conflict-based measures of correct first responses, then the two measurement types should strongly correlate with each other. This means that flanker effects on RT1 and time until error-correcting responses (RT2) should correlate, ice effects on RT1 and RT2 should correlate, and players’ SSRT and time-to-stop (TTS) should correlate. Thus, we expected players’ error-correcting responses to show ice and flanker effects on RT2 whose magnitudes were comparable to those observed on RT1. Furthermore, we expected players’ TTS to show a flanker effect comparable to the one observed on SSRT. If these expectations were met and error-correction measurements strongly correlated with first-response-based measurements, we concluded that cognitive continuity likely occurred between the two response types. Furthermore, whenever cognitive continuity likely held, we expected error-correcting responses to beneficially supplement first-response data. This would provide converging evidence regarding continuity between correct initial and error-correcting responses, as supplementing initial response data with dissimilar data should reduce reliability.

Results

Response tendencies

The prevalence of players’ error-correcting responses, described in Table 1, shows that error-correcting responses were common in Tunnel Runner, and were particularly common in lava trials. As most lava trials with an initial response were followed by a delayed stop followed by a correction, which allowed the measurement of TTS. Players’ mean RTs per trial and response types can be found in the Supplementary material.Table 1 Average number of trials per participant per measurement type across the studies. SSRT refers to stop-signal reaction time.

Measurement type	Study 1: trials per participant	Study 2: trials per participant	
Ice effect: correct first responses	242.1	242.6	
Ice effect: error-corrections	62.2	69.4	
Flanker effect: correct first responses	193.9	191.5	
Flanker effect: error-corrections	28.5	32.6	
SSRT: lava trials	84.0	84.0	
Time-to-stop in lava trials	28.8	35.8	

Do error-correcting responses show cognitive continuity with initial responses in conflict-based cognitive measurements?

Table 2 Conflict effects on error-correcting responses and their comparisons with conflict effects on first responses.

Source	Study 1: mean (95% CI)	Study 2: mean (95% CI)	
Flanker effect on RT2	44.2* (26.1–62.3)	58.7* (41.9–75.6)	
Difference between flanker effects on RT2 and RT1	−11.7 (−30.2–6.7)	−4.9 (−21.6–11.9)	
Ice effect on RT2	49.7* (35.2 − 64.2)	48.6* (34.5 − 62.8)	
Difference between ice effects on RT2 and RT1	−35.0* (−57.8–−12.2)	−39.9* (−62.7–−17.1)	
Flanker effect on TTS	15.2* (8.8–21.7)	15.5* (8.2–22.7)	
Difference between

flanker effects on TTS and SSRT

	−1.9 (−12.4–8.6)	−4.4 (−14.4–5.7)	
RT1 is the time to correct first responses, RT2 is the time from color assignment until error-correcting responses, SSRT is stop-signal reaction time, and TTS is the time between the stop-signal and the stopping of an initial response that is followed by a corrective response. Confidence intervals and significance levels were based on hierarchical regression models for all but the last row, which is based on paired-sample t-tests. All measurements are at the millisecond unit. * Significant difference.

Fig. 4 Correlations between first response and error-correction measurements. (a): The correlation between players’ flanker effects on correct first responses (RT1) and the time from color assignment until error-correcting responses (RT2) across both studies. (b): The correlation between players’ stop-signal reaction time (SSRT − transformed to the distribution of SSRTs in matching flanker trials) and the time between the stop-signal and the stopping of an initial response that is followed by a corrective response (TTS) across both studies. Dashed lines reflect the standard error around an estimate; all measurements are at the millisecond unit.

Given the assumption that players’ error-correcting responses reflect a continuation of the same cognitive functions measured with the ice and flanker effects, we expected the times from the central rat’s color assignment until correct initial responses (RT1) and from the central rat’s color assignment until error-correcting responses (RT2) to show comparable conflict effects. Furthermore, we expected conflict effects on RT1 to strongly correlate with conflict effects on RT2. To assess these expectations, we used joint hierarchical regression models with maximal random effect specification51. With random and fixed terms for the intercept, trial type, response type, and the interaction between response type and trial type.

As described in Table 2, hierarchical regression models showed significant flanker effects on RT2 in studies 1 (m = 44.2 ms, t(77.7) = 4.77, p < 0.001) and 2 (m = 58.7 ms, t(74.6) = 6.82, p < 0.001) which were not significantly different from (though quantitatively shorter than) flanker effects on RT1 in both studies 1 (m = −11.7 ms, t(78.3) = −1.25, p = 0.216) and 2 (m =−4.9 ms, t(74.7) = −0.57, p = 0.570). Furthermore, flanker effects on RT1 and RT2 strongly correlated in both studies 1 (r = 0.53, 95% CI: 0.35–0.66, t(83) = 5.62, p < 0.001) and 2 (r = 0.63, 95% CI: 0.48–0.75, t(78) = 7.22, p < 0.001), as shown in Fig. 4. These results suggested that the flanker effects on RT2 reflected a continuation of the cognitive processes measured by the game’s flanker effects on RT1.

To better understand how the flanker effect influenced RT2, we separated the flanker effect on RT2 into its constituent effects on the time it took participants to make mistaken initial responses and on the time from incorrect initial responses to error-correcting responses. Using hierarchical regression models with random and fixed terms for the intercept and trial type, we found significant flanker effects on the time to incorrect initial responses in both studies 1 (m = 39.4 ms, 95% CI: 23.9–55.0, t(71.9) = 4.97, p < .001) and 2 (m = 44.9 ms, 95% CI: 30.6–59.2, t(69.7) = 5.15, p < .001). Furthermore, we found an inconsistent flanker effect after the incorrect initial response, which significantly increased time from incorrect initial responses to error-correcting responses in mismatching trials in study 2 (m = 14.3 ms, 95% CI: 3.2–25.2, t(69.7) = 2.56, p = .013), though not in study 1 (m = 9.0 ms, 95% CI: −2.0–19.9, t(68.6) = 1.60, p = 0.114).

As described in Table 2, we found ice effects on RT2 in both studies 1 (m = 49.7 ms, t(92.4) = 6.73, p < 0.001) and 2 (m = 48.6 ms, t(87.1) = 6.73, p < 0.001), although the effects on RT2 were considerably shorter than the effects on RT1 in studies 1 (m = −35.0 ms, t(96.3) = −3.01, p = 0.003) and 2 (m = −39.9 ms, t(89.6) = −3.44, p <.001). Furthermore, ice effects on RT1 and RT2 showed a non-significant (though quantitatively negative) correlation in study 1 (r = -.15, 95% CI: −0.34–0.04, t(96) = −1.48, p = 0.141), and a significant negative correlation in study 2 (r = −0.39, 95% CI: −0.55–0.20, t(89) = −3.99, p < 0.001), precluding a strong positive correlation between the measures. These results led us to conclude that the ice effects on RT2 likely did not reflect a continuation of the cognitive processes measured by the ice effects on RT1.

Given our assumption that players’ late stopping responses reflect a continuation of the same cognitive functions responsible for their stop-signal reaction times (SSRTs), we expected players’ time-to-stop (TTS), calculated as the time from the stop-signal until the inhibition of an ongoing initial response, to show a flanker effect comparable to the one observed on SSRT. Furthermore, we expected SSRT and TTS measures to strongly correlate. To assess these expectations, we calculated SSRTs via the integration method, and players’ TTS via hierarchical regression models with fixed and random terms for intercept and trial type.

As shown in Table 2, we found flanker effects on players’ TTS in both studies 1 (m = 15.2 ms, t(90.4) = 4.61, p < 0.001) and 2 (m = 15.5 ms, t(87) = 4.2, p < 0.001), which paired-samples t-tests showed were not significantly different from (though quantitatively shorter than) the flanker effects on SSRT in studies 1 (m = −2.6 ms, t(96) = −0.46, p = 0.628) and 2 (m = −4.4 ms, t(88) = −0.88, p = 0.382). Furthermore, as shown in Fig. 4, players’ SSRT scores strongly correlated with their TTS scores in both studies 1 (r = 0.76, 95% CI: 0.65–0.83, t(95) = 11.24, p < 0.001) and 2 (r = 0.74, 95% CI: 0.63–0.82, t(87) = 10.39, p < 0.001). These results led us to conclude that players’ TTS likely reflected a continuation of the cognitive processes measured by their SSRT.

Can error-correcting responses be used alongside initial responses to improve the reliability of conflict-based cognitive measurements?

Table 3 Reliability of measurements based on first responses only, and on first responses combined with error-correcting responses.

Measurement type	Reliability: first responses only (95% CI)	Reliability: first responses and error-corrections (95% CI)	
Study 1: flanker effect	0.732 (0.605–0.805)	0.754 (0.624–0.827)	
Study 2: flanker effect	0.771 (0.619–0.850)	0.779 (0.624–0.857)	
Study 1: ice effect	0.874 (0.810–0.912)	0.847 (0.744–0.902)	
Study 2: ice effect	0.814 (0.718–0.869)	0.802 (0.686–0.864)	
Study 1: SSRT	0.831 (0.727–0.892)	0.877 (0.818–0.908)	
Study 2: SSRT	0.847 (0.728–0.898)	0.879 (0.797–0.918)	
SSRT is stop-signal reaction time. Reliability was estimated via odd-even split-halves for the ice and flanker effects and via McDonald’s ω for SSRT. Confidence intervals were estimated via bias-corrected accelerated bootstrapping with 100,000 iterations.

If error-correcting responses are indeed continuous with correct initial responses, then they should be able to enhance the psychometric properties of the measurements. Thus, we supplemented first-response-based measurements with error-correcting responses in two ways. For ice and flanker effects, we used hierarchical regression models with fixed and random effects of trial type and response type and only a fixed effect for their interaction. This model structure embodied the assumption that individual differences in conflict effects were shared between response types. For SSRT calculations, we z-transformed players’ TTS as estimated by hierarchical regression with fixed and random intercept and only a fixed trial type effect. We then averaged players’ z-transformed TTS alongside their z-transformed SSRTs calculated separately per matching and mismatching flanker trials.

As shown in Table 3, supplementing first-response-based measurements with error-correcting responses resulted in quantitatively modest increments to our measurements’ reliability that ranged from .008 to .046 for flanker effect and SSRT measurements. While we cannot generalize these results beyond these samples, they are in-line with the expectation of continuity between initial correct and error-correcting responses for these measures. The SSRT and flanker effect estimated via first responses were nearly identical to those obtained by combining first response and error-correction data (rs = 0.97). In contrast, quantitatively modest reliability reductions of −0.027 and −0.012 were seen in studies 1 and 2’s measures of the ice effects, and the two measurement approaches were not as strongly correlated (study 1: r = 0.86; study 2: r = 0.93) as before. This provides further support for discontinuity between the ice effects.

Since it is more difficult to improve the reliability of measurements with high initial reliability, the observed reliability increments require further elaboration. Using formula 3 from Kucina et al. 31 we translated our observed reliability increments into a number of additional trials and then calculated how this amount of trials would improve a typical31 reliability of 0.50. Crucially, the results of this procedure depend only on the initial reliability and how it changes. Following this procedure, we found that increments to the flanker effect’s reliability reflect as many additional trials as increments from 0.50 to 0.576 in study 1, and to 0.540 in study 2. Whereas increments to the SSRT’s reliability reflect as many additional trials as increments from 0.50 to 0.767 in study 1, and to 0.738 in study 2. These results illustrate the potential reliability benefits of cognitively-continuous error-correction data.

Can error-correcting responses be used as cognitive measurements on their own?

Table 4 Precision of the different measurements, defined as the ratio between the standard deviation of individual differences and of the measurement noise estimated by a statistical model.

Measurement	Study 1: η	Study 2: η	
Flanker effect on RT1	0.24	0.27	
Flanker effect on RT2	0.24	0.24	
Ice effect on RT1	0.50	0.44	
Ice effect on RT2	0.32	0.32	
TTS	0.73	0.66	
RT1 is the time to correct first responses, RT2 is the time to error-correcting responses, SSRT is stop-signal reaction time, and TTS is time-to-stop.

To assess the potential of error-correcting responses as separate cognitive measurements, we estimated individual differences and measurement noise using hierarchical regressions that considered first response data separately from error-correction data. For the ice and flanker effects on RT1 and RT2, the models contained fixed and random terms for intercept and trial type. For the TTS, the models included fixed and random intercepts, and only a fixed trial type term since individual differences due to trial type (the flanker effect on TTS) were minimal.

As shown in Table 4, RT2 measures of flanker and ice effects were no more precise than RT1 measures. This means that the game’s RT2 conflict effect measures, similarly to the RT1 conflict effect measures, require hundreds of response trials to achieve acceptable measurement reliability on their own. Consequently, it is not feasible to use these RT2 conflict effect measures on their own. In contrast, the TTS measure showed very high precision, sufficient to achieve excellent split-half reliability of 0.924 and 0.936 in studies 1 and 2. Which suggests that TTS can serve as a standalone measurement.

Discussion

We used behavioral data collected from Tunnel Runner to address key theoretical and psychometric challenges to the use of error-correcting responses in conflict-based cognitive assessment. Specifically, we assessed whether cognitive continuity held between first responses and error-correcting responses in the game’s conflict-based measurements, and we examined how error-correcting responses can be combined with initial responses, and whether error-correcting responses can measure cognitive functions on their own. Our results supported cognitive continuity between initial and error-correcting responses in the game’s measurements of interference control and response inhibition but not for response-rule switching. Furthermore, supplementing first-response data with error-correcting responses quantitatively increased the reliability of the flanker effect and SSRT measurements, and decreased the reliability of the ice effect measurement. Lastly, error-correcting responses showed the ability to measure response inhibition, via TTS, separately from first responses, which was not the case for interference control and response-rule switching. Overall, our results suggest that cognitive continuity between initial and error-correcting responses can extend to conflict effects, although this is not always the case. This key finding suggests that error-correcting responses can enhance or, in the case of response inhibition, even replace first-response-based conflict measures.

Explanation

We found cognitive continuity between first-response-based and error-correction-based measures of response inhibition and interference control, yet not for response-rule switching. The ice effects on error-correcting responses were nearly half the size of the effects on correct responses, and the ice effects on the two response types did not positively correlate. We speculate that this discontinuity can be explained in reference to neurophysiological indications that response-rule switching involves the initial enhancement of effortful control mechanisms that is later followed by reduced action monitoring36, and that the reconfiguration of response rules is delayed after task switching compared to other conditions52. Specifically, if delayed reduction in action monitoring translates into delayed reduction in response caution in ice trials, then less evidence would need to accumulate to initiate error-correcting responses compared to initial responses. Furthermore, delayed reconfiguration of response rules could mean that the evidence accumulation rate is increased after initial responses, as re-configuration was more likely to have already occurred. These processes could cause correct responses to be easier to make later in ice trials, potentially explaining the weaker ice effects on error-correcting responses. Furthermore, if individual differences in the impact of these processes are unrelated to or negatively associated with the initial ice effect, then these processes could explain why the ice effects on RT1 and RT2 did not positively correlate.

The continuity between first-response-based and error-correction-based measures of response inhibition and interference control should be interpreted with caution. Our results do not imply that error-correcting responses are identical to correct first responses, nor do they suggest that the cognitive processes involved in error-correcting responses are completely unchanged compared to initial responses. Our continuity results are compatible with suggestions that response boundaries37 and evidence accumulation rates38 can differ between error-correcting responses and initial responses33. What our results suggest is that if there are differences between the cognitive processes involved in the two response types, then these differences exert a small cumulative influence on conflict-based measures of interference control and response inhibition.

When cognitive continuity was otherwise shown between response types, supplementing initial responses with error-correcting responses repeatedly resulted in quantitative improvements to reliability. This provided further converging evidence for continuity, as additional data of the same type should increase reliability20. In the model-based approach we used to measure interference control via the flanker effect, error-correcting responses were treated as additional trials for measurement. Psychometrically, this means that any added value of the error-correcting responses would depend on the number and precision of these responses and the measurement’s initial reliability31. Such that the value of cognitively-continuous error-correction data might increase with the number of error-correcting responses and their measurement precision, yet diminish as the measurement’s initial reliability increases20,31. Therefore, since the flanker effects showed acceptable initial measurement reliability and the number of error-correcting responses was not large, only quantitatively modest reliability increments could be observed. However, if a measurement has high precision yet low reliability because of a limited number of correct initial responses, then cognitively-continuous error-correcting responses, if common enough, could provide considerable reliability improvements.

The combination of first-response-based response inhibition measurements with error-correcting responses was achieved by averaging three distinct measurements and thus brings its own psychometric considerations. Specifically, the TTS measurement needs to be highly reliable on its own, enabling it to strongly correlate with the other two SSRT measurements to form an internally consistent set of distinct measurements. In addition, the TTS measure, unlike the other error-correction measures, showed the ability to reliably measure response inhibition on its own using achievable amounts of error-correcting responses. This high reliability was driven by the measurement’s high precision, which was likely achieved because the TTS measure was not a difference score20,21.

Implications for theories of cognitition

While error-correcting behaviors are a common aspect of daily lives, they received limited theoretical attention in the domain of cognitive control, where the emphasis has been on highlighting continuities and discontinuities with initial responses33. Recent theoretical perspectives on cognitive control, such as Gated Cascade Diffusion53, and Binding and Retrieval in Action Control (BRAC)54, seek to explain the nature of initial responses, and how error corrections manifest earlier in the trial53 or in subsequent trials54. In particular, BRAC attempts to provide a direct explanation for various conflict effects and the relationship between performance in earlier and later trials, including post-error slowing and congruency sequence effects. However, it is difficult to apply these theoretical perspectives to error-correcting responses that follow mistaken initial responses in the same trial, as these perspectives neither attempt to explain nor are based on these types of error-correcting responses. This is unfortunate because individuals have a natural tendency to correct errors, which manifests in cognitive tasks even when error-corrections are not allowed5 and plays a role in shaping responses in subsequent trials55,56. Thus, error-correcting responses can help in understanding and linking between initial responses and trial sequence effects; crucially, error-correcting responses are worthy of consideration on their own and could play an important role in further understanding human cognition.

By examining cognitive continuity between conflict effects on initial and error-correcting responses, our results open a new path towards integrating error-correcting responses into theories of human cognition. Our results also showcase Tunnel Runner’s unique ability to evoke and measure error-correcting responses, and position it as a powerful tool for the study of error-correcting responses. This should help theories of cognition advance toward a better understanding of the dynamic nature of human cognition.

Implications for cognitive measurement

Typical conflict-based cognitive assessment tools face criticism for being boring1,4 and behaviorally restrictive3,4 while achieving limited psychometric reliability16. Cognitive games such as Tunnel Runner, which create a more fluid, dynamic, and engaging cognitive assessment experience4, can mitigate some of these issues while creating new challenges and opportunities. Since cognitive games tend to evoke a higher prevalence of incorrect first responses4,13, they lose more correct RT data and, therefore, show reduced efficiency. However, by enabling and using error-correcting responses, cognitive games may be able to compensate for much of, or even potentially exceed, the lost information. Furthermore, error-correcting responses can create new opportunities for data collection, as was the case with the time-to-stop measure. A similar measure could be implemented in other response inhibition paradigms, such as go/no-go, to efficiently provide an additional yet distinct57 way to measure response inhibition. Alternatively, time-to-stop could help assess theoretical constructs beyond SSRTs, such as failures to initiate the stop runner58.

The use of error-correction data could complement other approaches for improving the reliability of cognitive assessment. Increasing the difficulty of conflict tasks can enhance their measurement reliability31, although it may lead to more incorrect initial responses and thus greater data loss4. By considering error-correcting responses, this data loss can be minimized. Furthermore, the use of error-correcting responses requires careful use of hierarchical models, making it a natural fit for theoretically-informed hierarchical models of response data, which were shown to enhance reliability30. Overall, our results should encourage the allowance of error-correction in cognitive assessment and promote the use of the measurement opportunities created by error-correcting responses. This should help cognitive assessment tools better capture the dynamic nature of human cognition while providing participants with an improved experience4.

Implications for other types of behavioral measurements

Differences in initial responses to different types of stimuli are used in various fields and for different purposes, such as attitude assessment in consumer research and user modeling in human-computer interaction. Although these application areas often allow users to change their minds and reverse an initial decision they perceive as incorrect. Since there is no reason to expect continuity to be restricted only to cognitive measurements, our results imply that actions that reverse an initial decision (or change-of-mind data33,37) could be useful for purposes other than cognitive assessment. Thus, our results and the approach we took in defining, measuring, and utilizing change-of-mind responses and establishing continuity between different response types pave a path for the use of change of mind data in various research and application areas.

Limitations and future directions

Our studies contain several limitations, which we outline here. First, the number of players whose in-game responses could not be analyzed was high, although in line with online cognitive games studies4,13,31. Second, we did not assess the impact of using error-correcting responses on test-retest reliability nor on the prediction of other variables. We emphasize the structure of error-correcting responses and their relationship with initial responses, leaving the prediction of other measurements for future work. Third, our samples consisted mainly of gamers, meaning that our results might not generalize to non-gamers. This is a consequence of the self-selecting nature of online sampling and Tunnel Runner’s requirement for a functional graphical processing unit in a player’s computer. Fourth, we did not model error-correcting responses with evidence accumulation models, as these are not often used to assess individual differences7,59. Nevertheless, evidence accumulation models can be made to fit and explain error-correction data and could benefit from additional data5,7. Fifth, we did not consider error-correcting responses in relation to conflict effects on accuracy scores. Although accuracy levels are an integral part of many conflict-based measurements60, we did not find a good way to conceptualize the use of error-correcting responses for accuracy measures. Lastly, we did not model in-game task-switching costs, which could add noise to our continuity results. An approach to modeling task-switching with Tunnel Runner’s multiple elements would be complex and first needs to be established. Thus, future work could consider the impact of error-correcting responses on test-retest reliability, assess cognitive continuity using different behavioral measurements, and/or develop better methods to use error-correcting responses for cognitive assessment, including accuracy scores, potentially via computational modeling that account for trial-switching costs, or via time-to-event analyses.

Conclusions

We demonstrated that error-correcting responses can be cognitively continuous with initial responses for conflict-based measurements of interference control and response inhibition, such that error-correcting responses can supplement or even replace first-response-based cognitive measurements. However, cognitive discontinuity can also apply in some conflict-based measurements, as was the case for response-rule switching, in which case error-correcting responses should not be used alongside or instead of initial responses. These results improve our understanding of the dynamics of human cognition that go beyond initial responses, call further theoretical attention toward error-correcting responses, and pave a path toward the use of change of mind data for cognitive assessment and other research and application areas.

Supplementary Information

Supplementary Information.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-024-71762-z.

Acknowledgements

This publication is part of the project “Game-based Digital Biomarkers for Acute and Chronic Stress” (VI.Veni.202.171) of the research programme NWO Talent Programme VENI, which is financed by the Dutch Research Council (NWO).

Author contributions

B.M.: Conceptualization; conceptual visualization; data collection; data analysis; data visualization; writing—original draft. N.J.E.: Conceptualization; conceptual visualization; writing—reviewing and editing. M.V.B.: Conceptualization; supervision; resources; writing—reviewing and editing.

Data availability

Data and scripts are available as Supplementary material and at https://osf.io/qjhwv/. Questions should be addressed to the first author. A public demo of the game is available at https://tunnel-runner.itch.io/tunnel-runner-demo.

Competing interests

The authors declare no competing interests.

Ethical approval

All studies were carried out according to the Declaration of Helsinki and approved by the Ethical Review Board of the Department of Industrial Design at the University of Eindhoven.

Consent to participate

Informed consent was obtained from all participants.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Meier, M., Martarelli, C. & Wolff, W. Bored participants, biased data? How boredom can influence behavioral science research and what we can do about it. 10.31234/osf.io/hzfqr (2023).
2. Ono T Sakurai T Kasuno S Murai T Novel 3-D action video game mechanics reveal differentiable cognitive constructs in young players, but not in old Sci. Rep. 2022 12 11751 10.1038/s41598-022-15679-5 35864114
Ono, T., Sakurai, T., Kasuno, S. & Murai, T. Novel 3-D action video game mechanics reveal differentiable cognitive constructs in young players, but not in old. Sci. Rep. 12, 11751. 10.1038/s41598-022-15679-5 (2022).35864114 10.1038/s41598-022-15679-5
3. Shamay-Tsoory SG Mendelsohn A Real-life neuroscience: An ecological approach to brain and behavior research Perspect. Psychol. Sci. 2019 14 841 859 10.1177/1745691619856350 31408614
Shamay-Tsoory, S. G. & Mendelsohn, A. Real-life neuroscience: An ecological approach to brain and behavior research. Perspect. Psychol. Sci. 14, 841–859. 10.1177/1745691619856350 (2019).31408614 10.1177/1745691619856350
4. Markovitch, B., Markopoulos, P. & Birk, M. V. Tunnel Runner: a Proof-of-principle for the feasibility and benefits of facilitating players’ sense of control in cognitive assessment games. In Proceedings of the CHI Conference on Human Factors in Computing Systems, CHI ’24, 1–18. (Association for Computing Machinery, New York, NY, USA, 2024). 10.1145/3613904.3642418
5. Evans NJ Dutilh G Wagenmakers E-J Van Der Maas HL Double responding: A new constraint for models of speeded decision making Cognit. Psychol. 2020 121 101292 10.1016/j.cogpsych.2020.101292 32217348
Evans, N. J., Dutilh, G., Wagenmakers, E.-J. & Van Der Maas, H. L. Double responding: A new constraint for models of speeded decision making. Cognit. Psychol. 121, 101292. 10.1016/j.cogpsych.2020.101292 (2020).32217348 10.1016/j.cogpsych.2020.101292
6. Taylor GJ Nguyen AT Evans NJ Does allowing for changes of mind influence initial responses? Psychon. Bull. Rev. 2023 10.3758/s13423-023-02371-6 37884778
Taylor, G. J., Nguyen, A. T. & Evans, N. J. Does allowing for changes of mind influence initial responses?. Psychon. Bull. Rev. 10.3758/s13423-023-02371-6 (2023).37884778 10.3758/s13423-023-02371-6
7. Evans NJ Wagenmakers E-J Evidence accumulation models: Current limitations and future directions Quant. Methods Psychol. 2020 16 73 90 10.20982/tqmp.16.2.p073
Evans, N. J. & Wagenmakers, E.-J. Evidence accumulation models: Current limitations and future directions. Quant. Methods Psychol. 16, 73–90. 10.20982/tqmp.16.2.p073 (2020).10.20982/tqmp.16.2.p073
8. Friehs MA Dechant M Vedress S Frings C Mandryk RL Effective gamification of the stop-signal task: Two controlled laboratory experiments JMIR Serious Games 2020 8 e17810 10.2196/17810 32897233
Friehs, M. A., Dechant, M., Vedress, S., Frings, C. & Mandryk, R. L. Effective gamification of the stop-signal task: Two controlled laboratory experiments. JMIR Serious Games 8, e17810. 10.2196/17810 (2020).32897233 10.2196/17810
9. Lumsden J Skinner A Coyle D Lawrence N Munafo M Attrition from web-based cognitive testing: A repeated measures comparison of gamification techniques J. Med. Internet Res. 2017 19 e8473 10.2196/jmir.8473
Lumsden, J., Skinner, A., Coyle, D., Lawrence, N. & Munafo, M. Attrition from web-based cognitive testing: A repeated measures comparison of gamification techniques. J. Med. Internet Res. 19, e8473. 10.2196/jmir.8473 (2017).10.2196/jmir.8473
10. Lumsden J Skinner A Woods AT Lawrence NS Munafò M The effects of gamelike features and test location on cognitive test performance and participant enjoyment PeerJ 2016 4 e2184 10.7717/peerj.2184 27441120
Lumsden, J., Skinner, A., Woods, A. T., Lawrence, N. S. & Munafò, M. The effects of gamelike features and test location on cognitive test performance and participant enjoyment. PeerJ 4, e2184. 10.7717/peerj.2184 (2016).27441120 10.7717/peerj.2184
11. Miranda AT Palmer EM Intrinsic motivation and attentional capture from gamelike features in a visual search task Behav. Res. Methods 2014 46 159 172 10.3758/s13428-013-0357-7 23835649
Miranda, A. T. & Palmer, E. M. Intrinsic motivation and attentional capture from gamelike features in a visual search task. Behav. Res. Methods 46, 159–172. 10.3758/s13428-013-0357-7 (2014).23835649 10.3758/s13428-013-0357-7
12. Szalma JL Schmidt TN Teo GWL Hancock PA Vigilance on the move: Video game-based measurement of sustained attention Ergonomics 2014 57 1315 1336 10.1080/00140139.2014.921329 25001010
Szalma, J. L., Schmidt, T. N., Teo, G. W. L. & Hancock, P. A. Vigilance on the move: Video game-based measurement of sustained attention. Ergonomics 57, 1315–1336. 10.1080/00140139.2014.921329 (2014).25001010 10.1080/00140139.2014.921329
13. Wiley, K., Vedress, S. & Mandryk, R. L. How Points and Theme Affect Performance and Experience in a Gamified Cognitive Task. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, CHI ’20, 1–15, 10.1145/3313831.3376697 (Association for Computing Machinery, New York, NY, USA, 2020).
14. Wiley K Berger P Friehs MA Mandryk RL Measuring the reliability of a gamified stroop task: Quantitative experiment JMIR Serious Games 2024 12 e50315 10.2196/50315 38598265
Wiley, K., Berger, P., Friehs, M. A. & Mandryk, R. L. Measuring the reliability of a gamified stroop task: Quantitative experiment. JMIR Serious Games 12, e50315. 10.2196/50315 (2024).38598265 10.2196/50315
15. van den Berg R A common mechanism underlies changes of mind about decisions and confidence eLife 2016 5 e12192 10.7554/eLife.12192 26829590
van den Berg, R. et al. A common mechanism underlies changes of mind about decisions and confidence. eLife 5, e12192. 10.7554/eLife.12192 (2016).26829590 10.7554/eLife.12192
16. Hedge C Powell G Sumner P The reliability paradox: Why robust cognitive tasks do not produce reliable individual differences Behav. Res. Methods 2018 50 1166 1186 10.3758/s13428-017-0935-1 28726177
Hedge, C., Powell, G. & Sumner, P. The reliability paradox: Why robust cognitive tasks do not produce reliable individual differences. Behav. Res. Methods 50, 1166–1186. 10.3758/s13428-017-0935-1 (2018).28726177 10.3758/s13428-017-0935-1
17. Logan, G.D. On the ability to inhibit thought and action: a user's guide to the stop signal paradigm. in Inhibitory Processes in Attention, Memory and Language (eds. Dagenbach, D. & Carr, T.H.) 189–236 (Academic Press, San Diego, 1994).
18. Verbruggen F A consensus guide to capturing the ability to inhibit actions and impulsive behaviors in the stop-signal task eLife 2019 8 e46323 10.7554/eLife.46323 31033438
Verbruggen, F. et al. A consensus guide to capturing the ability to inhibit actions and impulsive behaviors in the stop-signal task. eLife 8, e46323. 10.7554/eLife.46323 (2019).31033438 10.7554/eLife.46323
19. Eriksen BA Eriksen CW Effects of noise letters upon the identification of a target letter in a nonsearch task Percept. Psychophys. 1974 16 143 149 10.3758/BF03203267
Eriksen, B. A. & Eriksen, C. W. Effects of noise letters upon the identification of a target letter in a nonsearch task. Percept. Psychophys. 16, 143–149. 10.3758/BF03203267 (1974).10.3758/BF03203267
20. Zorowitz S Niv Y Improving the reliability of cognitive task measures: A narrative review Biol. Psychiatry Cognit. Neurosci. Neuroimaging 2023 8 789 797 10.1016/j.bpsc.2023.02.004 36842498
Zorowitz, S. & Niv, Y. Improving the reliability of cognitive task measures: A narrative review. Biol. Psychiatry Cognit. Neurosci. Neuroimaging 8, 789–797. 10.1016/j.bpsc.2023.02.004 (2023).36842498 10.1016/j.bpsc.2023.02.004
21. Overall JE Woodward JA Unreliability of difference scores: A paradox for measurement of change Psychol. Bull. 1975 82 85 86 10.1037/h0076158
Overall, J. E. & Woodward, J. A. Unreliability of difference scores: A paradox for measurement of change. Psychol. Bull. 82, 85–86. 10.1037/h0076158 (1975).10.1037/h0076158
22. Diamond A Executive functions Ann. Rev. Psychol. 2013 64 135 168 10.1146/annurev-psych-113011-143750 23020641
Diamond, A. Executive functions. Ann. Rev. Psychol. 64, 135–168. 10.1146/annurev-psych-113011-143750 (2013).23020641 10.1146/annurev-psych-113011-143750
23. Rae CL Response inhibition on the stop signal task improves during cardiac contraction Sci. Rep. 2018 8 9136 10.1038/s41598-018-27513-y 29904123
Rae, C. L. et al. Response inhibition on the stop signal task improves during cardiac contraction. Sci. Rep. 8, 9136. 10.1038/s41598-018-27513-y (2018).29904123 10.1038/s41598-018-27513-y
24. Friehs MA No effects of 1 Hz offline TMS on performance in the stop-signal game Sci. Rep. 2023 13 11565 10.1038/s41598-023-38841-z 37463991
Friehs, M. A. et al. No effects of 1 Hz offline TMS on performance in the stop-signal game. Sci. Rep. 13, 11565. 10.1038/s41598-023-38841-z (2023).37463991 10.1038/s41598-023-38841-z
25. Brunetti M Zappasodi F Croce P Di Matteo R Parsing the Flanker task to reveal behavioral and oscillatory correlates of unattended conflict interference Sci. Rep. 2019 9 13883 10.1038/s41598-019-50464-x 31554881
Brunetti, M., Zappasodi, F., Croce, P. & Di Matteo, R. Parsing the Flanker task to reveal behavioral and oscillatory correlates of unattended conflict interference. Sci. Rep. 9, 13883. 10.1038/s41598-019-50464-x (2019).31554881 10.1038/s41598-019-50464-x
26. Montalti M Mirabella G Unveiling the influence of task-relevance of emotional faces on behavioral reactions in a multi-face context using a novel Flanker-Go/No-go task Sci. Rep. 2023 13 20183 10.1038/s41598-023-47385-1 37978229
Montalti, M. & Mirabella, G. Unveiling the influence of task-relevance of emotional faces on behavioral reactions in a multi-face context using a novel Flanker-Go/No-go task. Sci. Rep. 13, 20183. 10.1038/s41598-023-47385-1 (2023).37978229 10.1038/s41598-023-47385-1
27. Xie L Ren M Cao B Li F Distinct brain responses to different inhibitions: Evidence from a modified Flanker task Sci. Rep. 2017 7 6657 10.1038/s41598-017-04907-y 28751739
Xie, L., Ren, M., Cao, B. & Li, F. Distinct brain responses to different inhibitions: Evidence from a modified Flanker task. Sci. Rep. 7, 6657. 10.1038/s41598-017-04907-y (2017).28751739 10.1038/s41598-017-04907-y
28. Morris SE Cuthbert BN Research domain criteria: Cognitive systems, neural circuits, and dimensions of behavior Dialog. Clin. Neurosci. 2012 14 29 37 10.31887/DCNS.2012.14.1/smorris
Morris, S. E. & Cuthbert, B. N. Research domain criteria: Cognitive systems, neural circuits, and dimensions of behavior. Dialog. Clin. Neurosci. 14, 29–37. 10.31887/DCNS.2012.14.1/smorris (2012).10.31887/DCNS.2012.14.1/smorris
29. Research Domain Criteria (RDoC) - National Institute of Mental Health (NIMH). https://www.nimh.nih.gov/research/research-funded-by-nimh/rdoc.
30. Haines, N. et al. Theoretically informed generative models can advance the psychological and brain sciences: lessons from the reliability paradox. Preprint at PsyArXiv 10.31234/osf.io/xr7y3 (2020).
31. Kucina T Calibration of cognitive tests to address the reliability paradox for decision-conflict tasks Nat. Commun. 2023 14 2234 10.1038/s41467-023-37777-2 37076456
Kucina, T. et al. Calibration of cognitive tests to address the reliability paradox for decision-conflict tasks. Nat. Commun. 14, 2234. 10.1038/s41467-023-37777-2 (2023).37076456 10.1038/s41467-023-37777-2
32. Ratcliff R A theory of memory retrieval Psychol. Rev. 1978 85 59 108 10.1037/0033-295X.85.2.59
Ratcliff, R. A theory of memory retrieval. Psychol. Rev. 85, 59–108. 10.1037/0033-295X.85.2.59 (1978).10.1037/0033-295X.85.2.59
33. Stone C Mattingley JB Rangelov D On second thoughts: Changes of mind in decision-making Trends Cognit. Sci. 2022 26 419 431 10.1016/j.tics.2022.02.004 35279383
Stone, C., Mattingley, J. B. & Rangelov, D. On second thoughts: Changes of mind in decision-making. Trends Cognit. Sci. 26, 419–431. 10.1016/j.tics.2022.02.004 (2022).35279383 10.1016/j.tics.2022.02.004
34. Rabbitt P Rodgers B What does a man do after he makes an error? an analysis of response programming Q. J. Exp. Psychol. 1977 29 727 743 10.1080/14640747708400645
Rabbitt, P. & Rodgers, B. What does a man do after he makes an error? an analysis of response programming. Q. J. Exp. Psychol. 29, 727–743. 10.1080/14640747708400645 (1977).10.1080/14640747708400645
35. Vickers D Lee MD Dynamic models of simple judgments: II. Properties of a self-organizing PAGAN (Parallel, adaptive, generalized accumulator network) model for multi-choice tasks Nonlinear Dyn. Psychol. Life Sci. 2000 4 1 31 10.1023/A:1009571011764
Vickers, D. & Lee, M. D. Dynamic models of simple judgments: II. Properties of a self-organizing PAGAN (Parallel, adaptive, generalized accumulator network) model for multi-choice tasks. Nonlinear Dyn. Psychol. Life Sci. 4, 1–31. 10.1023/A:1009571011764 (2000).10.1023/A:1009571011764
36. Schroder HS Moran TP Moser JS Altmann EM When the rules are reversed: Action-monitoring consequences of reversing stimulus-response mappings Cognit. Affect. Behav. Neurosci. 2012 12 629 643 10.3758/s13415-012-0105-y 22797946
Schroder, H. S., Moran, T. P., Moser, J. S. & Altmann, E. M. When the rules are reversed: Action-monitoring consequences of reversing stimulus-response mappings. Cognit. Affect. Behav. Neurosci. 12, 629–643. 10.3758/s13415-012-0105-y (2012).22797946 10.3758/s13415-012-0105-y
37. Resulaj A Kiani R Wolpert DM Shadlen MN Changes of mind in decision-making Nature 2009 461 263 266 10.1038/nature08275 19693010
Resulaj, A., Kiani, R., Wolpert, D. M. & Shadlen, M. N. Changes of mind in decision-making. Nature 461, 263–266. 10.1038/nature08275 (2009).19693010 10.1038/nature08275
38. Bronfman ZZ Decisions reduce sensitivity to subsequent information Proc. R. Soc. B Biol. Sci. 2015 282 20150228 10.1098/rspb.2015.0228
Bronfman, Z. Z. et al. Decisions reduce sensitivity to subsequent information. Proc. R. Soc. B Biol. Sci. 282, 20150228. 10.1098/rspb.2015.0228 (2015).10.1098/rspb.2015.0228
39. Litman L Robinson J Conducting Online Research on Amazon Mechanical Turk and Beyond 2020 Washington SAGE Publications
Litman, L. & Robinson, J. Conducting Online Research on Amazon Mechanical Turk and Beyond (SAGE Publications, Washington, 2020).
40. Bates, D., Mächler, M., Bolker, B. & Walker, S. Fitting Linear Mixed-Effects Models using lme4. (2014). http://arxiv.org/abs/1406.5823.
41. Huang FL Li X Using cluster-robust standard errors when analyzing group-randomized trials with few clusters Behav. Res. Methods 2022 54 1181 1199 10.3758/s13428-021-01627-0 34505994
Huang, F. L. & Li, X. Using cluster-robust standard errors when analyzing group-randomized trials with few clusters. Behav. Res. Methods 54, 1181–1199. 10.3758/s13428-021-01627-0 (2022).34505994 10.3758/s13428-021-01627-0
42. Pustejovsky JE Tipton E Small-sample methods for cluster-robust variance estimation and hypothesis testing in fixed effects models J. Bus. Econ. Stat. 2018 36 672 683 10.1080/07350015.2016.1247004
Pustejovsky, J. E. & Tipton, E. Small-sample methods for cluster-robust variance estimation and hypothesis testing in fixed effects models. J. Bus. Econ. Stat. 36, 672–683. 10.1080/07350015.2016.1247004 (2018).10.1080/07350015.2016.1247004
43. Chen G Trial and error: A hierarchical modeling approach to test-retest reliability NeuroImage 2021 245 118647 10.1016/j.neuroimage.2021.118647 34688897
Chen, G. et al. Trial and error: A hierarchical modeling approach to test-retest reliability. NeuroImage 245, 118647. 10.1016/j.neuroimage.2021.118647 (2021).34688897 10.1016/j.neuroimage.2021.118647
44. Haines N Sullivan-Toole H Olino T From classical methods to generative models: Tackling the unreliability of neuroscientific measures in mental health research Biol. Psychiatry Cognit. Neurosci. Neuroimaging 2023 10.1016/j.bpsc.2023.01.001
Haines, N., Sullivan-Toole, H. & Olino, T. From classical methods to generative models: Tackling the unreliability of neuroscientific measures in mental health research. Biol. Psychiatry Cognit. Neurosci. Neuroimaging 10.1016/j.bpsc.2023.01.001 (2023).10.1016/j.bpsc.2023.01.001
45. Littman R Hochman S Kalanthroff E Reliable affordances: A generative modeling approach for test-retest reliability of the affordances task Behav. Res. Methods 2023 10.3758/s13428-023-02131-3 37127802
Littman, R., Hochman, S. & Kalanthroff, E. Reliable affordances: A generative modeling approach for test-retest reliability of the affordances task. Behav. Res. Methods 10.3758/s13428-023-02131-3 (2023).37127802 10.3758/s13428-023-02131-3
46. Schielzeth H Robustness of linear mixed-effects models to violations of distributional assumptions Methods Ecol. Evol. 2020 11 1141 1152 10.1111/2041-210X.13434
Schielzeth, H. et al. Robustness of linear mixed-effects models to violations of distributional assumptions. Methods Ecol. Evol. 11, 1141–1152. 10.1111/2041-210X.13434 (2020).10.1111/2041-210X.13434
47. Dunn TJ Baguley T Brunsden V From alpha to omega: A practical solution to the pervasive problem of internal consistency estimation Br. J. Psychol. 2014 105 399 412 10.1111/bjop.12046 24844115
Dunn, T. J., Baguley, T. & Brunsden, V. From alpha to omega: A practical solution to the pervasive problem of internal consistency estimation. Br. J. Psychol. 105, 399–412. 10.1111/bjop.12046 (2014).24844115 10.1111/bjop.12046
48. Drost EA Validity and reliability in social science research Educ. Res. Perspect. 2020 38 105 123 10.3316/informit.491551710186460
Drost, E. A. Validity and reliability in social science research. Educ. Res. Perspect. 38, 105–123. 10.3316/informit.491551710186460 (2020).10.3316/informit.491551710186460
49. Hayes AF Coutts JJ Use omega rather than Cronbach’s alpha for estimating reliability. But... Commun. Methods Meas. 2020 14 1 24 10.1080/19312458.2020.1718629
Hayes, A. F. & Coutts, J. J. Use omega rather than Cronbach’s alpha for estimating reliability. But.... Commun. Methods Meas. 14, 1–24. 10.1080/19312458.2020.1718629 (2020).10.1080/19312458.2020.1718629
50. Kelley K Pornprasertmanit S Confidence intervals for population reliability coefficients: Evaluation of methods, recommendations, and software for composite measures Psychol. Methods 2016 21 69 92 10.1037/a0040086 26962759
Kelley, K. & Pornprasertmanit, S. Confidence intervals for population reliability coefficients: Evaluation of methods, recommendations, and software for composite measures. Psychol. Methods 21, 69–92. 10.1037/a0040086 (2016).26962759 10.1037/a0040086
51. Barr DJ Levy R Scheepers C Tily HJ Random effects structure for confirmatory hypothesis testing: Keep it maximal J. Mem. Lang. 2013 68 255 278 10.1016/j.jml.2012.11.001
Barr, D. J., Levy, R., Scheepers, C. & Tily, H. J. Random effects structure for confirmatory hypothesis testing: Keep it maximal. J. Mem. Lang. 68, 255–278. 10.1016/j.jml.2012.11.001 (2013).10.1016/j.jml.2012.11.001
52. Steinhauser M Maier ME Ernst B Neural correlates of reconfiguration failure reveal the time course of task-set reconfiguration Neuropsychologia 2017 106 100 111 10.1016/j.neuropsychologia.2017.09.018 28939202
Steinhauser, M., Maier, M. E. & Ernst, B. Neural correlates of reconfiguration failure reveal the time course of task-set reconfiguration. Neuropsychologia 106, 100–111. 10.1016/j.neuropsychologia.2017.09.018 (2017).28939202 10.1016/j.neuropsychologia.2017.09.018
53. Dendauw E The gated cascade diffusion model: An integrated theory of decision making, motor preparation, and motor execution Psychol. Rev. 2024 10.1037/rev0000464 38386394
Dendauw, E. et al. The gated cascade diffusion model: An integrated theory of decision making, motor preparation, and motor execution. Psychol. Rev. 10.1037/rev0000464 (2024).38386394 10.1037/rev0000464
54. Frings C Binding and retrieval in action control (BRAC) Trends Cognit. Sci. 2020 24 375 387 10.1016/j.tics.2020.02.004 32298623
Frings, C. et al. Binding and retrieval in action control (BRAC). Trends Cognit. Sci. 24, 375–387. 10.1016/j.tics.2020.02.004 (2020).32298623 10.1016/j.tics.2020.02.004
55. Steinhauser M How to correct a task error: Task-switch effects following different types of error correction J. Exp. Psychol. Learn. Mem. Cognit. 2010 36 1028 1035 10.1037/a0019340 20565218
Steinhauser, M. How to correct a task error: Task-switch effects following different types of error correction. J. Exp. Psychol. Learn. Mem. Cognit. 36, 1028–1035. 10.1037/a0019340 (2010).20565218 10.1037/a0019340
56. Beatty PJ Buzzell GA Roberts DM Voloshyna Y McDonald CG Subthreshold error corrections predict adaptive post-error compensations Psychophysiology 2021 58 e13803 10.1111/psyp.13803 33709470
Beatty, P. J., Buzzell, G. A., Roberts, D. M., Voloshyna, Y. & McDonald, C. G. Subthreshold error corrections predict adaptive post-error compensations. Psychophysiology 58, e13803. 10.1111/psyp.13803 (2021).33709470 10.1111/psyp.13803
57. Littman R Takacs A Do all inhibitions act alike? A study of go/no-go and stop-signal paradigms PLOS One 2017 12 e0186774 10.1371/journal.pone.0186774 29065184
Littman, R. & Takacs, A. Do all inhibitions act alike? A study of go/no-go and stop-signal paradigms. PLOS One 12, e0186774. 10.1371/journal.pone.0186774 (2017).29065184 10.1371/journal.pone.0186774
58. Matzke D Love J Heathcote A A Bayesian approach for estimating the probability of trigger failures in the stop-signal paradigm Behav. Res. Methods 2017 49 267 281 10.3758/s13428-015-0695-8 26822670
Matzke, D., Love, J. & Heathcote, A. A Bayesian approach for estimating the probability of trigger failures in the stop-signal paradigm. Behav. Res. Methods 49, 267–281. 10.3758/s13428-015-0695-8 (2017).26822670 10.3758/s13428-015-0695-8
59. Evans NJ Steyvers M Brown SD Modeling the covariance structure of complex datasets using cognitive models: An application to individual differences and the heritability of cognitive ability Cognit. Sci. 2018 42 1925 1944 10.1111/cogs.12627
Evans, N. J., Steyvers, M. & Brown, S. D. Modeling the covariance structure of complex datasets using cognitive models: An application to individual differences and the heritability of cognitive ability. Cognit. Sci. 42, 1925–1944. 10.1111/cogs.12627 (2018).10.1111/cogs.12627
60. Liesefeld HR Janczyk M Combining speed and accuracy to control for speed-accuracy trade-offs(?) Behav. Res. Methods 2019 51 40 60 10.3758/s13428-018-1076-x 30022459
Liesefeld, H. R. & Janczyk, M. Combining speed and accuracy to control for speed-accuracy trade-offs(?). Behav. Res. Methods 51, 40–60. 10.3758/s13428-018-1076-x (2019).30022459 10.3758/s13428-018-1076-x
