==== Front PLoS One PLoS One plos plosone PLoS ONE 1932-6203 Public Library of Science San Francisco, CA USA 10.1371/journal.pone.0243435 PONE-D-19-34461 Research Article Social Sciences Linguistics Phonetics Vowels Biology and Life Sciences Neuroscience Cognitive Science Cognitive Psychology Academic Skills Literacy Biology and Life Sciences Psychology Cognitive Psychology Academic Skills Literacy Social Sciences Psychology Cognitive Psychology Academic Skills Literacy Social Sciences Linguistics Grammar Phonology People and Places Population Groupings Age Groups Children People and Places Population Groupings Families Children Biology and Life Sciences Neuroscience Cognitive Science Cognitive Psychology Learning Human Learning Biology and Life Sciences Psychology Cognitive Psychology Learning Human Learning Social Sciences Psychology Cognitive Psychology Learning Human Learning Biology and Life Sciences Neuroscience Learning and Memory Learning Human Learning Biology and Life Sciences Neuroscience Cognitive Science Cognitive Neuroscience Learning Disabilities Dyslexia Biology and Life Sciences Neuroscience Cognitive Neuroscience Learning Disabilities Dyslexia Computer and Information Sciences Software Engineering Computer Software Apps Engineering and Technology Software Engineering Computer Software Apps People and Places Population Groupings Professions Teachers Annotating digital text with phonemic cues to support decoding in struggling readers ANNOTATING DIGITAL TEXT WITH PHONEMIC CUEShttps://orcid.org/0000-0002-5850-7248Donnelly Patrick M. ConceptualizationData curationFormal analysisInvestigationMethodologyProject administrationResourcesSoftwareSupervisionValidationVisualizationWriting – original draftWriting – review & editing12* Larson Kevin ConceptualizationResourcesSoftware3 Matskewich Tanya ConceptualizationSoftware3 Yeatman Jason D. ConceptualizationFunding acquisitionMethodologyProject administrationResourcesSupervisionWriting – review & editing45 1 Department of Speech & Hearing Sciences, University of Washington, Seattle, Washington, United States of America 2 Institute for Learning & Brain Sciences, University of Washington, Seattle, Washington, United States of America 3 Microsoft Corporation, Redmond, Washington, United States of America 4 Graduate School of Education, Stanford University, Stanford, California, United States of America 5 Division of Developmental-Behavioral Pediatrics, Stanford University School of Medicine, Stanford, California, United States of America Männel Claudia Editor Max-Planck-Institut fur Kognitions- und Neurowissenschaften, GERMANY Competing Interests: Two authors [K.L. and T.M.] are employed by Microsoft Corporation. Microsoft corporation only provided financial support in the form of salaries [K.L. and T.M.] and research materials. This does not alter our adherence to PLOS ONE policies on sharing data and materials. * E-mail: pdonne@uw.edu 7 12 2020 2020 15 12 e024343513 12 2019 23 11 2020 © 2020 Donnelly et al2020Donnelly et alThis is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.An advantage of digital media is the flexibility to personalize the presentation of text to an individual’s needs and embed tools that support pedagogy. The goal of this study was to develop a tablet-based reading tool, grounded in the principles of phonics-based instruction, and determine whether struggling readers could leverage this technology to decode challenging words. The tool presents a small icon below each vowel to represent its sound. Forty struggling child readers were randomly assigned to an intervention or control group to test the efficacy of the phonemic cues. We found that struggling readers could leverage the cues to improve pseudoword decoding: after two weeks of practice, the intervention group showed greater improvement than controls. This study demonstrates the potential of a text annotation, grounded in intervention research, to help children decode novel words. These results highlight the opportunity for educational technologies to support and supplement classroom instruction. http://dx.doi.org/10.13039/100000169Division of Behavioral and Cognitive Sciences1551330Yeatman Jason D. http://dx.doi.org/10.13039/100000071National Institute of Child Health and Human DevelopmentR21HD092771Yeatman Jason D. http://dx.doi.org/10.13039/100006112Microsoft ResearchYeatman Jason D. http://dx.doi.org/10.13039/501100003986Jacobs FoundationYeatman Jason D. This work was funded by NSF BCS 1551330, NICHD R21HD092771, Microsoft Research Grants and Jacobs Foundation Research Fellowship to J.D.Y. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Two authors [K.L. and T.M], are employed by Microsoft Corporation. Microsoft Corporation provided support in the form of salaries for authors [K.L. and T.M.] and research materials but did not have any additional role in the study design, data collection and analysis, decision to publish, or preparation of the manuscript. The specific roles of these authors are articulated in the ‘author contributions’ section. Data AvailabilityAll data and analysis code associated with this manuscript are publicly accessible at the following link: [https://github.com/patdonnelly/Donnelly_2019_PLOSONE]Data Availability All data and analysis code associated with this manuscript are publicly accessible at the following link: [https://github.com/patdonnelly/Donnelly_2019_PLOSONE] ==== Body Introduction The discrepancy between the value that society places on literacy and reading achievement levels in American youth [1] is a source of concern both among policy makers and scientists [2, 3]. Developmental dyslexia, a learning disability that impacts reading, is widespread, affecting between 5–17% of the population [4, 5]. Beyond dyslexia, poor literacy rates are a nationwide issue, with 34% of fourth graders performing below the Basic level on national achievement tests [6]. Together, these results paint a troubling landscape of literacy achievement and illuminate a non-trivial need for expanding access to evidence-based instruction and intervention. For many families of struggling readers, access to high quality, evidence-based interventions outside of school are not only limited, but also represent a significant financial burden [7]. Even with a diagnosis, children with reading disabilities struggle to find the support they need in their typical classrooms, necessitating supplemental, after-school programs [8]. For this reason, and with the ever-growing landscape of educational technologies, families are turning to digital alternatives. Many apps and technologies are now widely marketed to augment, or even replace, the teacher in delivering intervention for struggling readers. These new tools are advertised as educational, presented as games, and, in general, are fun to use. Although promising, many of these tools ignore the evidence base on effective instruction and intervention techniques and, of the many educational technologies that are available today, very few have scientific studies testing their efficacy [9–11]. In 2012, it was estimated that hundreds of thousands of educational apps had been released on the Apple iOS app store [10]. According to a RAND report, children between the ages of three and five years-old spend, on average, four hours per day interacting with communications technology (i.e., smart phones, tablets, etc.) [12]. The Joan Ganz Cooney Center reported, based on a survey of parents, that 35% of children aged two to ten years-old use educational apps at least once per week. Furthermore, federal funds are being devoted to bringing Internet and devices to our nation’s schools, providing infrastructure for even greater involvement with digital media in education. However, despite the exciting potential offered by educational technology, parents and educators alike feel overwhelmed by the plethora of options and the lack of guidelines surrounding technologies advertised as educational [10]. Decades of scientific research into the behavioral and neural mechanisms of literacy learning has led to the development and testing of effective intervention programs for struggling readers, and established comprehensive guidelines and best-practices for implementation of an effective curriculum [2, 13–15]. Unfortunately, these evidence-based practices (e.g., phonemic awareness, phonics) are largely not being incorporated into the current technological boom. For example, a consistent finding in the intervention literature is that children with dyslexia benefit from direct instruction in phonological awareness and curricula that make clear links between orthography and phonology whereas children with stronger reading skills can often infer grapheme-phoneme correspondences without direct instruction (for review of the extensive literature on the importance of phonics/phonemic awareness see [14, 16–22]). From a sample of 184 apps compiled from online lists of award-winning or highly-rated apps, researchers discovered that although they were entertaining, they lacked scientific backing and “their content, design, production, and distribution are […] an incomplete response to children’s literacy needs, especially for struggling readers” [10]. Thus, there is great need for researchers studying literacy development and reading disabilities to work with tech developers on the design of tools that are grounded in the extensive scientific literature on what works for struggling readers, systematically test their effectiveness, and contribute to the development of standards of practice for educational apps targeted at literacy. Recent metanalyses demonstrate much promise for digital solutions in the context of literacy, yet also describe the multitude of ways that technology is an inappropriate substitute for many aspects of pedagogy [23, 24]. Namely, these meta analyses demonstrate that technologies focused on supplementing what is provided in the evidence-based classroom (i.e. explicit phonics), rather than restructuring at the classroom level, have demonstrated the most promise in the digital landscape. These findings, however, should be interpreted with caution as the authors further contend that the preponderance of studies in this area are characterized by small samples and poor study design [24]. In parallel with research outside of technology, the onus for researchers is to rigorously test digital solutions to discover “what works best in which programs for what students under which conditions” [25]. As a means of scaffolding learning in the classroom, technological tools provide limitless practice [26–28] that can be individualized [24, 29, 30] to optimize for an individual reader’s strengths and weaknesses. Additionally, modern computational tools utilizing speech recognition and synthesis application programming interfaces (APIs), allow for embedded tools to be provided in real-time for any given text. One promising avenue for such technology has been the use of embedded support features such as visual images to facilitate learning sound-symbol correspondence. For example, Trainertext, a program in the United Kingdom which provides visual mnemonics above each phoneme in a given text, saw significant improvement after 10 months of exposure and at-home practice [31]. By providing a visual scaffold, the authors concluded that readers could leverage a phonics-based text adaptation to improve their decoding skill. Moreover, SeeWord Reading app, a digital tool which uses a picture-embedded font to demonstrate grapheme-phoneme correspondence [32], has demonstrated efficacy in studies performed both in Singapore and the United States [33]. Both Trainertext and SeeWord Reading, demonstrate the utility of an embedded support to provide concrete visual relationships between written text and spoken language: a foundational skill for literacy. Together these studies demonstrate that technologies focusing at the level of the phoneme and the syllable (or rime) not only benefit reading performance, but also adhere to known learning and pedagogical principles [33, 34]. This paper outlines a collaboration between the Brain Development & Education Lab at the University of Washington and the Learning Tools team at Microsoft to develop Sound It Out, a web-based app that annotates text with phonemic image cues to assist in decoding. This tool is the product of collaborative goals to: (1) create an app informed by the literature–in this case, explicit phonics instruction [13], (2) focus on an adaptation over an intervention–a supplemental tool that would assist, not replace the teacher, and allow children to bring skills from the classroom to at-home practice with reading [35–37], (3) design a fun and whimsical interface that children would want to use [25, 38], and (4) enable children to confront their challenges and build the skills to decode more complex words. Our focus on explicit phonics instruction is based on decades of research detailing its importance in literacy education [19–21, 25, 39–41], and the unique role that technology affords to provide limitless exposure and practice inside and outside the classroom. The tool was designed through collaboration between the researchers at the University of Washington (P.M.D. and J.D.Y) and Microsoft (K.L., T.M.), and then tested (independently) in a laboratory study using a pre-registered randomized control trial (RCT) design to determine whether a conceptually simple digital tool can lead to improved word reading outcomes for struggling readers. As reflected in the preregistered report (available at https://osf.io/q8tpz), the study aimed to answer the following questions: Can struggling readers use phonemic cues to improve reading fluency?, Do struggling readers benefit from phonemic cues when decoding difficult words without time constraints?, and Can we predict which individuals will benefit from the tool based on a standardized battery of reading-related assessments? We hypothesized that the phonemic cue would aid struggling readers in more accurately reading short passages, and with repeated reading, will increase their reading rate. Further, we hypothesized that the phonemic cue would aid struggling readers in more accurately decoding individual real and pseudo-words and that this benefit will relate to the amount of practice they have had with the tool. Finally, regarding individual differences, we hypothesized that those struggling readers that have a specific impairment in phonological processing would benefit the most from the support provided by this tool. Methods Pre-registration The methods, including study design, hypotheses, and analysis plan, were pre-registered using the Open Science Framework (OSF) open-access, pre-registration pipeline. We obtained initial reviews and feedback on this pre-registration from an independent OSF reviewer, revised and re-submitted our methodology, and then adhered (with some minor deviations) to this pre-registered plan throughout the duration of the study. Deviations are noted and explained within this manuscript and are compiled in a ‘Transparent Changes” document in the project repository. Documentation is available at https://osf.io/q8tpz. Participants Forty children between the ages of 8 and 12 were recruited from the University of Washington Reading & Dyslexia Research Program database, an online repository of families interested in dyslexia research in the Puget Sound region. The participants (19 females; 21 males) were classified as struggling readers based on a battery of behavioral measurements administered within one year prior to participation in the present study. Here, we use the term “struggling reader” rather than “dyslexia” because there is substantial variability in diagnostic criteria for dyslexia and our goal was to design a tool that would support literacy development for anyone that was struggling, regardless of a dyslexia diagnosis. To be considered a struggling reader, participants needed to have reading skills that were more than one standard deviation (SD) below the mean on either the Woodcock-Johnson IV Tests of Achievement Basic Reading Skills composite (WJ BRS) or the Test of Word Reading Efficiency—2 Index (TOWRE Index), and scores above 1 SD below the mean on the Wechsler Abbreviated Scale of Intelligence Full Scale-2 composite (WASI FS-2). A threshold of 1 SD, rather than the 1.5 SD threshold defined in our preregistration, was adopted to better account for the heterogeneity in the struggling reading population and to expedite participant recruitment. Further, phonological processing abilities were measured for each of the participants using the Comprehensive Test of Phonological Processing– 2 (CTOPP 2) but was not used as an enrollment criterion in the study. Together with age and IQ (WASI-II), CTOPP-2 scores were collected for the purpose of analyzing individual differences in intervention effects to determine if the tool is particularly effective for subjects with certain characteristics. Participants were previously screened for potential speech/language/hearing disorders, neurological impairments, and psychiatric disorders and had none. ADHD was not a disqualifying factor as there is a high co-occurrence with reading disability [42]. In our sample 12 children had a diagnosis of ADHD (6 Control, 6 Intervention). Demographic information on the sample can be found in Table 1. 10.1371/journal.pone.0243435.t001Table 1 Demographic information for study participants. Characteristic Intervention (N = 20) Control (N = 20) Mean (SD) Mean (SD) Age (years) 10.34 (1.3) 9.79 (1.1) Female (proportion) 0.5 0.45 WJ Basic Reading Skills Composite 78.5 (12.92) 81.15 (8.43) TOWRE-2 TWRE Index Composite 70.25 (7.48) 70.65 (7.10) WASI Full-Scale 2 Composite 97.6 (8.75) 100.15 (17.20) CTOPP Phonological Awareness Composite 87.2 (9.7) 83.4 (11.94) CTOPP Rapid Naming Composite 79.05 (8.57) 77.75 (6.7) Demographic information for participants in the Intervention and Control groups. See Methods for descriptions of the individual characteristics. For each characteristic, the mean is provided with the standard deviation within parentheses. Independent t-tests—and Wilcoxon signed-rank tests for Age/Gender—demonstrated no significant differences across all characteristics. The parents of all participants in the study provided written and informed consent under a protocol that was approved by the University of Washington Institutional Review Board and all procedures, including recruitment, child assent, and testing, were carried out under the stipulations of the University of Washington Human Subjects Division. App design Sound It Out is a web-based application (web app) that annotates text passages with visual phonemic cues to assist decoding. When a passage is viewed using the web app, the vowels appear in blue font with the image cues (located just underneath) indicating the associated phoneme. Each image cue is a highly recognizable symbol whose name contains the target vowel sound for the vowel above. For example, in the word “cow”, “o” would appear in blue, with the symbol of a house below. The house cues the child that the letter “o” in “cow” makes the same sound as the /aʊ/ in word “house”. Fig 1shows three sentences taken from a grade level passage and annotated with Sound It Out. 10.1371/journal.pone.0243435.g001Fig 1 Sound It Out example text. The vowels in this excerpt from Aesop’s ‘The Fox and the Crow” fable appear blue with phonemic image cues provided below. The name of the symbols cues the reader to the sound of the target vowel. Reprinted under a CC BY license, with permission from Microsoft Corporation, original copyright 2019. To aid in symbol recognition and retention, the app also integrates a voice cue; when a child presses the phonemic cue symbol, a voice narrates the symbol name followed by an isolated presentation of the target vowel sound. For example, in the “cow” example, when a child presses the image of the house, below “o”, a voice will say, “house, /aʊ/”. The vowel sounds were recorded by a native English speaker with training in phonetics. The recordings were judged by the three native English-speaking authors to be typical examples of the given vowel sounds and, during the training period, participants were exposed to all the vowel sounds and were able to correctly identify each vowel. The goal of this app is for the cues to provide helpful hints that aid in decoding and provide children the support they need to attempt to decode difficult words and, eventually, learn the highly inconsistent grapheme-phoneme correspondences of vowels in English. Instead of simply reading challenging words, as is typical in a speech-to-text tool [43, 44], Sound It Out focuses on a particularly difficult task for struggling readers: vowel decoding. Because vowels in English represent the majority of ‘mutually interfering discriminations’—or multiple sound associations - that young readers must master [45], learning the rules for vowels represents a significant portion of phonics curricula [3, 15, 46]. With the understanding that vowels may not be the only challenge for the reader, Sound It Out was designed to provide a tool that scaffolds learning and empowers struggling readers to read more complex passages independently. Sound It Out was designed as a feature that can be turned on or off when a child is reading. In this study, both intervention and control participants were given a tablet and taught how to use it for reading. The primary manipulation was whether the Sound It Out feature was turned on. Beyond this manipulation, the text in the table-based reader was identical. Procedure Study design In a randomized pre-post design, participants were randomized to a control or intervention condition. Randomization was unconstrained with group assignment determined at time of consent; however, sibling participants were assigned to the same group to better control participant adherence. Both intervention and control participants completed an initial, baseline session that collected all outcome measures using the normal text condition presented on a Kindle fire tablet without the Sound It Out cue (see Outcome Measures). Control participants completed a brief training in use of the tablet then a two-week, at-home practice program without the Sound It Out tool. Intervention participants completed the full training period that included both the use of the tablet as well as a formalized introduction to and practice with the Sound It Out cue prior to an identical at-home practice program with the tool turned on. After two weeks, all participants returned for a post-intervention session where the intervention participants were tested with the cue, while control participants experienced the normal text condition for all study stimuli. As the goal of this study was to test proof-of-concept for a digitally embedded phonetic cue in a small scale RCT, generalization to un-cued reading in the intervention group was not tested. Training program For the intervention group, each child was first oriented to the presence of the cue in an example passage. The researcher would show a passage with the cue and say the following: “Now we are going to use a cool tool that we made to help you read the tricky words. Underneath each word we have symbols that help you figure out the sound that the blue letters make.” The researcher would walk through a word and demonstrate how the cues could be used to help decode words. Then the researcher explained when to use the cues: “When you come to a word you don’t know, just look at the symbols and that should help you figure out the sounds that the blue letters make.” The child was then instructed to read through the passage. When they came to tricky words the researcher alternated between clicking on the symbols and naming the symbols to help sound out the words that cause difficulty. If the child read through the first example passage with ease, a more challenging passage was added to ensure that the child had demonstrated efficient and correct usage. Then, with the use of flashcards, the researcher reviewed each symbol explicitly with the child. For the control group, participants were introduced to the tablet and instructed on how to navigate to the various passages for at home practice. They were not shown the phonemic cues, but otherwise followed an identical procedure. At-home practice. After training with the application, and when the baseline testing session was completed, participants were provided tablets to take home for reading practice. Participants were asked to read at least one story per day over the course of two weeks using the app (at home). For the intervention group, the Sound It Out feature was enabled so that phonemic cues showed up in the passages. For the control group, text was rendered in the same font but without the cues. Tablets were pre-loaded with 36 first, second, and third grade supplementary passages for children to read with or without their parents. To encourage meaningful practice, parents were provided with a brief introduction to the app prior to taking home the tablet and were given instructions to allow the child to work through difficult words using the image cues prior to providing any additional guidance. For those that adhered to these instructions, and assuming each passage would take approximately ten minutes to complete, each child experienced at least 100 minutes of exposure over the two-week practice period. Adherence to the practice schedule was measured via short, three-question comprehension quizzes completed after reading a story (through a web interface). A participant was only credited for a practice passage if they received a comprehension score of at least two. Supplementary passages for at-home reading were from ReadWorks.org, an online library of grade-level passages (used with permission of ReadWorks). Passages were phonetically coded manually by the research team based on the most common pronunciation in the Pacific Northwest dialect of English. The code was verified by both P.M.D & K.L., with inconsistencies discussed and decided via consensus. All passages were displayed on an Amazon Kindle Fire 8 tablet at a set font size and resolution. Comprehension questions, to gauge practice adherence, were created by a trained, certified teacher at a local school that specializes in working with children with developmental dyslexia and were grade-level matched to the individual passages. Passages and comprehension questions can be found in the supplementary material (see S1 File). Outcome measures Measures of reading performance were collected at baseline and after the two-week period of practice. Real and pseudo word decoding accuracy was measured by having participants read lists of words that were loaded into the web app. Four unique lists of 30 real and 30 pseudo words were created with two lists being delivered at each session. All lists were developed using the orthographic wordform database MCWord [47]. Real word lists consisted of the most frequent words in the English language with five instances each of three-letter to eight-letter words. Pseudo word lists consisted of the most frequent bigrams in English with an identical progression from three-letter to eight-letter pseudo words. (See S2 and S3 Tables for detailed word statistics). All lists were unique and were rated for consistent difficulty using timed reading in ten typically reading adults. Lists were administered in a counterbalanced order for participants in each group. Lists were phonetically coded manually by the first author based on the most common pronunciation in the Pacific Northwest dialect of English. Accuracy in pronunciation was not limited to the code used for the phonemic cues but extended to acceptable pronunciations in English. At the start of administration, all participants were reminded that they were not being timed and encouraged to read as accurately as possible. For intervention participants exposed to the image cue, participants were additionally reminded that the symbols were there to help them should they come to a challenging word. Post-hoc analysis revealed that performance was highly reliable across the different word lists: performance was highly correlated for the two lists of real words (r = 0.86, p < 0.001) and pseudo words (r = 0.79, p < 0.001) in each session. Accuracy of real and pseudo word decoding was our primary outcome measure (number of words read correctly on each list akin to the Woodcock Johnson Word ID and Word Attack (both untimed measures)). Passage reading rate and accuracy was measured by having participants read grade-level passages that were loaded into the web app. All testing passages used were from the Dynamic Indicators of Basic Early Literacy Skills (DIBELS) 6th Edition library, used with permission of the University of Oregon Center on Teaching & Learning. These passages are commonly used as benchmark assessments in schools and have been extensively used in reading research. Only passages rated at second, third, and fourth grade were used. For each testing session, every passage was presented by a research assistant in the Brain Development & Education Lab and audio recorded for accurate scoring and coding. As with the decoding measures, participants were reminded at the start of administration that they were not being timed, encouraged to read as accurately as possible, and (for intervention participants) that the symbols were there should they come to a challenging word. Instead of constraining the oral reading to the one-minute limit of the DIBELS protocol, all passages were read to completion twice in a repeated reading design. Each passage reading yielded four measures: accuracy (number of words pronounced correctly) on the first and second read; and rate of reading (number of accurate words per minute) on the first and second read. To test the influence of Sound It Out on connected text reading, analyses focused on the word-reading accuracy of the first read and the word-reading rate of the second read. Second reading rate was used to control for the added time that might be associated with using the phonemic cues to sound out difficult words on the first read. Testing passages were coded and rated in the same manner as the at-home practice passages. Statistics Due to the presence of missing data, data were analyzed with linear mixed effects (LME) models, as specified in our pre-registration. Missing data consisted of word reading accuracy and rate information for twelve passages from seven participants due to testing fatigue or inability to complete the passages. For each outcome measure, we fit an LME model with fixed effects of: (1) time (pre-intervention / post-intervention as a categorical variable); (2) group (intervention / control groups as a categorical variable); (3) the group by time interaction. The models included a random intercept for participant, to account for individual variation in baseline performance. To account for differences between the individual, lab-created word lists, we added a random intercept for word list to those models. Practice data were used to ensure that all participants engaged with the tool at home and were also used in correlational analyses to examine the impact of at-home exposure on improvement. Due to issues collecting reliable usage statistics for the at home reading practice, prediction analyses were not appropriate. Instead, exploratory correlation analyses were performed using the Pearson correlation coefficient between post-pre difference scores and the three subject characteristics collected at baseline: age, WASI-II and the CTOPP-2. This analysis differs from that described in the preregistration due to the small number of reading variables collected and inability to collect robust measures of exposure, making methods of dimensionality reduction not appropriate. Analyses were carried out using the NumPy and SciPy libraries of Python and the MATLAB Statistics Toolbox (2019a) [48]. All data and analysis code associated with this manuscript are publicly accessible at the following link: [https://github.com/patdonnelly/Donnelly_2019_PLOSONE] Results Phonemic cues improve decoding accuracy for real and pseudo words For our primary outcome measure (as specified in our pre-registration [link]), children were assessed on their ability to decode lists of increasingly more complex real and pseudo words prior to, and immediately following, the two week intervention period. For this measure of decoding accuracy, words were displayed in a list (Fig 2A and 2B). Fig 2 also includes bar plots of difference scores as well as violin plots of the full score distribution. 10.1371/journal.pone.0243435.g002Fig 2 Untimed decoding performance on real and pseudo words. Example stimuli from the real-word (A) and the pseudo-word (B) lists with Sound it Out phonemic cues below the highlighted vowels. Bar plots show difference scores (number of words read correctly) from the first and second sessions for both the control and intervention groups on real word (C) and pseudo word (D) lists. Bar heights represent the additional words read on the second session. Error bars reflect +/- 1 SEM. Below the bar plots, violin density plots show group performance on these measures for both sessions with superimposed line plots of individual performance. For real-word decoding accuracy, although both groups did show some growth, the group by time interaction was not significant (β = 1.3, t(156) = 1.923, p = 0.056) indicating that the growth in the intervention group was not statistically different from the control group. For pseudo-word decoding accuracy the group by time interaction was significant (β = 3.175, t(156) = 2.99, p = 0.003) with the intervention group showing significantly greater improvement than the control group (a threshold of 0.0125 was defined in the preregistered report to adjust for multiple comparisons). At pretest, despite randomization, the intervention group by chance had lower scores than the control group: this is evidenced by the significant main effect of group in the mixed effects model (β = -6.2, t(156) = -2.65, p = 0.009). To determine the influence of practice on individual outcomes, we examined the correlation between at-home practice and individual growth. All participants completed at-home practice (intervention Mean = 13 stories, SD = 6; control Mean = 12, SD = 6), but the correlation between amount of practice and growth was not significant for the intervention group (real words r = -0.14, p = 0.55; pseudo words r = 0.23, p = 0.35) or the control group (real words r = 0.01, p = 0.97; pseudo word r = -0.37, p = 0.10). Regarding our final question on subject characteristics that predict individual response, a correlation analysis revealed no significant relationships between our baseline predictors (age, WASI-II Full Scale-2, and CTOPP-2) and improvement in real-word decoding. For pseudo-word decoding, there were negative correlations with the CTOPP-2 Phonological Awareness (PA) (r = -0.52, p = 0.018) and Phonological Memory (PM) (r = -0.48, p = 0.034) composite measures as well as the WASI-II Full Scale– 2 score (r = -0.49, p = 0.027), but these effects were not significant after correcting for multiple comparisons. Due to the heterogeneity of our sample, we tested a model with added covariates for age and initial phonological awareness ability: model fit comparison revealed no benefit to the more complex model and no significant main effects for the added covariates (S1 File). These findings show that without the constraints of time during testing, there was a beneficial effect of access to the phonemic cue for single word decoding. This benefit was observed in the case of pseudoword decoding where children were asked to pronounce novel words in isolation. Moreover, correlation analyses suggest that this effect is more pronounced for those participants with more significant impairments in phonological processing and lower IQ (though these effects did not surpass our adjusted significance threshold of p < 0.0125). Although the effect sizes were moderate (Cohen’s d = 0.74 for pseudoword decoding, d = 0.57 for real word decoding), these results suggest that children, with practice, can incorporate a novel cue to scaffold independent decoding. Phonemic cues for connected text reading To determine whether the phonemic cue confers a benefit for reading connected text, we assessed word reading accuracy and rate on grade level passages before and after intervention. Fig 3depicts bar plots of difference scores and violin density plots for the control and intervention participants in terms of (a) reading accuracy: number of words read correctly in the passage on the first read; and (b) reading rate: number of correct words per minute in the second read. 10.1371/journal.pone.0243435.g003Fig 3 Accuracy and rate for connected text reading. Bar graphs depict difference scores from the first and second sessions for both the control and intervention groups for word reading accuracy (A) and rate (B). Bar heights represent the additional number of words read correctly and additional accurate words per second on the second session. Error bars reflect +/- 1 SEM. Below the bar plots, violin density plots show group performance on these measures for both sessions with superimposed line plots of individual performance. For word reading accuracy the group by time interaction was not significant (β = 0.014, t(65) = 1.1, p = 0.275). For word reading rate there was a non-significant group by time interaction (β = 0.014, t(64) = 0.368, p = 0.714). Effect sizes were d = 0.36 for accuracy and 0.13 for rate. To examine the effect of heterogeneity in our sample, we tested a model with added covariates for age and initial phonological awareness ability: model fit comparison and analysis of added fixed effects demonstrated no significant effects of these covariates (S1 File). Correlation analyses revealed only a near-significant negative relationship between age and word reading accuracy (r = -0.55, p = 0.028) suggesting that younger children may benefit more from Sound it Out. Discussion Using a small scale RCT design, we tested the hypotheses that struggling readers could leverage a phonemic image cue placed below the vowels in digitally presented text to improve reading accuracy for isolated words and connected text, and that this benefit would be more pronounced for those readers with lower performance on measures on phonological processing. Data collected after a two-week period of unsupervised (but digitally monitored) practice demonstrated that struggling readers could read more complex words using the tool: compared to the control group, the intervention group showed a significantly larger improvement in decoding accuracy specifically for pseudo-words. As depicted in the results, this benefit did not extend to either measure of connected-text reading (accuracy and rate) and the improvement in real-word decoding did not differ significantly between intervention and control groups. Although there was no benefit, stable performance on measures of connected text reading was observed for all participants with no significant difference between groups. The lack of benefits for connected text reading might reflect the limited training period or the increased cognitive demands of a novel approach to reading. These are important questions for future studies as generalization to connected text is of key importance. Correlation analyses, after multiple comparison correction, revealed no significant relationships between our variables of interest and benefit of the cue. Due to unreliable practice data (see Statistics), analyses cannot support any conclusions regarding the relationship between subject characteristics and benefits conferred by phonemic cues. However, results suggest that the tool may benefit those participants who are younger and/or have lower phonological processing scores (see Results). Together, although most analyses failed to meet our adjusted significance threshold, data suggests that participants were able to effectively use the cues in isolated situations (i.e. pseudoword reading), but the tool did not become sufficiently automatic to produce significant gains in passage reading fluency. Decades of dyslexia research has been devoted to developing and systematically testing interventions designed for struggling readers. In the digital age, devices provide ever-expanding access to a plethora of educational apps and resources advertised as educational. Many families seek out these resources to supplement their child’s education. Unfortunately, most of this educational technology lacks scientific backing [49]. A goal of this project was to embed a core feature of evidence-based practice in literacy instruction into a digital tool to scaffold learning to read. Inspired by the key tenants of phonics instruction (e.g., explicit and clear instruction in letter-sound correspondence, repeated exposure, and systematic practice), the phonemic cue was designed to provide struggling readers with a hint to aid the decoding of novel words. We focused on vowels because, in English, the highly inconsistent grapheme-phoneme mapping is a major hurdle for struggling readers. In a landscape of digital tools that provide instant corrections at the whole word level, the phonemic cue annotation in this study is one of a only a few learning aids that provides an element of instruction (at the phoneme level) to support generalizable skill [31, 33, 43, 50]. Instead of being given the answer at the first sign of struggle, the child can utilize the phonemic cue to learn part of the word, yet still needs to exercise the building blocks of literacy to get the answer. In a similar vein to this work, researchers in the UK developed Trainertext, a scaffolding program that provides whimsical, visual mnemonics for scaffolding letter-sound correspondence. Instead of focusing on vowels, Trainertext provides a visual cue above every phoneme in connected text that is associated with short rhyme. For example, above the word “gas”, Trainertext has images related to the phrases “Goat in a Boat”, “The Ant in Pink Pants”, and “The Snake with a Shake” to represent the grapheme-phoneme correspondence of each letter. A RCT with individualized instruction and 10 months of exposure demonstrated significant improvements (Cohen’s d = ~0.80) with the largest effects seen for decoding and phonological awareness (Messer & Nash, 2018). This work demonstrated how phonics-based text annotation could be leveraged by struggling readers to bootstrap their decoding skills, and that, with extended exposure, that benefit could generalize to decoding without the cues. The present study built on this work by (a) employing a simplified symbol set focused on vowel sounds, with the hope of building a tool that would be quicker to learn and less cognitively demanding and (b) could be used immediately without requiring months of a resource-intensive intervention program. Taken together, these two studies emphasize that text annotation is a promising approach, either in combination with an in-person intervention (as in [31]), or as a tool to support at-home practice (as in the present study). As was the case with Trainertext, Sound It Out requires children to learn a new, albeit intuitively designed, symbol system and practice sufficiently for the associations to become automatic. The symbols chosen were optimized for recognizability, but the challenge remains in teaching the child to associate a portion of the symbol name with a discrete sound segment in an often-unrelated word. Moreover, the sound-symbol association must be fast enough to not impede short term memory with increased cognitive load [51–53]. These two dimensions, effective use and automaticity, are captured by the two areas of measurement: the untimed word lists and the connected text reading. The finding of improved pseudoword decoding performance indicates that without temporal constraints or the cognitive demands of connected text reading, children may be able to use the phonemic cue to improve decoding performance. This suggests that a brief, two-week practice period was enough and the tool sufficiently intuitive to have an impact. As the limited effects in passage reading accuracy and rate reveal, however, the tool did not extend to situations when time constraints were re-introduced. Either due to the limited practice period, limited supervised practice, or conflict with existing strategies children use when approaching challenging words, children did not similarly benefit from the phonemic cues in connected text. Future studies should incorporate qualitative and metacognitive methods to identify factors and circumstances that encourage struggling readers to adopt a novel strategy. Albeit promising, these results should be interpreted cautiously: Our power analysis indicated that we only had sufficient power to detect relatively large effects and many of the analyses (e.g., individual differences) were likely underpowered. Also, as there was a significant difference at pre-test for our sole finding with pseudo word decoding, future studies are needed to rule out the role of regression toward the mean and possible ceiling effects. Thus, future work is needed with larger sample sizes to provide more conclusive results. Moreover, two additional points merit further investigation. First, future experiments should more efficiently, and quantitatively and qualitatively, monitor practice adherence and cadence at home to better explore the relationship between exposure and reading-related measures. Second, given the short intervention period, we did not examine generalization to reading improvements without the cue and across different aspects of skilled reading. We only investigated whether the cue could be effectively used to decode more complex words. Thus, examining long-term learning effects and generalization to a variety of different contexts, as well as the role of parental involvement/participation is an important future direction. We had anticipated that the cue would require limited exposure, but our results are in stark contrast with previous studies that have instituted a more comprehensive, extended training program and observed significant benefit to fluent reading [31, 54]. Thus, although the training and practice periods were enough to encourage effective use, they were not enough to ensure automatic and fluent use in a natural setting, and by extension, were insufficient to make general claims on efficacy of Sound it Out for supporting long-term growth in reading skills. On the other hand, the limited effects in accuracy and rate performance suggest that either the cue did not adversely impact reading performance or that it was underutilized given the increased cognitive demands of real-time reading. Many participants in the study have received supplemental instruction previously and have learned strategies for approaching new words. As a novel strategy, the phonemic cue may have been overridden when children were asked to read continuously and for comprehension. In line with the corpus of research on strategy instruction for literacy [39, 51, 55, 56], although the children in our sample demonstrated competency in the use of the cue in isolated, single-word decoding, strategy adoption would require more sustained exposure. A strength of educational technologies for literacy is their ability to empower parents, teachers and other advocates to support and supplement their child’s learning outside the classroom [31, 57–60]. Sound it Out is unique in that it provides a tool that gives parents a strategy for reinforcing phonics principles with their child. Many parents, when confronted with the stress and challenge of raising a child who struggles with learning to read, are told to read more to their child, but not given the knowledge base needed to provide meaningful support [61]. Post-study feedback from parents in the study were overwhelmingly positive, with a majority of parents noting interest in using the tool into the future (see S1 Table). Relatedly, a study by Ronimus and colleagues demonstrated increased efficacy of GraphoGame, a digital literacy program, for reading performance when children engaged with the tool with parental involvement [57]. Thus, with parental support and potential alignment with the teacher and in-school curriculum [24, 49, 62–66], Sound it Out represents a promising venture bridging research and practice. In aggregate, these findings represent a small scale proof-of-concept for this novel approach to assisting struggling readers by merging the extensive evidence base on effective literacy instruction and the affordances that technology lends to the educational arena. Not only did it prove promising in improving decoding performance with very limited practice, but it also was observed to be non-detrimental to passage reading, meaning that it was not too cognitively demanding. Future research should focus on optimizing training and practice to produce gains that will extend beyond isolated single word decoding and lead to more confident, fluent readers. Supporting information S1 Table Parent/child responses to post-study questionnaire for the intervention group. Listed are the responses to the post-study questionnaire for the intervention group participants and their parent. After completing the study, children were asked to answer honestly to the following questions: Did you like the app? And would you like to use the app again in the future? Parents were then asked if they enjoyed using the app. Those adults who did not respond did not participate in the practice to comment on the app. (DOCX) Click here for additional data file. S2 Table Real word frequency statistics. Frequency information for real word stimuli. Words were retrieved from MCWord Orthographic Wordform Database (http://www.neuro.mcw.edu/mcword/). According to the database, the frequency is a measure of how often the wordform occurred in 1,000,000 presentations in the CELEX database. (DOCX) Click here for additional data file. S3 Table Pseudoword frequency statistics. Frequency information for pseudoword stimuli. Pseudowords were retrieved from MCWord Orthographic Wordform Database (http://www.neuro.mcw.edu/mcword/). According to the database, the constrained bigram frequency is a measure of how often the bigram wordform occurred in 1,000,000 presentations in the CELEX database. (DOCX) Click here for additional data file. S1 File Additional analyses, at-home practice passages and comprehension questions. Each passage is provided as well as the comprehensions that were used to determine passage completion. Additional analyses, with explanations, are also provided. (PDF) Click here for additional data file. We would like to thank Greg Hitchcock and both the Advanced Reading Technologies and the Learning Tools teams for their aid in the design and development of Sound It Out. This was an ideal collaboration between research and industry and would not have been possible without the Microsoft team's support and commitment to the pursuit of science. We would also like to thank Taylor Madsen and the members of the Brain Development & Education Lab at the University of Washington for their input, feedback, and support. 10.1371/journal.pone.0243435.r001 Decision Letter 0 Männel Claudia Academic Editor © 2020 Claudia Männel2020Claudia MännelThis is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.Submission Version0 26 May 2020 PONE-D-19-34461 Annotating digital text with phonemic cues to support decoding in struggling readers PLOS ONE Dear Dr. Donnelly, Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. As you can see from the detailed comments of both reviewers (which I approve from my reading of the manuscript), they above all request a more comprehensive review of existing studies on embedded digital support tools and a clearer motivation of changes from the preregistration and particular methodological decisions.    Please submit your revised manuscript by Jul 10 2020 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file. Please include the following items when submitting your revised manuscript: A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'. An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'. If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter. If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: http://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols We look forward to receiving your revised manuscript. Kind regards, Claudia Männel, PhD Academic Editor PLOS ONE Journal Requirements: When submitting your revision, we need you to address these additional requirements. 1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf 2.  Thank you for stating the following in the Financial Disclosure section: "This work was funded by NSF BCS 1551330, NICHD R21HD092771, Microsoft Research Grants and Jacobs Foundation Research Fellowship to J.D.Y. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript" We note that one or more of the authors are employed by a commercial company: "Microsoft Corporation" a) Please provide an amended Funding Statement declaring this commercial affiliation, as well as a statement regarding the Role of Funders in your study. If the funding organization did not play a role in the study design, data collection and analysis, decision to publish, or preparation of the manuscript and only provided financial support in the form of authors' salaries and/or research materials, please review your statements relating to the author contributions, and ensure you have specifically and accurately indicated the role(s) that these authors had in your study. You can update author roles in the Author Contributions section of the online submission form. Please also include the following statement within your amended Funding Statement. “The funder provided support in the form of salaries for authors [insert relevant initials], but did not have any additional role in the study design, data collection and analysis, decision to publish, or preparation of the manuscript. The specific roles of these authors are articulated in the ‘author contributions’ section.” If your commercial affiliation did play a role in your study, please state and explain this role within your updated Funding Statement. b) Please also provide an updated Competing Interests Statement declaring this commercial affiliation along with any other relevant declarations relating to employment, consultancy, patents, products in development, or marketed products, etc.  Within your Competing Interests Statement, please confirm that this commercial affiliation does not alter your adherence to all PLOS ONE policies on sharing data and materials by including the following statement: "This does not alter our adherence to  PLOS ONE policies on sharing data and materials.” (as detailed online in our guide for authors http://journals.plos.org/plosone/s/competing-interests) . If this adherence statement is not accurate and  there are restrictions on sharing of data and/or materials, please state these. Please note that we cannot proceed with consideration of your article until this information has been declared. Please include both an updated Funding Statement and Competing Interests Statement in your cover letter. We will change the online submission form on your behalf. Please know it is PLOS ONE policy for corresponding authors to declare, on behalf of all authors, all potential competing interests for the purposes of transparency. PLOS defines a competing interest as anything that interferes with, or could reasonably be perceived as interfering with, the full and objective presentation, peer review, editorial decision-making, or publication of research or non-research articles submitted to one of the journals. Competing interests can be financial or non-financial, professional, or personal. Competing interests can arise in relationship to an organization or another person. Please follow this link to our website for more details on competing interests: http://journals.plos.org/plosone/s/competing-interests 3. Please include captions for your Supporting Information files at the end of your manuscript, and update any in-text citations to match accordingly. Please see our Supporting Information guidelines for more information: http://journals.plos.org/plosone/s/supporting-information. [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions Comments to the Author 1. Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. Reviewer #1: Partly Reviewer #2: No ********** 2. Has the statistical analysis been performed appropriately and rigorously? Reviewer #1: I Don't Know Reviewer #2: No ********** 3. Have the authors made all data underlying the findings in their manuscript fully available? The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified. Reviewer #1: Yes Reviewer #2: Yes ********** 4. Is the manuscript presented in an intelligible fashion and written in standard English? PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here. Reviewer #1: Yes Reviewer #2: Yes ********** 5. Review Comments to the Author Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters) Reviewer #1: There is very limited literature on Embedded Support Features, which can provide critical supports to struggling readers for understanding new content and participating in learning activities. Therefore, this study, which explores the efficacy of a tablet-based text annotation reading tool in intervention research, to help children decode novel words, fills a gap in the literature on embedded digital supports. The study makes a convincing case that while the edtech market has grown tremendously as new apps are being released every year and federal funds are being devoted to provide infrastructure for even greater involvement with digital media in schools, many of the edtech tools “lacked scientific backing”. While the study rightfully argues that edtech tools should go beyond just being entertaining and should be backed by research-based best practices, it fails to provide some literature on the best-practices. The current study lacks important and needed detail regarding how the existing literature would be evaluated to arrive at meaningful conclusions. For example, the study should refer to the extensive scientific literature on what works for struggling readers in terms of best practices as well as existing studies on embedded digital support tools. A review of the existing literature on digital supports would have provided a better understanding of what is known--which digital features or supports have been found to have promise to support struggling students’ learning and under what conditions. The study made some references to such tools/studies in the Discussion section, which should have been presented in the Literature Review section. I agree with the authors who argue that “there is great need for researchers studying literacy development and reading disabilities to work with tech developers on the design of tools that are grounded in the extensive scientific on what works for struggling readers, systematically test their effectiveness, and contribute to the development of standards of practice for educational apps targeted at literacy”. However, even though we have “evidence-based practices (e.g., phonemic awareness, phonics)”, we cannot assume that the best practices identified in non-digital contexts may easily apply or be expanded within digital learning contexts. The study makes such an assumption. The study states that “Intervention participants completed an extended training period” but it is not clear what “extended” means. One of the weaknesses of the study is to conclude that 100 minutes of exposure over the two-week practice period is enough to make conclusions about the efficacy of a tablet-based text annotation reading tool. The result could simply be due to the novelty effects, which pose a threat to external validity because they make it difficult to know if the results of the study are due to a treatment that works or due to the novelty of a treatment. Reviewer #2: The authors describe a small-scale randomized control trial with 40 struggling readers aged 8-12 who are split across two intervention groups. The aim was to investigate the effect of annotating text with phonemic cues in a tablet app by allowing children to play for a two-week intervention phase. Their results indicate that there is a benefit on untimed pseudoword decoding accuracy, but no effects on word decoding accuracy or connected text reading accuracy and speed. The manuscript is well written, the experimental design was preregistered and all data was made available. There are some shortcomings in this study, some of which relate to the experimental setup itself, and others to unjustified deviations from the study preregistration. Major: 1) The manuscript itself does not contain any research questions and hypothesis and only refers to the preregistration on OSF. This should be changed and implemented into the manuscript. Furthermore, the preregistration contains an additional question relating to individual phonological skills which is not touched upon in the manuscript. 2) As far as I understand from the manuscript the word lists and reading passages at post-training assessment were carried out with the phonemic cues present as depicted in Figure 2. It remains unclear whether this feature was only activated for the intervention group (1) or both groups (2). For 1) this would mean that what is being measured is the effect of cueing but not the generalization of the training on un-cued reading. Option 2) would mean that control children are confronted with this annotation for the first time at an assessment moment, which would increase their cognitive load and put them at a disadvantage. In either case, the impact of training should be assessed without cueing in a real-world reading scenario. 3) P13:L21: "a threshold of 0.025 was defined in the preregistered report to adjust for multiple comparisons". The report contains a threshold of 0.0125 which would render the sole significant finding not significant. Where does this deviation come from? Minor: Introduction: Adding phonemic cues in form of symbols above text is not necessarily bound to digital media per se and could equally be done on paper. It remains unclear why there is a focus on implementing this inside an app. How does this relate to previous literacy intervention work? Has this been done on paper and were there comparable findings? Method: P6:L17: The preregistration report mentions 1.5SD below the mean as an inclusion criterion for being a struggling reader, here it is 1SD. Where does this deviation come from? The group characteristics provided in Table 1 do not contain p-values to show that the groups do not differ from one another. Could you also explain the randomization/matching used for group assignment? Were there any constraints so that the groups would end up being comparable? In the discussion you mention that many participants were receiving supplemental instruction for their reading difficulties. Can you specify how many and how their distribution among the two groups were? Potential gains may ultimately also stem from simultaneous speech therapy if the groups are not balanced. Given that your sample is quite heterogenous with an age range of 8-12 years and the intention to include a broad range of struggling readers it might be worthwhile to add additional co-variates to your models to support a more causal and generalizable interpretation of your intervention (e.g. for age or pre-test phonological awareness skills - which was also an unmentioned research question). Could you provide additional information on the custom word lists you created? E.g. frequency information for the words and neighborhood density for the pseudowords. Furthermore, reliability information would also be relevant, i.e. Cronbach's alpha, split-half reliability or correlation between lists. Matlab script: Is there a specific reason why the models where fit with method 'ML'? This is only useful for stepwise model comparison and final models for publication should be fitted with REML as this gives a less biased estimation and more robustness towards outliers. Results: You mention missing data. Could you specify which variables from how many subjects are affected by this? I much appreciate access to the raw data and code! It would be nice if you added a code book indicating what measure each variable contains. P14:L3: Looking at the numbers of played stories (mean=13/12, SD=6) it seems that intervention fidelity was not very high. Could you provide more information on that? How many minutes of training would this translate to? What was the range of minimum and maximum exposure? Did children adhered to the instruction of playing one story per day? What happened to children that played less? Did they start playing each day and then stop after a few days? Were there children that did not play for a week and then played 10 session on the last day? P14:L3: Instead of looking into correlations I would suggest including app exposure as a covariate into your models. P14:L10: Thank you for providing an effect size. It would be informative if you could also provide these for all other discussed effects. The preregistration report states: "Based on pilot data it is reasonable to expect that the intervention will help subjects read 5-10 additional words on the untimed word lists." as well as "subjects differ, on average, by +/- 3 words on repeated administrations of the word lists". Based on these numbers a statistical power of 0.9 was estimated. In the actual study the results are much lower: 1 additional word and 3 additional pseudowords, which falls into the range of the expected test-retest reliability you mentioned. Which makes me wonder how robust is this effect? The actual power will be much lower and the reported effect size of d=0.74 seems way too large for such a short intervention period. I would be very careful at overinterpreting these findings as they are probably not generalizable. Discussion: Given the comments above, possible effects should be discussed with more caution. It also remains unclear what was actually measured with the experimental design (see major point 2 above). This should be made clearer when discussing potential effects. The lack of research questions in the manuscript makes the discussion rather unspecific. Here you should get back to your initial questions and hypothesis. P15:L11: "In connected text reading, growth was observed for all participants but without a significant difference between groups." These effects are essentially zero (0.01 - 0.04 words per minute extra), so they don’t have a practical relevance. P16:L16: There are many different effect size measures which differ in their scales. Please mention which one is meant. I'm missing a limitations section where you discuss possible lack of power, adherence to daily app use and the need for replication in a bigger sample. Summary: In sum, I agree that text annotation is a promising approach, but it seems that this study describes null results. Those are still relevant to make publicly available and I would recommend to further clarify the experimental design and conduct an exploratory analysis by adding relevant co-variates to the models (pre-test phonological awareness skills, age, app exposure). ********** 6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy. Reviewer #1: No Reviewer #2: Yes: Toivo Glatz [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.] While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step. 10.1371/journal.pone.0243435.r002 Author response to Decision Letter 0 Submission Version1 9 Jul 2020 A detailed response to reviewers has been uploaded with this submission. We humbly thank the reviewers for their thoughtful comments and critiques that we believe have led to a stronger manuscript overall. Attachment Submitted filename: Response to Reviewers.docx Click here for additional data file. 10.1371/journal.pone.0243435.r003 Decision Letter 1 Männel Claudia Academic Editor © 2020 Claudia Männel2020Claudia MännelThis is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.Submission Version1 18 Aug 2020 PONE-D-19-34461R1 Annotating digital text with phonemic cues to support decoding in struggling readers PLOS ONE Dear Dr. Donnelly, Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. Please submit your revised manuscript by Oct 02 2020 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file. Please include the following items when submitting your revised manuscript: A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'. An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'. If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter. If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: http://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols We look forward to receiving your revised manuscript. Kind regards, Claudia Männel, PhD Academic Editor PLOS ONE [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions Comments to the Author 1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation. Reviewer #1: (No Response) Reviewer #2: (No Response) ********** 2. Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. Reviewer #1: (No Response) Reviewer #2: Partly ********** 3. Has the statistical analysis been performed appropriately and rigorously? Reviewer #1: (No Response) Reviewer #2: No ********** 4. Have the authors made all data underlying the findings in their manuscript fully available? The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified. Reviewer #1: Yes Reviewer #2: Yes ********** 5. Is the manuscript presented in an intelligible fashion and written in standard English? PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here. Reviewer #1: Yes Reviewer #2: Yes ********** 6. Review Comments to the Author Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters) Reviewer #1: I would like to thank the authors for the changes they have made in this version. The authors made an effort to address my comments in the previous round of review. For example, I commented that the study should refer to the extensive scientific literature on what works for struggling readers in terms of best practices as well as existing studies on embedded digital support tools. While this version provides some literature review of existing studies on embedded digital supports, it still doesn't provide any literature review on research-based best practices that the authors refer to such as (e.g., phonemic awareness, phonics etc.). Additionally, in this version, they referred to 3 meta-analysis (9, 16, 17) that demonstrate promise for digital solutions in the context of literacy. However, all these 3 studies were conducted pretty much by the same authors. It looks as if the authors published the same or similar meta-analyses in three different publications. If that is the case, I would advise referring to only one of the meta-analysis rather than all three and cite some other existing studies on embedded digital support tools. Rather than explaining each study, synthesize results into a summary of what is and is not known, identifying areas of controversy in the literature. Reviewer #2: Review for PONE-D-19-34461R1: "Annotating digital text with phonemic cues to support decoding in struggling readers" I would like to thank the authors for transparently addressing the raised points and providing a much-improved manuscript. Below is a list of mostly minor points which still require further attention or clarification. My only major concern relates to the data analysis and provided code. Introduction: Regarding research questions (P6:L14-17) and hypotheses (P6:L17-23) none of the research questions contains exposure and you added the hypothesis of an intervention x exposure interaction to the second research question, although it would also be valid for the first one, and ultimately perhaps be a better fit with the third one. Is there a specific reason for that? Method: Following up on your clarification of my mayor point 2 in the previous version of the manuscript, you added that at post-intervention only intervention participants saw the cues and control participants had cues disabled. Could you further clarify how the assessment looked at pre-test? Did both groups read without cues? Or was it already enabled for the intervention group? And did the assessment take place before or after the introduction training with the app, or was the pre-assessment done offline in a pencil and paper style? Considering that grapheme-phoneme associations also depends on availability of high-quality phoneme representations, could you clarify whether the recordings you added to the app were judged for prototypicality, and whether children were instructed to play with headphones or not? P7:L16: I assume classification as struggling reader happened within one year "prior" to participation? P8:L11: If ADHD was not an exclusion criterion could you provide numbers on how many children had this diagnosis? P11:L22: The instruction "When you come a word you don't know..." does not appear to be grammatical? P13:L1-3: Given that you do not have logs for exposure and credit exposure based on the comprehension questions I wanted to point out that the comprehension questions are all True/False responses and therefore the probability of getting 2 or more correct replies by just guessing (and therefore being credited with the exposure) is rather high at 0.5. Did you consider to analyze these comprehension data for intervention effects as well? P14:L7: Please avoid exponential notation for p-values and shorten to p < 0.001 according to common convention. P14:L9: Could you clarify whether the Woodcock Johnson test measures timed or untimed decoding? Data analysis: P8:L13: In the caption of Table 1 you say that you conducted independent t-tests. However, not all variables are normally distributed, and you must use the wilcox/mann-whitney test in these cases. It does not influence the fact that there are not differences between the groups though. Thank you for carrying out additional analyses with age and initial phonological awareness as covariates. Could you still add these models to the supplementary information and add a sentence in the results section indicating that this was done and yielded comparable results? P15:L6-L15: Thank you for providing the code book which allowed me to have a look at the data myself. There are still mayor issues with the data analysis though. In the manuscript you state that you added independent random intercepts of time and participant - which would suggest (1|time) and (1|participant), in the Matlab code you have session nested within participants (1-session|participant) as well as a reading list random intercept (1|acc_Indicator). The former is not well specified though, because "1-" ignores the session part. It should be specified as (1+session|participant) or because the 1 is implied (session|participant). Was there a specific intention behind using "1-"? In any case, this changes the coefficients only minimally but changes the confidence intervals quite a bit. I'm also concerned that this is a very complex random effects structure for the little data you have and you might be overfitting. I would suggest to do model comparison (with AIC/BIC) to see which random effects structure is required. Did you also check that model assumptions were fulfilled after fitting (normality and homoscedasticity of residuals?) Furthermore, the statistics which are presented in the results section do only partially stem from the provided Matlab code. It is, for example, not apparent where the main effects of group and session (P16:L21-L22) come from and why they only have 78 degrees of freedom when these effects have 156 DF in the model. Please make sure to provide all code for results which you provide in the manuscript (also for possible post-hoc or correlation analyses as well as calculations of effect sizes). With the pseudoword model (P17:L3-L8) there is also the issue that it describes a big and significant difference between the two groups at pretest (with the intervention group scoring 6 words lower) while at post-test the two groups are at the same level again. This is not mentioned in the results section and sheds a different light on the sole significant effect. This is made worse by plotting only pre-post differences and the use of bar charts to represent the data as it distorts the perception of observed values and draws attention to unimportant aspects (i.e. bar height rather than difference between means). See, e.g.: https://doi.org/10.1371/journal.pbio.1002128 Consider using dotplots or boxplots of the raw pre-post data with an indicator of the means. P16:L17: I understand that due to unreliable usage statistics the correlation analysis was the best you can do, but this is not a valid approach for a prediction analysis. It's best to be transparent about the unreliable data by adding a few sentences from the response letter to the statistics part of the methods section and also bring this up in the limitations. Or remove this analysis altogether after mentioning that this part of the data collection did not work as intended. Results: P16:L21: You are using ß (latin sharp s) instead of β (greek beta), as well as a p-value with 5 decimals. P18:L1-3: Repeated use of "pronounced". Furthermore, the effect was not "particularly pronounced" for the pseudoword decoding, but it was exclusively there. P18:L3-5: Please present the statistics if you refer to effects which suggest something. Also in the Matlab script. P19:L3: Significance and effect sizes are completely unrelated. We can observe giant yet non-significant effects, as well as extremely small (and thus irrelevant) yet highly significant effects. Discussion: P19:L22-P20:L4: I find this part a bit misleading as it jumps back and forth: things are inconclusive; a relation between pretest phonological awareness and intervention outcome is mentioned for which no statistic is provided in the results; two mentions that after p-value adjustment there were no effects left, but it is still discussed that these effects might suggest something. Especially at the start of the discussion there should be a clear summary of what was found first and then one can discuss what the presence and absence of effects might indicate. P20:L16: [...] highly inconsistent grapheme-phoneme mapping "of English" [...] P22:L3-L5: Throughout the manuscript there are sentences which use many commas and therefore appear encapsulated and make it unnecessarily hard for the reader to grasp the most relevant information. In this case you can straight out write that there is an improved decoding of pseudowords and remove the subordinate clause. As this is the only comment regarding writing style, I also wanted to note that there appear to be a lot of double spaces in manuscript. Content wise it could be highlighted here that this improvement appeared after being trained with the app for (only) 2 weeks. P23:L5-L13: This is an interesting observation. They appeared to have learned to make use of the cues, but this disappears in the testing situation. I feel it’s quite relevant to understand this better. Given that there is no performance penalty it seems like they might be able to ignore the cues altogether? One could take a more qualitative approach and investigate how children use it in these different scenarios or alternatively add an assessment mode into the app itself which blends with the intervention? P23:L23: Given that parents also got a brief introduction to the app and you come back to it in the discussion it might be worth to recommend measuring parental involvement in future studies? Summary: In summary, the manuscript has much improved and will be a relevant addition to the field, but the provided analysis and code are not yet up to the required standards. ********** 7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy. Reviewer #1: No Reviewer #2: Yes: Toivo Glatz [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.] While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step. 10.1371/journal.pone.0243435.r004 Author response to Decision Letter 1 Submission Version2 25 Sep 2020 We would like to thank the reviewers once again for their thoughtful and insightful comments on this study. We have made further revisions and agree with the reviewers’ points that these changes will make this manuscript suitable for publishing. The reviewers comments are below in black and our responses and specific actions are noted in blue. Reviewer #1: I would like to thank the authors for the changes they have made in this version. The authors made an effort to address my comments in the previous round of review. For example, I commented that the study should refer to the extensive scientific literature on what works for struggling readers in terms of best practices as well as existing studies on embedded digital support tools. While this version provides some literature review of existing studies on embedded digital supports, it still doesn't provide any literature review on research-based best practices that the authors refer to such as (e.g., phonemic awareness, phonics etc.). We thank the reviewer for this added clarification on the missing references in our literature review section. We have added necessary citations to the statement above “ (e.g., phonemic awareness, phonics etc.)” on page 4. Beyond explanation of our phonemic cue in the context of other embedded supports, we believe a more thorough review of specific instructional techniques for phonics/phonemic awareness is beyond the scope of this manuscript. In this revision we do, however, refer the reader to the most comprehensive review papers and meta-analyses. Additionally, explanation of the focus of the app is provided in the App design section (page 9-10) detailing its use in the context of broader phonics curricula. The section of the introduction now reads: “Decades of scientific research into the behavioral and neural mechanisms of literacy learning has led to the development and testing of effective intervention programs for struggling readers, and established comprehensive guidelines and best-practices for implementation of an effective curriculum [2,13–15]. Unfortunately, these evidence-based practices (e.g., phonemic awareness, phonics) are largely not being incorporated into the current technological boom. For example, a consistent finding in the intervention literature is that children with dyslexia benefit from direct instruction in phonological awareness and curricula that make clear links between orthography and phonology whereas children with stronger reading skills can often infer grapheme-phoneme correspondences without direct instruction (for review of the extensive literature on the importance of phonics/phonemic awareness see [16–23]). ” (P4, L8-17) Additionally, in this version, they referred to 3 meta-analysis (9, 16, 17) that demonstrate promise for digital solutions in the context of literacy. However, all these 3 studies were conducted pretty much by the same authors. It looks as if the authors published the same or similar meta-analyses in three different publications. If that is the case, I would advise referring to only one of the meta-analysis rather than all three and cite some other existing studies on embedded digital support tools. Rather than explaining each study, synthesize results into a summary of what is and is not known, identifying areas of controversy in the literature. Although the meta analyses cited have very similar authorship, they provide slightly different perspectives on the literature that we believe provide a better picture of the state of research in this area. However, we do agree that the first (Cheung & Slavin, 2011) substantially overlaps with the Slavin, et al., 2011 and have removed the first from the citation. The primary difference in contribution between the Slavin, et al., 2011 and the Cheung & Slavin, 2013 is that the latter provides a specific context for struggling readers whereas the former applies to all elementary school youth. We believe that both perspectives are highly relevant and have amended the section to the following with some added synthesis: “Recent metanalyses demonstrate much promise for digital solutions in the context of literacy, yet also describe the multitude of ways that technology is an inappropriate substitute for many aspects of pedagogy [24,25]. Namely, these meta analyses demonstrate that technologies focused on supplementing what is provided in the evidence-based classroom (i.e. explicit phonics), rather than restructuring at the classroom level, have demonstrated the most promise in the digital landscape. These findings, however, should be interpreted with caution as the authors further contend that the preponderance of studies in this area are characterized by small samples and poor study design [25]. ” (P5, L3-10) Reviewer #2: Review for PONE-D-19-34461R1: "Annotating digital text with phonemic cues to support decoding in struggling readers" I would like to thank the authors for transparently addressing the raised points and providing a much-improved manuscript. Below is a list of mostly minor points which still require further attention or clarification. My only major concern relates to the data analysis and provided code. Thank you Introduction: Regarding research questions (P6:L14-17) and hypotheses (P6:L17-23) none of the research questions contains exposure and you added the hypothesis of an intervention x exposure interaction to the second research question, although it would also be valid for the first one, and ultimately perhaps be a better fit with the third one. Is there a specific reason for that? We thank the reviewer for identifying this point. We indicated an exposure effect in the second question only because we believed that children would require significantly more practice than the 2-week period to motivate effective use during connected text reading (question 1). As the second question related to single word decoding in a more “assessment” context we believed we would see more active use and, in extension, more effect of exposure. Method: Following up on your clarification of my mayor point 2 in the previous version of the manuscript, you added that at post-intervention only intervention participants saw the cues and control participants had cues disabled. Could you further clarify how the assessment looked at pre-test? Did both groups read without cues? Or was it already enabled for the intervention group? And did the assessment take place before or after the introduction training with the app, or was the pre-assessment done offline in a pencil and paper style? We apologize for the confusion that was met in parsing our study design. The first sentence of the study design section indicates that “all participants completed an initial, baseline session that collected all outcome measures using a normal text condition presented on a Kindle fire tablet without the Sound It Out cue” (L6-8). The training period for both groups (intervention and control) occurred after this baseline testing with the normal text condition as detailed in the remaining paragraph. All testing was completed on a Kindle Fire tablet. We have further adjusted the text to clarify this point. Considering that grapheme-phoneme associations also depends on availability of high-quality phoneme representations, could you clarify whether the recordings you added to the app were judged for prototypicality, and whether children were instructed to play with headphones or not? The recordings were created by a member of the study team (KL) and were judged to be typical examples of the vowel sounds by PD & JD. All three are native English speakers. Furthermore, the vowel sounds were played during the training period and kids were given the opportunity to indicate if a given recording was unclear. The following has been added to the manuscript for clarity on this point: “The vowel sounds were recorded by a native English speaker with training in phonetics. The recordings were judged by the three native English-speaking authors to be typical examples of the given vowel sounds and, during the training period, participants were exposed to all the vowel sounds and were able to correctly identify the vowels.” (P10, L20-P11,L2) P7:L16: I assume classification as struggling reader happened within one year "prior" to participation? Yes - the statement has been amended to the following: “The participants (19 females; 21 males) were classified as struggling readers based on a battery of behavioral measurements administered within one year prior to participation in the present study.” (P8, L4-6) P8:L11: If ADHD was not an exclusion criterion could you provide numbers on how many children had this diagnosis? Yes - the following statement has been added: “In our sample 12 children had a diagnosis of ADHD (6 Control, 6 Intervention).” (P9, L1-2) P11:L22: The instruction "When you come a word you don't know..." does not appear to be grammatical? Thank you for catching this typo - we have amended the statement to the following: “When you come to a word you don’t know, just look at the symbols and that should help you figure out the sounds that the blue letters make.” (P12, L19-21) P13:L1-3: Given that you do not have logs for exposure and credit exposure based on the comprehension questions I wanted to point out that the comprehension questions are all True/False responses and therefore the probability of getting 2 or more correct replies by just guessing (and therefore being credited with the exposure) is rather high at 0.5. Did you consider to analyze these comprehension data for intervention effects as well? We did look to see if there were any differences in practice between control and intervention participants, and found that there was no difference between the distributions of practice using an independent t test (t = -0.592, p=0.558). P14:L7: Please avoid exponential notation for p-values and shorten to p < 0.001 according to common convention. We have amended where applicable. P14:L9: Could you clarify whether the Woodcock Johnson test measures timed or untimed decoding? We have added clarification - the statement now reads: “Accuracy of real and pseudo word decoding was our primary outcome measure (number of words read correctly on each list akin to the Woodcock Johnson Word ID and Word Attack (both untimed measures)).” (P15, L4-6) Data analysis: P8:L13: In the caption of Table 1 you say that you conducted independent t-tests. However, not all variables are normally distributed, and you must use the wilcox/mann-whitney test in these cases. It does not influence the fact that there are not differences between the groups though. Thank you for pointing this out - I have performed the Wilcoxon signed-rank test on Age/Gender to conclude that all characteristics demonstrate no significant differences. The caption now reads: “Demographic information for participants in the Intervention and Control groups. See Methods for descriptions of the individual characteristics. For each characteristic, the mean is provided with the standard deviation within parentheses. Independent t-tests - and Wilcoxon signed-rank tests for Age/Gender - performed demonstrate no significant differences across all characteristics.” Thank you for carrying out additional analyses with age and initial phonological awareness as covariates. Could you still add these models to the supplementary information and add a sentence in the results section indicating that this was done and yielded comparable results? We have added the following to the manuscript to indicate that this analysis was performed with a parenthetical reference to a more in-depth explanation in the supplement: “Due to the heterogeneity of our sample, we tested a model with added covariates for age and initial phonological awareness ability: model fit comparison revealed no benefit to the more complex model and no significant main effects for the added covariates (S1 File).” (P18, L23 - P19, L3) “Similarly concerned with the heterogeneity of our sample, we tested a model with added covariates for age and initial phonological awareness ability: model fit comparison and analysis of added fixed effects demonstrated no significant effect (S1 File).” (P20, L10-13) The following has been added to the Supplement: “Role of Age and Phonological Awareness as Covariates to Mixed Effects Models Due to the heterogeneity of our struggling reader population in both reading ability and age, we performed an exploratory analysis of these results testing a model that adds covariates for participant age and initial phonological awareness ability (as measured using the CTOPP phonological awareness composite measure). Model fits were compared using AIC/BIC values and revealed that in no case was the model with added covariates a superior model to the simpler model used in the manuscript. Looking at the results of these models, moreover, revealed no significant effects of age or phonological awareness. Interaction effects were unchanged from the simpler models reported in the manuscript. The model fit workflow and code associated with this analysis can be found in the associated project repository on GitHub.” P15:L6-L15: Thank you for providing the code book which allowed me to have a look at the data myself. There are still mayor issues with the data analysis though. In the manuscript you state that you added independent random intercepts of time and participant - which would suggest (1|time) and (1|participant), in the Matlab code you have session nested within participants (1-session|participant) as well as a reading list random intercept (1|acc_Indicator). The former is not well specified though, because "1-" ignores the session part. It should be specified as (1+session|participant) or because the 1 is implied (session|participant). Was there a specific intention behind using "1-"? In any case, this changes the coefficients only minimally but changes the confidence intervals quite a bit. I'm also concerned that this is a very complex random effects structure for the little data you have and you might be overfitting. I would suggest to do model comparison (with AIC/BIC) to see which random effects structure is required. Did you also check that model assumptions were fulfilled after fitting (normality and homoscedasticity of residuals?) We thank the reviewer for drawing attention to this issue in our statistical model. We had in fact run model comparisons to determine the ideal random effects structure. The notation issue raised by the reviewer was due to incorrect documentation in a previous version of MATLAB. The best fitting model was, in fact, the simple one suggested by the reviewer with a random intercept for participant. We now provide a more detailed model fit workflow. For the word list data, we also added a random intercept for list number (1|acc_Indicator) because we found it to vastly improve model fit due to slight variations between word lists (though including or excluding this random effect does not impact the effects of interest). We have repeated all analyses using this random effects structure with no major changes to our main findings. We also add to our uploaded code checks for normality in residual distribution/heteroscedasticity. Analyses reveal that in all model fits there is a normal distribution of residuals and no evidence of heteroscedasticity. The following has been amended to the manuscript: “For each outcome measure, we fit an LME model with fixed effects of: (1) time (pre-intervention / post-intervention as a categorical variable); (2) group (intervention / control groups as a categorical variable); (3) the group by time interaction. The models included a random effect for participant, to account for individual variation in baseline performance. To account for differences between the individual, lab-created word lists, we added a random effect for word list to those models. ” (P16, L6-11) We have also amended the “Transparent Changes” document in our preregistration. Furthermore, the statistics which are presented in the results section do only partially stem from the provided Matlab code. It is, for example, not apparent where the main effects of group and session (P16:L21-L22) come from and why they only have 78 degrees of freedom when these effects have 156 DF in the model. Please make sure to provide all code for results which you provide in the manuscript (also for possible post-hoc or correlation analyses as well as calculations of effect sizes). The mentioned analyses were performed post-hoc, and for clarity have been changed to paired t-tests on subsets of the dataset to help illustrate the differences between groups and provide perspective for the interaction results. Discussion of these results in the manuscript have been amended to clarify that distinction: “For real-word decoding accuracy the group by time interaction was not significant (β=1.3, t(156) =1.923, p=0.056) indicating that the growth in the intervention group was not statistically different from the control group. Post-hoc, paired t-test analyses dividing the data at the group level revealed a significant increase in in the intervention group (t(19) = 3.75, p = 0.001) and a non-significant increase in the control group (t(19) = 1.10, p = 0.285).” “For pseudo-word decoding accuracy the group by time interaction was significant (β=3.175, t(156)=2.99, p=0.003) with the intervention group showing significantly greater improvement than the control group (a threshold of 0.0125 was defined in the preregistered report to adjust for multiple comparisons). At pretest, despite randomization, the intervention group by chance had lower scores than the control group: this is evidenced by the significant main effect of group in the mixed effects model. Post-hoc, paired t-test analyses at the group level revealed there was a significant increase in the intervention group (t(19) = 2.176, p = 0.042) and a small but non-significant decrease in the control group (t(19) = -1.149, p = 0.265).” “For word reading accuracy the group by time interaction was not significant (β=0.014, t(65)=1.1, p=0.275). Post-hoc, paired t-test analyses dividing the data at the group level revealed a non-significant increase in the intervention group (t(15) = 2.325, p = 0.035) and a non-significant increase in the control group (t(17)=0.518, p = 0.611). [JY1] For word reading rate there was a non-significant group by time interaction (β=0.014, t(64)=0.368, p=0.714). Post-hoc, paired t-test analyses at the group level revealed a non-significant increase in both the intervention (t(14)=1.468, p = 0.164) and control groups (t(17)=1.000, p = 0.331).” We have also ensured that these and other analyses are available in the public repository. With the pseudoword model (P17:L3-L8) there is also the issue that it describes a big and significant difference between the two groups at pretest (with the intervention group scoring 6 words lower) while at post-test the two groups are at the same level again. This is not mentioned in the results section and sheds a different light on the sole significant effect. This is made worse by plotting only pre-post differences and the use of bar charts to represent the data as it distorts the perception of observed values and draws attention to unimportant aspects (i.e. bar height rather than difference between means). See, e.g.: https://doi.org/10.1371/journal.pbio.1002128 Consider using dotplots or boxplots of the raw pre-post data with an indicator of the means. This is a very important point and we have added the following to the results section to add clarity and transparency: “For pseudo-word decoding accuracy the group by time interaction was significant (β=3.175, t(156)=2.99, p=0.003) with the intervention group showing significantly greater improvement than the control group (a threshold of 0.0125 was defined in the preregistered report to adjust for multiple comparisons). At pretest, despite randomization, the intervention group by chance had lower scores than the control group: this is evidenced by the significant main effect of group in the mixed effects model. Post-hoc, paired t-test analyses at the group level revealed there was a significant increase in the intervention group (t(19) = 2.176, p = 0.042) and a small but non-significant decrease in the control group (t(19) = -1.149, p = 0.265).” (P18, L3-10) We have also added violin plots to the Supplemental information (S1 Fig) to provide more detailed visualizations of the data. P16:L17: I understand that due to unreliable usage statistics the correlation analysis was the best you can do, but this is not a valid approach for a prediction analysis. It's best to be transparent about the unreliable data by adding a few sentences from the response letter to the statistics part of the methods section and also bring this up in the limitations. Or remove this analysis altogether after mentioning that this part of the data collection did not work as intended. Thank you for this insight. We agree that it is important to be more transparent about this limitation and have amended the statistics section to read: “Due to issues collecting reliable usage statistics for the at home reading practice, prediction analyses were not appropriate. Instead, post-hoc correlation analyses were performed using the Pearson correlation coefficient between post-pre difference scores and the three subject characteristics collected at baseline: age, WASI-II and the CTOPP-2. This analysis differs from that described in the preregistration due to the small number of reading variables collected and inability to collect robust measures of exposure, making methods of dimensionality reduction not appropriate.” (P16, L14-20) We have also amended the Discussion section to specify Correlation rather than Prediction analysis and draw the reader’s attention again to the limited nature of this dataset: “Correlation analyses, moreover, were inconclusive but suggested that the tool may benefit those participants who are younger and/or have lower phonological processing scores (see Results). However, these results were not significant (after multiple comparison correction) and, due to unreliable practice data (see Statistics), insufficient to support any conclusions regarding the relationship between subject characteristics and benefits conferred by phonemic cues” (P21, L11-16) We have also added the following to the limitations section: “First, future studies should more efficiently, and quantitatively and qualitatively, monitor practice adherence and cadence at home to better explore the relationship between exposure and reading-related measures.” (P24, L7-9) Results: P16:L21: You are using ß (latin sharp s) instead of β (greek beta), as well as a p-value with 5 decimals. We have amended accordingly. P18:L1-3: Repeated use of "pronounced". Furthermore, the effect was not "particularly pronounced" for the pseudoword decoding, but it was exclusively there. We have amended the paragraph to read: “These findings show that without the constraints of time during testing, there was a beneficial effect of access to the phonemic cue for single word decoding. This benefit was observed in the case of pseudoword decoding where children were asked to pronounce novel words in isolation. Moreover, prediction analyses suggest that this effect is more pronounced for those participants with more significant impairments in phonological processing and lower IQ (though these effects did not surpass our adjusted significance threshold of p < 0.0125). “ P18:L3-5: Please present the statistics if you refer to effects which suggest something. Also in the Matlab script. Now when we say suggest, we provide relevant statistics that suggest this idea. In cases that we do not have a statistic, but suggestion of a hypothesis in future work, we note that in the paper. P19:L3: Significance and effect sizes are completely unrelated. We can observe giant yet non-significant effects, as well as extremely small (and thus irrelevant) yet highly significant effects. We have amended the statement to read: “Further, data reflected effect sizes of d = 0.36 for accuracy and 0.26 for rate.” Discussion: P19:L22-P20:L4: I find this part a bit misleading as it jumps back and forth: things are inconclusive; a relation between pretest phonological awareness and intervention outcome is mentioned for which no statistic is provided in the results; two mentions that after p-value adjustment there were no effects left, but it is still discussed that these effects might suggest something. Especially at the start of the discussion there should be a clear summary of what was found first and then one can discuss what the presence and absence of effects might indicate. We thank the reviewer for this comment and agree that the start of the discussion section is not as clear as it should be. We have amended the first paragraph to provide a better summary of the findings before proceeding on to our interpretation. Furthermore, we have amended the wording to make clear that the relationship between phonological awareness and age is related to the correlation post-hoc analysis in the results section. “Using a RCT design, we tested the hypotheses that struggling readers could leverage a phonemic image cue placed below the vowels in digitally presented text to improve reading accuracy for isolated words and connected text, and that this benefit would be more pronounced for those readers with lower performance on measures on phonological processing. Data collected after a two-week period of unsupervised (but digitally monitored) practice demonstrated that struggling readers could read more complex words using the tool: compared to the control group, the intervention group showed a significantly larger improvement in decoding accuracy specifically for pseudo-words. As depicted in the results, this benefit did not extend to real words or either measure related to connected-text reading (accuracy and rate). Although there was no benefit, stable performance on measures of connected text read was observed for all participants with no significant difference between groups. The lack of benefits for connected text reading might reflect the limited training period or the increased cognitive demands of a novel approach to reading. These are important questions for future studies as generalization to connected text is of key importance. Correlation analyses, moreover, were inconclusive but suggest that the tool may benefit those participants who are younger and/or have lower phonological processing scores (see Results)” P20:L16: [...] highly inconsistent grapheme-phoneme mapping "of English" [...] The sentence now reads: “We focused on vowels because, in English, the highly inconsistent grapheme-phoneme mapping is a major hurdle for struggling readers.” P22:L3-L5: Throughout the manuscript there are sentences which use many commas and therefore appear encapsulated and make it unnecessarily hard for the reader to grasp the most relevant information. In this case you can straight out write that there is an improved decoding of pseudowords and remove the subordinate clause. As this is the only comment regarding writing style, I also wanted to note that there appear to be a lot of double spaces in manuscript. Content wise it could be highlighted here that this improvement appeared after being trained with the app for (only) 2 weeks. Thank you for this feedback - we certainly want to make sure that the manuscript is clear and readable. We have deleted the subordinate clause in that sentence and have removed double spaces. P23:L5-L13: This is an interesting observation. They appeared to have learned to make use of the cues, but this disappears in the testing situation. I feel it’s quite relevant to understand this better. Given that there is no performance penalty it seems like they might be able to ignore the cues altogether? One could take a more qualitative approach and investigate how children use it in these different scenarios or alternatively add an assessment mode into the app itself which blends with the intervention? This is a great point and represents a significant hurdle for any pedagogy to motivate a challenged reader even when they know it works/helps. We had a constant debate about how to measure use while reading and will contend that it is possible the children were ignoring the symbols altogether. During testing the children were reminded quite frequently to utilize the symbols, but as we discuss in the manuscript: ‘Either due to the limited practice period, limited supervised practice, or conflict with existing strategies children use when approaching challenging words,’ they didn’t feel as motivated to adopt the cues. We agree that future studies should work to collect better measures to piece this apart and have added the following to that paragraph in the discussion: “Future studies should incorporate qualitative and metacognitive methods to identify factors and circumstances that encourage struggling readers to adopt a novel strategy.” (P23, L22 - P24, L2). P23:L23: Given that parents also got a brief introduction to the app and you come back to it in the discussion it might be worth to recommend measuring parental involvement in future studies? We agree. The paragraph now reads: “Albeit promising, these results should be interpreted cautiously: Our power analysis indicated that we only had sufficient power to detect relatively large effects and many of the analyses were likely underpowered. Thus, future studies are needed with larger sample sizes to provide more conclusive results. Moreover, two additional points merit further investigation. First, future studies should more efficiently monitor practice adherence and cadence at home to better explore the relationship between exposure and reading-related measures. Second, given the short intervention period, we did not examine generalization to reading improvements without the cue and across different aspects of skilled reading. We only investigated whether the cue could be effectively used to decode more complex words. Thus, examining long-term learning effects and generalization to a variety of different contexts, as well as the role of parental involvement/participation is an important future direction.” Summary: In summary, the manuscript has much improved and will be a relevant addition to the field, but the provided analysis and code are not yet up to the required standards. We thank the reviewer for these relevant and meaningful recommendations for our manuscript. We believe the work is now much stronger for it. Attachment Submitted filename: Response To Reviewers.pdf Click here for additional data file. 10.1371/journal.pone.0243435.r005 Decision Letter 2 Männel Claudia Academic Editor © 2020 Claudia Männel2020Claudia MännelThis is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.Submission Version2 19 Oct 2020 PONE-D-19-34461R2 Annotating digital text with phonemic cues to support decoding in struggling readers PLOS ONE Dear Dr. Donnelly, Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. As you will see below, only Reviewer 2 suggests some further revisions. Please submit your revised manuscript by Dec 03 2020 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file. Please include the following items when submitting your revised manuscript: A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'. An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'. If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter. If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: http://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols We look forward to receiving your revised manuscript. Kind regards, Claudia Männel, PhD Academic Editor PLOS ONE [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions Comments to the Author 1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation. Reviewer #2: (No Response) ********** 2. Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. Reviewer #2: Partly ********** 3. Has the statistical analysis been performed appropriately and rigorously? Reviewer #2: No ********** 4. Have the authors made all data underlying the findings in their manuscript fully available? The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified. Reviewer #2: Yes ********** 5. Is the manuscript presented in an intelligible fashion and written in standard English? PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here. Reviewer #2: Yes ********** 6. Review Comments to the Author Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters) Reviewer #2: I thank the authors once more for addressing the points I raised and providing an improved manuscript. A few minor points remain to be solved. First, some comments relating to your response letter: You mentioned that you provide violin plots as S1, but these were not visible in the resubmission. Furthermore, I would strongly suggest to replace the barplots in figures 1 and 2 with these violin plot. In the last revision, I saw the document/section with the “transparent changes” but it was not part of this new resubmission and is neither in the repository (as is mentioned in the text) nor the pre-registration. You wrote: "During testing the children were reminded quite frequently to utilize the symbols, but as we discuss in the manuscript: ‘Either due to the limited practice period, limited supervised practice, or conflict with existing strategies children use when approaching challenging words,’ they didn’t feel as motivated to adopt the cues." Depending on whether these reminders happened during or in between testing trials this might be problematic, as it shifts the concept you measured away from the benefit of Sound It Out in a naturalistic setting, where children show what they learned, to one where adults instruct children what to do. This might also relate to the previously discussed finding that the children appeared to have learned to make use of the cues, but that this disappears in the testing situation. Can you clarify the exact procedure when children were “frequently reminded” to make use of the cues and add this to the method/discussion? P14:L21: Could you clarify whether the randomization of the reading lists happened at the individual or intervention group level? Please replace random effect with random intercept (as opposed to random slope, which are both a type of random effect): P16:L11 a random intercept per participant P16:L13 a random intercept for word list P16:L17: this is not "post hoc" analysis P17-P19: Thank your for explaining and improving the mixed models, which now seem appropriately done. The rest of the statistical analyses remain problematic though. First of all, only significant effects should be followed up with post-hoc tests. Otherwise, you run the risk of discovering spurious significant effects, and you have to adjust for multiple comparisons as well. The mixed model approach is much superior to t-tests because it takes into account the variance that can be attributed to test items and subjects in the random effects structure. In sum, the result of the model weight more than the t-tests, and if the model does not describe a significant interaction, this is the end of the analysis. If you wanted to base your discussion on the results of the t-tests, you would not have to run any models in the first place. If you want to do a post-hoc analysis of a mixed model you should furthermore not revert to t-tests but to approaches like least square means for multiple comparisons. Usually this is only done when you have factors with more than 2 levels and the pairwise comparisons cannot be read out of the model summary anymore (which does not apply to your model where group and time have 2 levels each). Instead, or on top, of plotting the raw data you might also want to plot the effects that the mixed model describes. P19:L7 – In the response letter you indicated to replace prediction analysis with correlation analysis, but it still says prediction analyses here. P21:L2-5: As mentioned in my last revision, while it is true that the intervention group had a bigger gain in pseudoword decoding than the control group, this has to be seen in the light of a significantly lower starting level. At post-test the two groups read equally well and from my perspective this should be the main result. Looking purely at increase from pre- to post-test is usually not relevant because both, regression to the mean as well as ceiling effects can explain your findings. All of these things are not discussed so far, and this is also why it is so important to provide better figures that tell the entire story. Furthermore, if you want to draw a causal conclusion (i.e. that the game improves pseudoword decoding) the two intervention groups have to fulfil the requirement of exchangeability. Usually in an RCT this is given or assumed, but in small samples it often does not hold and cannot be corrected or controlled for, due to the limited number of covariates such a small regression analysis can afford. I also only noticed now that the randomization took place after the (pre-test) outcome measures had already been collected (P11:L22). So you were in the position to avoid this scenario and produce exchangeable groups during randomization (e.g. by matching). In any case, I would be very careful to draw causal conclusions in this scenario. P21:L11-19. I still think this paragraph needs to be restructured as I suggested previously. First of all, it should be mentioned that there were no significant effects of the correlation analysis. Then you can discuss possible reasons (low sample size, multiple comparisons, etc.) and the trends you observed in the data at hand, before concluding that this is not enough evidence to support any firm claims. In sum, I think the manuscript is on a good way. The introduction and method sections are ok. For the results I recommend removing the post-hoc t-tests and replacing the barplots with violin plots or model predictions. The discussion still has to be adjusted to reflect the actual findings, as some claims are still too strong and are not backed up by the data and the experimental design. ********** 7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy. Reviewer #2: Yes: Toivo Glatz [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.] While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step. 10.1371/journal.pone.0243435.r006 Author response to Decision Letter 2 Submission Version3 6 Nov 2020 We would like to thank the reviewers once again for their thoughtful and insightful comments on this study. We have made further revisions and agree with the reviewers’ points that these changes will make this manuscript suitable for publishing. The reviewers comments are below in black and our responses and specific actions are noted in blue. Reviewer #2: I thank the authors once more for addressing the points I raised and providing an improved manuscript. A few minor points remain to be solved. First, some comments relating to your response letter: You mentioned that you provide violin plots as S1, but these were not visible in the resubmission. Furthermore, I would strongly suggest to replace the barplots in figures 1 and 2 with these violin plot. We apologize that this figure was not readily visible to you. In line with the PLOS ONE guidelines, it was in the zip file of Supplemental Information. We agree that the violin plots are a very important addition to transparency in our data visualizations. We have amended figures 1 and 2 to now include these violin plots in addition to the bar plots. The figure also now displays differences scores for each individual as lines on the violin plots. In the last revision, I saw the document/section with the “transparent changes” but it was not part of this new resubmission and is neither in the repository (as is mentioned in the text) nor the pre-registration. We apologize for this oversight. The document can be found in our pre registration file repository, here: https://osf.io/cy5ms/. You wrote: "During testing the children were reminded quite frequently to utilize the symbols, but as we discuss in the manuscript: ‘Either due to the limited practice period, limited supervised practice, or conflict with existing strategies children use when approaching challenging words,’ they didn’t feel as motivated to adopt the cues." Depending on whether these reminders happened during or in between testing trials this might be problematic, as it shifts the concept you measured away from the benefit of Sound It Out in a naturalistic setting, where children show what they learned, to one where adults instruct children what to do. This might also relate to the previously discussed finding that the children appeared to have learned to make use of the cues, but that this disappears in the testing situation. Can you clarify the exact procedure when children were “frequently reminded” to make use of the cues and add this to the method/discussion? We apologize that our previous explanation was ambiguous and agree that frequent reminders and researcher-motivations would impact our study’s findings and interpretations. To clarify, at the start of each testing session children were reminded: (1) that they were not being timed so they could take their time with each word to be as accurate as possible and (2) that the symbols were there to help them should they come to a challenging word. The following has been added to the methods section: “At the start of administration, all participants were reminded that they were not being time and encouraged to read as accurately as possible. For intervention participants exposed to the image cue, participants were additionally reminded that the symbols were there to help them should they come to a challenging word.” (p15, 2-5) “As with the decoding measures, participants were reminded at the start of administration that they were not being timed, encouraged to read as accurately as possible, and (for intervention participants) that the symbols were there should they come to a challenging word.” (p 15, 18-21) P14:L21: Could you clarify whether the randomization of the reading lists happened at the individual or intervention group level? Randomization of reading lists happened at the individual level. Please replace random effect with random intercept (as opposed to random slope, which are both a type of random effect): P16:L11 a random intercept per participant P16:L13 a random intercept for word list We have updated the terminology accordingly. P16:L17: this is not "post hoc" analysis As these correlation analyses weren’t planned and we considered them an interesting extension of the observed data, we feel that they satisfy the definition of post-hoc (we also realize that there are a variety of definitions of “post-hoc”). To make clear that they were performed with limited data, we have revised the manuscript to reflect that these are considered exploratory in nature. P17-P19: Thank your for explaining and improving the mixed models, which now seem appropriately done. The rest of the statistical analyses remain problematic though. First of all, only significant effects should be followed up with post-hoc tests. Otherwise, you run the risk of discovering spurious significant effects, and you have to adjust for multiple comparisons as well. The mixed model approach is much superior to t-tests because it takes into account the variance that can be attributed to test items and subjects in the random effects structure. In sum, the result of the model weight more than the t-tests, and if the model does not describe a significant interaction, this is the end of the analysis. If you wanted to base your discussion on the results of the t-tests, you would not have to run any models in the first place. If you want to do a post-hoc analysis of a mixed model you should furthermore not revert to t-tests but to approaches like least square means for multiple comparisons. Usually this is only done when you have factors with more than 2 levels and the pairwise comparisons cannot be read out of the model summary anymore (which does not apply to your model where group and time have 2 levels each). Instead, or on top, of plotting the raw data you might also want to plot the effects that the mixed model describes. Although we do not agree with the reviewers points here, we have, nonetheless, removed the t-statistics from the results section. It is our opinion that many readers will be left wondering about those analyses and that they provide an interesting, contextual look into the results that are couched within the interaction effects. The reason that we had included those statistics is in response to specific requests from others that have given feedback on this work. But, once again, those statistics are not critical to the paper and we have removed them. P19:L7 – In the response letter you indicated to replace prediction analysis with correlation analysis, but it still says prediction analyses here. Thanks for catching this - have amended accordingly. P21:L2-5: As mentioned in my last revision, while it is true that the intervention group had a bigger gain in pseudoword decoding than the control group, this has to be seen in the light of a significantly lower starting level. At post-test the two groups read equally well and from my perspective this should be the main result. Looking purely at increase from pre- to post-test is usually not relevant because both, regression to the mean as well as ceiling effects can explain your findings. All of these things are not discussed so far, and this is also why it is so important to provide better figures that tell the entire story. Furthermore, if you want to draw a causal conclusion (i.e. that the game improves pseudoword decoding) the two intervention groups have to fulfil the requirement of exchangeability. Usually in an RCT this is given or assumed, but in small samples it often does not hold and cannot be corrected or controlled for, due to the limited number of covariates such a small regression analysis can afford. I also only noticed now that the randomization took place after the (pre-test) outcome measures had already been collected (P11:L22). So you were in the position to avoid this scenario and produce exchangeable groups during randomization (e.g. by matching). In any case, I would be very careful to draw causal conclusions in this scenario. We thank the reviewer for this comment and agree that care needs to be taken to ensure that our interpretations are supported by our results. This study represents a small-scale ‘proof of concept’ study with encouraging results, but we take care to describe our findings without use of causal language. As regression to the mean and ceiling effects are important factors to keep in mind, the following has been adjusted in the discussion: “Albeit promising, these results should be interpreted cautiously: Our power analysis indicated that we only had sufficient power to detect relatively large effects and many of the analyses (e.g., individual differences) were likely underpowered. Also, as there was a significant difference at pre-test for our sole finding with pseudo word decoding, future studies are needed to rule out the role of regression toward the mean and possible ceiling effects. Thus, future studies are needed with larger sample sizes to provide more conclusive results.” (p24.,10-12) As to randomization, the participants were randomly assigned to groups prior to participation in the pre-test for the study. This is explained in the sentence following the one you reference: “Randomization was unconstrained with group assignment determined at time of consent” We understand that the language in the Study Design section is ambiguous - the intention there was to make it clear that all participants had identical pre-testing prior to deviating in study protocol for their assigned group. We have amended that section to instead read: “In a randomized pre-post design, participants were randomized to a control or intervention condition. Randomization was unconstrained with group assignment determined at time of consent; however, sibling participants were assigned to the same group to better control participant adherence. Both intervention and control participants completed an initial, baseline session that collected all outcome measures using the normal text condition presented on a Kindle fire tablet without the Sound It Out cue (see Outcome Measures).” (p11-12, 20-2) P21:L11-19. I still think this paragraph needs to be restructured as I suggested previously. First of all, it should be mentioned that there were no significant effects of the correlation analysis. Then you can discuss possible reasons (low sample size, multiple comparisons, etc.) and the trends you observed in the data at hand, before concluding that this is not enough evidence to support any firm claims. For clarity with our interpretations, we have amended that paragraph to read: “Correlation analyses, after multiple comparison correction, revealed no significant relationships between our variables of interest and benefit of the cue. Due to unreliable practice data (see Statistics), analyses cannot support any conclusions regarding the relationship between subject characteristics and benefits conferred by phonemic cues. However, results suggest that the tool may benefit those participants who are younger and/or have lower phonological processing scores (see Results). Together, although most analyses failed to meet our adjusted significance threshold, data suggests that participants were able to effectively use the cues in isolated situations (i.e. pseudoword reading), but the tool did not become sufficiently automatic to produce significant gains in passage reading fluency.” (p.21-2, 16-2) In sum, I think the manuscript is on a good way. The introduction and method sections are ok. For the results I recommend removing the post-hoc t-tests and replacing the barplots with violin plots or model predictions. The discussion still has to be adjusted to reflect the actual findings, as some claims are still too strong and are not backed up by the data and the experimental design. We greatly appreciate the reviewer’s support and commitment to making this work a more impactful contribution to the literature. 10.1371/journal.pone.0243435.r007 Decision Letter 3 Männel Claudia Academic Editor © 2020 Claudia Männel2020Claudia MännelThis is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.Submission Version3 20 Nov 2020 PONE-D-19-34461R3 Annotating digital text with phonemic cues to support decoding in struggling readers PLOS ONE Dear Dr. Donnelly, Thank you for submitting your revised manuscript to PLOS ONE. We invite you to submit a revised version of the manuscript that incorporates the final minor changes requested by Reviewer 2. I will then be happy to accept your manuscript for publication. Please submit your revised manuscript by Jan 04 2021 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file. Please include the following items when submitting your revised manuscript: A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'. An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'. If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter. If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: http://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols We look forward to receiving your revised manuscript. Kind regards, Claudia Männel, PhD Academic Editor PLOS ONE [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions Comments to the Author 1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation. Reviewer #2: (No Response) ********** 2. Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. Reviewer #2: Yes ********** 3. Has the statistical analysis been performed appropriately and rigorously? Reviewer #2: Yes ********** 4. Have the authors made all data underlying the findings in their manuscript fully available? The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified. Reviewer #2: Yes ********** 5. Is the manuscript presented in an intelligible fashion and written in standard English? PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here. Reviewer #2: Yes ********** 6. Review Comments to the Author Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters) Reviewer #2: Thank you for your detailed answers and providing a, once more, much improved manuscript. I would suggest incorporating the following minor changes prior to publication: You now mention the differences in pseudoword decoding at pre-test in the discussion, but these are not yet reported in the results section. Please add the relevant statistics to the results section, if possible with effect size. I would suggest to add "small scale RCT" (on P21:L11) and/or "small scale proof of concept" (on P26:L17). P24:L2-5: The repeated use of "future studies" reads a bit awkward. I'd suggest to slightly rephrase. ********** 7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy. Reviewer #2: Yes: Toivo Glatz [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.] While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step. 10.1371/journal.pone.0243435.r008 Author response to Decision Letter 3 Submission Version4 20 Nov 2020 We would like to thank the reviewer and have provided responses for each concern. The reviewers comments are below in black and our responses and specific actions are noted in blue. ------------------------------------------------------ Reviewer #2: Thank you for your detailed answers and providing a, once more, much improved manuscript. I would suggest incorporating the following minor changes prior to publication: You now mention the differences in pseudoword decoding at pre-test in the discussion, but these are not yet reported in the results section. Please add the relevant statistics to the results section, if possible with effect size. The results has been added and now reads: “At pretest, despite randomization, the intervention group by chance had lower scores than the control group: this is evidenced by the significant main effect of group in the mixed effects model (β=-6.2, t(156) = -2.65, p=0.009).” I would suggest to add "small scale RCT" (on P21:L11) and/or "small scale proof of concept" (on P26:L17). We have amended the text to reflect your suggestions. Those lines are now amended to read: “Using a small scale RCT design, we tested the hypotheses that …”, and “In aggregate, these findings represent a small scale proof-of-concept for this …” P24:L2-5: The repeated use of "future studies" reads a bit awkward. I'd suggest to slightly rephrase. We agree and have amended to the following: “Also, as there was a significant difference at pre-test for our sole finding with pseudo word decoding, future studies are needed to rule out the role of regression toward the mean and possible ceiling effects. Thus, future work is needed with larger sample sizes to provide more conclusive results. Moreover, two additional points merit further investigation. First, future experiments should more efficiently, and quantitatively and qualitatively, monitor practice adherence and cadence at home to better explore the relationship between exposure and reading-related measures.” Attachment Submitted filename: Response to Reviewers.pdf Click here for additional data file. 10.1371/journal.pone.0243435.r009 Decision Letter 4 Männel Claudia Academic Editor © 2020 Claudia Männel2020Claudia MännelThis is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.Submission Version4 23 Nov 2020 Annotating digital text with phonemic cues to support decoding in struggling readers PONE-D-19-34461R4 Dear Dr. Donnelly, We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements. Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication. An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org. If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org. Kind regards, Claudia Männel, PhD Academic Editor PLOS ONE Additional Editor Comments (optional): Reviewers' comments: 10.1371/journal.pone.0243435.r010 Acceptance letter Männel Claudia Academic Editor © 2020 Claudia Männel2020Claudia MännelThis is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. 25 Nov 2020 PONE-D-19-34461R4 Annotating digital text with phonemic cues to support decoding in struggling readers Dear Dr. Donnelly: I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now with our production department. If your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information please contact onepress@plos.org. If we can help with anything else, please email us at plosone@plos.org. Thank you for submitting your work to PLOS ONE and supporting open access. Kind regards, PLOS ONE Editorial Office Staff on behalf of Dr. Claudia Männel Academic Editor PLOS ONE ==== Refs References 1 National Assessment of Educational Progress. The Nation’s Report Card. Reading 2005. 2006. Available: http://files.eric.ed.gov/fulltext/ED486463.pdf 2 Committee on the Prevention of Reading Difficulties in Young Children . Preventing reading difficulties in young children . Snow CE , Burns MS , Griffin P , editors. Washington, DC : National Academy Press ; 1998 10.2307/495689 3 Moats LC . Teaching Reading Is Rocket Science: What Expert Teachers of Reading Should Know and Be Able To Do . American Federation of Teachers ; 1999 6 Available: https://eric.ed.gov/?id=ED445323 4 Lyon GR , Shaywitz SE , Shaywitz BA . A definition of dyslexia . Ann Dyslexia . 2003 ;53 : 1 –14 . 10.1007/s11881-003-0001-9 5 Wolf M . Proust and the Squid: The Story and Science of the Reading Brain . New York : HarperCollins ; 2007 . 6 National Assessment of Educational Progress. NAEP Report Card: Reading. 2019. Available: https://www.nationsreportcard.gov/reading?grade=4 7 Delany K . The Experience of Parenting a Child With Dyslexia: An Australian perspective . J Student Engagem Educ Matters . 2017 ;7 Available: http://ro.uow.edu.au/jseem/vol7/iss1/6 8 Beach KD , Ellen M , Zoi A. P , Maryann M , Paola P , Vintinner JP . Effects of a Summer Reading Intervention on Reading Skills for Low-Income Black and Hispanic Students in Elementary School . Read Writ Q . 2018 ;34 : 1 –18 . 10.1080/10573569.2018.1446859 9 Cheung ACK , Slavin RE . The effectiveness of education technology for enhancing reading achievement: A meta-analysis . Best Evid Encycl . 2011 ;97 : 1 –48 . 10.3354/dao02401 22303625 10 Guernsey L , Levine MH . Tap, Click, Read: Growing Readers in a World of Screens . San Francisco : Jossey-Bass ; 2015 . 11 Shaheen NL , Lohnes Watulak S . Bringing Disability Into the Discussion: Examining Technology Accessibility as An Equity Concern in the Field of Instructional Technology . J Res Technol Educ . 2019 ;51 : 187 –201 . 10.1080/15391523.2019.1566037 12 Daugherty L , Dossani R , Johnson E-E , Wright C . Getting on the Same Page: Identifying Goals for Technology Use in Early Childhood Education . RAND Corporation ; 2014 Available: https://www.rand.org/pubs/research_reports/RR673z1.html 13 National Reading Panel . Teaching Children to Read: An Evidence-Based Assessment of the Scientific Research Literature on Reading and Its Implications for Reading Instruction . Washington, D.C. ; 2000 . 14 Lovett MW , Frijters JC , Wolf M , Steinbach KA , Sevcik RA , Morris RD . Early intervention for children at risk for reading disabilities: The impact of grade at intervention and individual differences on intervention outcomes . J Educ Psychol . 2017 ;109 : 889 –914 . 10.1037/edu0000181 15 Castles A , Rastle K , Nation K . Ending the Reading Wars: Reading Acquisition From Novice to Expert . Psychol Sci Public Interes . 2018 ;19 : 5 –51 . 10.1177/1529100618772271 29890888 16 Bradley L , Bryant PE . Categorizing sounds and learning to read—a causal connection . Nature . 1983 ;301 : 419 –421 . 10.1038/301419a0 17 Storch S a , Whitehurst GJ . Oral language and code-related precursors to reading: evidence from a longitudinal structural model . Dev Psychol . 2002 ;38 : 934 –947 . 10.1037/0012-1649.38.6.934 12428705 18 Foorman BR , Francis DJ , Fletcher JM , Schatschneider C , Mehta PD . The role of instruction in learning to read: Preventing reading failure in at-risk children . J Educ Psychol . 1998 ;90 : 37 –55 . 10.1037/0022-0663.90.1.37 19 Torgesen JK , Alexander AW , Wagner RK , Rashotte CA , Voeller KKS , Conway T . Intensive Remedial Instruction for Children with Severe Reading Disabilities: Immediate and Long-term Outcomes From Two Instructional Approaches . J Learn Disabil . 2001 ;34 : 33 –58 . 10.1177/002221940103400104 15497271 20 Scanlon DM , Vellutino FR . A Comparison of the Instructional Backgrounds and Cognitive Profiles of Poor, Average, and Good Readers Who Were Initially Identified as At Risk for Reading Failure . Sci Stud Read . 1997 ;1 : 191 –215 . 10.1207/s1532799xssr0103_2 21 Wanzek J , Stevens EA , Williams KJ , Scammacca NK , Vaughn S , Sargent K . Current Evidence on the Effects of Intensive Early Reading Interventions . J Learn Disabil . 2018 ; 002221941877511. 10.1177/0022219418775110 29779424 22 Vaughn S , Fletcher JM , Francis DJ , Denton CA , Wanzek J , Wexler J , et al Response to Intervention with Older Students with Reading Difficulties . Learn Individ Differ . 2008 ;18 : 338 –345 . 10.1016/j.lindif.2008.05.001 19129920 23 Slavin RE , Lake C , Davis S , Madden NA . Effective programs for struggling readers: A best-evidence synthesis . Educ Res Rev . 2011 ;6 : 1 –26 . 10.1016/j.edurev.2010.07.002 24 Cheung ACK , Slavin RE . Effects of Educational Technology Applications on Reading Outcomes for Struggling Readers: A Best-Evidence Synthesis . Read Res Q . 2013 ;48 : 277 –299 . 10.1002/rrq.50 25 Morris RD , Lovett MW , Wolf M , Sevcik RA , Steinbach KA , Frijters JC , et al Multiple-Component Remediation for Developmental Reading Disabilities . J Learn Disabil . 2012 ;45 : 99 –127 . 10.1177/0022219409355472 20445204 26 Clark DB , Tanner-Smith EE , Killingsworth SS . Digital Games, Design, and Learning: A Systematic Review and Meta-Analysis . Rev Educ Res . 2016 ;86 : 79 –122 . 10.3102/0034654315582065 26937054 27 Laurillard D . Learning “Number Sense” through Digital Games with Intrinsic Feedback . Australas J Educ Technol . 2016 ;32 : 32 –44 . 28 de Souza GN , Brito YP dos S , Tsutsumi MMA , Marques LB , Goulart PRK , Monteiro DC , et al The Adventures of Amaru: Integrating Learning Tasks Into a Digital Game for Teaching Children in Early Phases of Literacy . Front Psychol . 2018 ;9 : 2531 10.3389/fpsyg.2018.02531 30618954 29 Dowker A . Early identification and intervention for students with mathematics difficulties . J Learn Disabil . 2005 ;38 : 324 –32 . 10.1177/00222194050380040801 16122064 30 Rose DH , Strangman N . Universal Design for Learning: Meeting the challenge of individual learning differences through a neurocognitive perspective Universal Access in the Information Society . Springer ; 2007 pp. 381 –391 . 10.1007/s10209-006-0062-8 31 Messer D , Nash G . An evaluation of the effectiveness of a computer-assisted reading intervention . J Res Read . 2018 ;41 : 140 –158 . 10.1111/1467-9817.12107 32 Seward R , O’Brien B , Breit-Smith AD , Meyer B . Linking Design Principles with Educational Research Theories to Teach Sound to Symbol Correspondence with Multisensory Type . Visible Lang . 2014 ;48 : 87 –108 . Available: http://web.a.ebscohost.com.offcampus.lib.washington.edu/ehost/detail/detail?vid=0&sid=7dd32118-a4fd-4bff-a325-2eb983ba9092%40sessionmgr4006&bdata=JnNpdGU9ZWhvc3QtbGl2ZQ%3D%3D#AN=99932728&db=aft 33 O’Brien BA , Habib M , Onnis L . Technology-Based Tools for English Literacy Intervention: Examining Intervention Grain Size and Individual Differences . Front Psychol . 2019 ;10 : 2625 10.3389/fpsyg.2019.02625 31849754 34 Kyle FE , Kujala J , Richardson U , Lyytinen H , Goswami U . Assessing the effectiveness of two theoretically motivated computerassisted reading interventions in the United Kingdom: GG Rime and GG Phoneme . Read Res Q . 2013 ;48 : 61 –76 . 10.1002/rrq.038 35 Chall JS . Learning to Read: The Great Debate . New York : McGraw-Hill Book Company ; 1983 . 36 Hutchison A , Beschorner B , Schmidt-Crawford D . Exploring the Use of the iPad for Literacy Learning . Read Teach . 2012 ;66 : 15 –23 . 10.1002/TRTR.01090 37 Kim MK , McKenna JW , Park Y . The Use of Computer-Assisted Instruction to Improve the Reading Comprehension of Students With Learning Disabilities: An Evaluation of the Evidence Base According to the What Works Clearinghouse Standards . Remedial Spec Educ . 2017 ;38 : 233 –245 . 10.1177/0741932517693396 38 Wolf M , Barzillai M , Gottwald S , Miller L , Spencer K , Norton ES , et al The RAVE-O intervention: Connecting neuroscience to the classroom . Mind, Brain, Educ . 2009 ;3 : 84 –93 . 10.1111/j.1751-228X.2009.01058.x 39 Lovett MW , Lacerenza L , Borden SL , Frijters JC , Steinbach KA , De Palma M . Components of effective remediation for developmental reading disabilities: Combining phonological and strategy-based instruction to improve outcomes . J Educ Psychol . 2000 ;92 : 263 –283 . 10.1037/0022-0663.92.2.263 40 Oakland T , Black JL , Stanford G , Nussbaum NL , Balise RR . An Evaluation of the Dyslexia Training Program: A Multisensory Method for Promoting Reading in Students with Reading Disabilities . J Learn Disabil . 1998 ;31 : 140 –147 . 10.1177/002221949803100204 9529784 41 Mandel Morrow L , Asbury E . What Should We Do About Phonics? In: Gambrell LB , Mandel Morrow L , Neuman SB , Pressley M , editors. Best Practices in Literacy Instruction . New York, NY : The Guilford Press ; 1999 pp. 68 –89 . 42 Pennington BF . From single to multiple deficit models of developmental disorders . Cognition . 2006 ;101 : 385 –413 . 10.1016/j.cognition.2006.04.008 16844106 43 Olson RK , Wise BW . Reading on the computer with orthographic and speech feedback . Read Writ . 1992 ;4 : 107 –144 . 10.1017/CBO9781107415324.004 44 Takacs ZK , Swart EK , Bus AG . Benefits and Pitfalls of Multimedia and Interactive Features in Technology-Enhanced Storybooks . Rev Educ Res . 2015 ;85 : 698 –739 . 10.3102/0034654314566989 26640299 45 Bryant ND . Some Principles of Remedial Instruction for Dyslexia . Read Teach . 1965 ;18 : 567 –572 . 46 Hornung C , Martin R , Fayol M . The power of vowels: Contributions of vowel, consonant and digit RAN to clinical approaches in reading development . Learn Individ Differ . 2017 ;57 : 85 –102 . 10.1016/j.lindif.2017.06.006 47 Medler DA, Binder JR. MCWord: An Orthographic Wordform Database. 2005. Available: http://www.neuro.mcw.edu/mcword/ 48 The MathWorks I . MATLAB Statistics and Machine Learning Toolbox . Natick, MA, USA ; 2017 . 49 Slavin RE , Cheung ACK , Groff C , Lake C . Effective Reading Programs for Middle and High Schools: A Best-Evidence Synthesis . Read Res Q . 2008 ;43 : 290 –322 . 10.1598/RRQ.43.3.4 50 Wise BW , Olson RK . Computer-based phonological awareness and reading instruction . Ann Dyslexia . 1995 ;45 : 97 –122 . 10.1007/BF02648214 24234190 51 Wolf M , Bowers PG . Naming-Speed Processes and Developmental Reading Disabilities: An Introduction to the Special Issue on the Double-Deficit Hypothesis . J Learn Disabil . 1997 ;33 : 322 –324 . Available: http://journals.sagepub.com/doi/pdf/10.1177/002221940003300404 52 Norton ES , Wolf M . Rapid Automatized Naming (RAN) and Reading Fluency: Implications for Understanding and Treatment of Reading Disabilities . Annu Rev Psychol . 2012 ;63 : 427 –52 . 10.1146/annurev-psych-120710-100431 21838545 53 Wolf M , Katzir-Cohen T . Reading Fluency and Its Intervention . Sci Stud Read . 2001 ;5 : 211 –239 . 10.1207/S1532799XSSR0503 54 Torgesen JK , Wagner RK , Rashotte CA , Herron J , Lindamood P . Computer-assisted instruction to prevent early reading difficulties in students at risk for dyslexia: Outcomes from two instructional approaches . Ann Dyslexia . 2010 ;60 : 40 –56 . 10.1007/s11881-009-0032-y 20052566 55 Berninger VW , Lee YL , Abbott RD , Breznitz Z . Teaching children with dyslexia to spell in a reading-writers’ workshop . Ann Dyslexia . 2013 ;63 : 1 –24 . 10.1007/s11881-011-0054-0 21845501 56 Farkas WA , Jang BG . Designing, Implementing, and Evaluating a School-Based Literacy Program for Adolescent Learners With Reading Difficulties: A Mixed-Methods Study . Read Writ Q . 2019 ; 1 –17 . 10.1080/10573569.2018.1541770 57 Ronimus M , Lyytinen H . Is School a Better Environment Than Home for Digital Game-Based Learning? The Case of GraphoGame . Hum Technol An Interdiscip J Humans ICT Environ . 2015 ;11 : 123 –147 . 10.17011/ht/urn.201511113637 58 Francom GM . Barriers to technology integration: A time-series survey study . J Res Technol Educ . 2019 ; 1 –16 . 10.1080/15391523.2019.1679055 59 Dexter S , Richardson JW . What does technology integration research tell us about the leadership of technology? J Res Technol Educ . 2019 ; 1 –20 . 10.1080/15391523.2019.1668316 60 Lindeblad E , Nilsson S , Gustafson S , Svensson I . Assistive technology as reading interventions for children with reading impairments with a one-year follow-up . Disabil Rehabil Assist Technol . 2017 ;12 : 713 –724 . 10.1080/17483107.2016.1253116 27924656 61 Goodall J . Narrowing the achievement gap: Parental engagement with children’s learning . London : Routledge ; 2017 Available: https://content.taylorfrancis.com/books/download?dac=C2015-0-60071-2&isbn=9781317373247&format=googlePreviewPdf 62 Lynch J , Anderson J , Anderson A , Shapiro J . Parents’ Beliefs About Young Children’s Literacy Development And Parents’ Literacy Behaviors . Read Psychol . 2006 ;27 : 1 –20 . 10.1080/02702710500468708 63 Hannon P , James S . Parents’ and Teachers’ Perspectives on Preschool Literacy Development . Br Educ Res J . 1990 ;16 : 259 –272 . 10.1080/0141192900160304 64 Tichnor-Wagner A , Garwood JD , Bratsch-Hines M , Vernon-Feagans L . Home Literacy Environments and Foundational Literacy Skills for Struggling and Nonstruggling Readers in Rural Early Elementary Schools . Learn Disabil Res Pract . 2016 ;31 : 6 –21 . 10.1111/ldrp.12090 65 Auerbach ER . Toward a Social-Contextual Approach to Family Literacy . Harv Educ Rev . 1989 ;59 : 165 –182 . 10.17763/haer.59.2.h23731364l283156 66 Kraft MA , Monti-Nussbaum M . Can Schools Enable Parents to Prevent Summer Learning Loss? A Text-Messaging Field Experiment to Promote Literacy Skills . Ann Am Acad Pol Soc Sci . 2017 ;674 : 85 –112 . 10.1177/0002716217732009