==== Front PLoS One PLoS One plos PLOS ONE 1932-6203 Public Library of Science San Francisco, CA USA 10.1371/journal.pone.0288027 PONE-D-23-09235 Research Article Medicine and Health Sciences Mental Health and Psychiatry Mood Disorders Depression Medicine and Health Sciences Mental Health and Psychiatry Neuropsychiatric Disorders Anxiety Disorders Medicine and Health Sciences Mental Health and Psychiatry Neuroses Anxiety Disorders Social Sciences Linguistics Semantics Biology and Life Sciences Psychology Emotions Anxiety Social Sciences Psychology Emotions Anxiety Biology and Life Sciences Neuroscience Cognitive Science Cognitive Psychology Clinical Psychology Biology and Life Sciences Psychology Cognitive Psychology Clinical Psychology Social Sciences Psychology Cognitive Psychology Clinical Psychology Medicine and Health Sciences Mental Health and Psychiatry Biology and Life Sciences Psychology Emotions Social Sciences Psychology Emotions Medicine and Health Sciences Diagnostic Medicine Have the concepts of ‘anxiety’ and ‘depression’ been normalized or pathologized? A corpus study of historical semantic change Have the concepts of ‘anxiety’ and ‘depression’ been normalized or pathologized? https://orcid.org/0009-0007-9050-258X Xiao Yu Conceptualization Formal analysis Investigation Methodology Project administration Visualization Writing – original draft Writing – review & editing 1 https://orcid.org/0000-0003-3873-5021 Baes Naomi Data curation Formal analysis Investigation Methodology Project administration Visualization Writing – review & editing 1 Vylomova Ekaterina Data curation Formal analysis Investigation Methodology Project administration Writing – review & editing 2 https://orcid.org/0000-0002-1913-2340 Haslam Nick Conceptualization Formal analysis Funding acquisition Methodology Project administration Resources Supervision Writing – original draft Writing – review & editing 1 * 1 School of Psychological Sciences, The University of Melbourne, Melbourne, Australia 2 School of Computing and Information Systems, The University of Melbourne, Melbourne, Australia Ptaszynski Michal Editor Kitami Institute of Technology, JAPAN Competing Interests: The authors have declared that no competing interests exist. * E-mail: nhaslam@unimelb.edu.au 29 6 2023 2023 18 6 e028802727 3 2023 16 6 2023 © 2023 Xiao et al 2023 Xiao et al https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Research on concept creep indicates that the meanings of some psychological concepts have broadened in recent decades. Some mental health-related concepts such as ‘trauma’, for example, have acquired more expansive meanings and come to refer to a wider range of events and experiences. ‘Anxiety’ and ‘depression’ may have undergone similar semantic inflation, driven by rising public attention and awareness. Critics have argued that everyday emotional experiences are increasingly pathologized, so that ‘depression’ and ‘anxiety’ have broadened to include sub-clinical experiences of sadness and worry. The possibility that these concepts have expanded to include less severe phenomena (vertical concept creep) was tested by examining changes in the emotional intensity of words in their vicinity (collocates) using two large historical text corpora, one academic and one general. The academic corpus contained >133 million words from psychology article abstracts published 1970–2018, and the general corpus (>500 million words) consisted of diverse text sources from the USA for the same period. We hypothesized that collocates of ‘anxiety’ and ‘depression’ would decline in average emotional severity over the study period. Contrary to prediction, the average severity of collocates for both words increased in both corpora, possibly due to growing clinical framing of the two concepts. The study findings therefore do not support a historical decline in the severity of ‘anxiety’ and ‘depression’ but do provide evidence for a rise in their pathologization. http://dx.doi.org/10.13039/501100000923 Australian Research Council DP170104948 https://orcid.org/0000-0002-1913-2340 Haslam Nick http://dx.doi.org/10.13039/501100000923 Australian Research Council DP210103984 https://orcid.org/0000-0002-1913-2340 Haslam Nick Australian Research Council Discovery Projects awarded to Nick Haslam supported this research (DP170104948 and DP210103984). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Data AvailabilityThe datasets generated and analyzed for this study can be found in the Open Science Framework Repository: https://osf.io/u9mve/?view_only=f877aa015c78454fb62d898e51b47158. Data Availability The datasets generated and analyzed for this study can be found in the Open Science Framework Repository: https://osf.io/u9mve/?view_only=f877aa015c78454fb62d898e51b47158. ==== Body pmcIntroduction Studies of concept creep [1] indicate that in recent decades many harm-related concepts in psychology have undergone semantic broadening. Concepts such as abuse, bullying, prejudice, and trauma have come to refer to an increasingly wide range of actions and experiences. In the 1970s, for example, bullying was defined as a specific kind of childhood peer aggression that was repeated, intentional, and perpetrated downward in a power hierarchy. Over time, these criteria have loosened. In addition to such prototypical cases, bullying now commonly refers to aggression that is unrepeated, unintentional, and directed upward or laterally rather than exclusively downward, and is invoked to refer to adult as well as childhood behavior. [1] documented this pattern of semantic inflation in several concepts and distinguished two main forms. Concepts creep horizontally when they stretch to refer to qualitatively different phenomena, such as when ‘bullying’ expands to include workplace misbehavior or to aggression carried out via social media rather than in person. They creep vertically when they expand to include less severe phenomena, such as when the threshold for identifying bullying is lowered by no longer requiring behavior to be repeated. In theory, concept meanings could shift vertically by excluding less severe phenomena, but by Haslam’s definition this would represent concept deflation or contraction rather than creep. The two forms of concept creep may co-occur. Research on concept creep, reviewed by Haslam et al. [2], points to several potential contributors and consequences of the phenomenon. Its drivers may include a declining objective prevalence of harm and a rising cultural sensitivity to harm, both resulting in less severe harms being recognized. Its consequences may be mixed. On the one hand, by recognizing a wider range of harms, concept creep may allow previously ignored forms of suffering to be legitimated and addressed, and previously tolerated forms of maltreatment to be problematized. Recognizing gambling problems as addictions, unconscious bias as prejudice, or online abuse as bullying signals that these undesirable phenomena should be taken seriously. On the other hand, concept creep may have some disadvantages. Broadening concepts of harm may trivialize those concepts [3], lead to disproportionate responses to relatively mild cases, and promote a sense of personal vulnerability and fragility [4, 5]. If the concept of trauma dilutes or normalizes to include everyday adversities, for example, the concept may become trivialized, victims of “small-t trauma” may seek unnecessary clinical interventions, and increasing numbers of people may identify as lastingly damaged by relatively innocuous events. Concept creep occurs for a wide assortment of harm-related concepts, but it may be especially germane to concepts of mental ill-health. Broadened concepts of mental illness have been studied extensively under the headings of diagnostic inflation [6], pathologization [7], and psychiatrization [8, 9]. Critics of the Diagnostic and Statistical Manual for Mental Disorders (DSM), for example, have argued that successive editions have expanded the range of experiences and behaviors that are identified as mental disorders (horizontal concept creep), and that it has loosened the criteria for some disorders so that a lower severity threshold applies for their diagnosis (vertical concept creep). Although the evidence for wholesale diagnostic inflation is weak [10], there is strong evidence that certain diagnoses have expanded [e.g., ADHD: 11], with clear implications for over-diagnosis, over-medication, and resulting misallocation of clinical resources. Alongside any such expansion of concepts of mental ill-health within psychiatry, there are also concerns that laypeople’s concepts may also be broadening, resulting in apparent epidemics of self-diagnosed conditions. In view of the timeliness of these concerns about the semantic broadening of mental health-related concepts, it is important to investigate them in academic and general discourse. To date, research on the creep of mental health-related concepts has only addressed ‘trauma’ [12]. Studies have shown that this concept has risen steeply in use within academic psychology, and that this rise in usage has been accompanied by a broadening of the range of semantic contexts in which the term is employed (horizontal creep) [13, 14]. This research further indicates that the increase in the use of the term plays a causal role in its semantic inflation. A recent study by Baes, Vylomova, Zyphur, and Haslam [15] yielded similar findings for the vertical creep of ‘trauma’. Using a novel method, Baes et al. showed that over a five-decade period beginning in the 1970s, ‘trauma’ came to be associated with less emotionally severe words within a large corpus of psychology article abstracts. Again, analysis indicated that the rising usage of ‘trauma’ contributed to this broadened meaning of the term, which experimental studies suggest can lead to people judging ‘trauma’ to be less severe and less societally important [3]. In view of the paucity of research on the creep of mental health-related concepts beyond trauma, we conducted a study on possible semantic shifts in the concepts of anxiety and depression. It is particularly important to investigate shifts in these two concepts because they are central to some of the most common psychological problems in the general population and are often discussed together due to their symptomatic similarities [16] and comorbidity [17, 18]. Critics have argued that people experiencing normal worry, fear, sadness, or sorrow are increasingly diagnosed with anxiety and depressive disorders [19]. Horwitz and Wakefield [7] made the case that psychiatry had transformed everyday sadness into major depression, and Horwitz and Wakefield [20] extended their analysis to anxiety disorders, arguing that DSM commonly misdiagnoses adaptive anxieties as clinical conditions. By proposing that the diagnostic threshold for these disorders is too low, creating numerous false positives, this critique relates most strongly to the vertical form of concept creep. The controversy over the removal of the bereavement exemption for the diagnosis of major depression in DSM-5 [21] exemplifies this concern over the loosening of diagnostic criteria, as does DSM-5’s [22] removal of the requirement that people receiving some anxiety disorder diagnoses must recognize that their anxiety is unreasonable. Concerns over increasingly liberal diagnostic criteria for anxiety and depressive disorders relate to professional concepts of disorder, embodied in formal diagnostic systems. However, it is equally important to evaluate shifts in concepts of anxiety and depression in the wider culture. There is evidence that these concepts may have also undergone semantic inflation among laypeople. Bröer and Besseling [23], for example, argue that ‘depression’ can refer to ordinary sadness or low mood in colloquial language rather than exclusively to pathological conditions. If this is the case, broadened lay concepts of anxiety and depression might contribute to excessive self-diagnosis and inappropriate treatment [24, 25]. For example, people experiencing sub-clinical distress might strain mental health services, divert resources from people with more severe conditions [26], and experience side-effects from unnecessary medications. Given evidence and concerns regarding the vertical concept creep of ‘anxiety’ and depression’ in both professional and everyday language, we carried out a study of historical shifts in the emotional severity associated with each term in two large text corpora: professional-academic (psychology article abstracts) and general (a curated corpus of diverse USA-derived texts). Employing the methodology developed in Baes et al.’s [15] analysis of ‘trauma’, we examined whether words occurring in the vicinity of instances of ‘anxiety’ or ‘depression’ (i.e., collocates) tended to become less negative and intense in connotation–based on established norms for the affective meaning of words [27]–from 1970 to 2018. Consistent with the declining pattern observed for ‘trauma’ by Baes et al. [15] and our expectation that ‘anxiety’ and ‘depression’ have come to refer to increasingly mild and everyday experiences, we hypothesized that these concepts would appear in the context of less emotionally severe words over time (vertical concept creep). In addition to testing this primary hypothesis in each corpus, we planned to conduct post hoc analyses to clarify whether specific collocates of ‘anxiety’ and ‘depression’ might account for any observed trends. Method Materials The psychology corpus comprised 871,340 abstracts from 875 psychology journals, collected from E-Research and PubMed databases, covering the period 1930 to 2019 [28]. The journal set was distributed broadly across the psychology discipline, including the following number of journals from each of the following, non-mutually exclusive Scimago journal classifications: 216 (24.7%) developmental and educational psychology, 171 (19.5%) clinical psychology, 158 (18.1%) social psychology, 156 (17.8%) psychology (miscellaneous), 144 (16.5%) applied psychology, 99 (11.3%) experimental and cognitive psychology, and 45 (5.1%) neuropsychology and physiological psychology. Abstracts were used due to copyright restrictions and the fact that abstracts can effectively capture the main content of scientific articles [29]. The final corpus of psychology abstracts was limited to 1970 to 2018 data, due to the relatively small number of abstracts outside this bracket, yielding 133,082,240 words. The general corpus was constructed by combining two existing corpora: the Corpus of Historical American English [CoHA; 30] and the Corpus of Contemporary American English [CoCA; 31]. CoHA contains 400 million words from the 1810s to the early 2000s, drawn from 115,000 texts evenly distributed across several kinds of everyday publication, including fiction, magazines, newspapers, and non-fiction books. CoCA contains 560 million words from 1990 to 2019 drawn from approximately 500,000 texts extracted from spoken language, TV shows, academic journals, fiction, magazines, newspapers, and blogs. To prevent potential overlap with the psychology corpus, we excluded CoCA texts that were sourced from academic journals, as well as excluding blogs due to their missing year data, before merging CoCA texts with CoHA to form a general corpus. This combined corpus has previously been demonstrated to be reliable [13]. We further extracted texts containing the phrase ‘Great Depression’ from both corpora in an effort to restrict usages of ‘depression’ to its psychiatric meaning. The phrase was rare (0.5% of all texts containing ’depression’) in texts in the psychology corpus but common (14.2%) in the combined CoHA/CoCA corpus. Finally, only CoHA/CoCA texts between 1970 and 2018 were extracted to match the psychology corpus time period. In total, “anxiety” and “depression” appeared 47,324 times and 52,010 times, respectively, in the psychology corpus, and 7,959 times and 7,878 times in the CoHA/CoCA corpus (excluding texts with instances of “Great Depression”). Fig 1 presents the relative frequency of these centre terms by year in the two restricted corpora. The two centre terms appear much more frequently in the psychology corpus and become more prevalent in it over the study period. 10.1371/journal.pone.0288027.g001 Fig 1 Relative frequency of “anxiety” and “depression” in the corpora. Warriner norms dataset Affective meaning norms published by Warriner and colleagues [27] were used to evaluate the emotional severity of the contexts in which target words (i.e., ‘centre terms’) appeared. This dataset provides norms for valence, arousal, and dominance ratings of 13,915 English lemmas (i.e., the canonical or dictionary form of a set of word forms) provided by 1,827 United States residents (aged 16 to 87 years; 60% female). Participants rated how they felt while reading a word on a series of scales ranging from 1 (low) to 9 (high). For the valence rating (n = 723, M = 5.1, SD = 1.7), 1 corresponded to feeling extremely "annoyed", "bored", "despaired", "melancholic", "unhappy", or, "unsatisfied", and 9 corresponded to feeling extremely "contented", "happy", "hopeful", "pleased", or "satisfied". For the arousal rating (n = 745, M = 4.2, SD = 2.3), 1 represented feeling "calm", "dull", "relaxed", "sleepy", "sluggish", or "unaroused", while 9 indicated feeling "agitated", "aroused", "excited", "frenzied", "jittery", "stimulated", or "wide-awake". The dominance ratings were not used. Measures Contexts of the centre terms The study evaluated semantic changes in the centre terms ‘anxiety’ and ‘depression’ by examining shifts in the emotional severity of collocated words occurring in their immediate context within the corpora. Following established practice [32], and consistent with Baes et al. [15], collocates were defined as individual words occurring within a ±5-word context window of each centre term. Before extracting the collocates, the text corpora were pre-processed by converting all words into their lemmas to reduce variations of word forms (e.g., “go” for “gone”, “going” and “went”). Including repetitions of specific words, the collocate extraction procedure in the psychology corpus (1970–2018) resulted in 1,030,314 anxiety collocates, and 1,093,926 depression collocates; in the general (CoHA/CoCA) corpus (1970–2018), there were 107,611 anxiety collocates and 118,542 depression collocates. Severity index To compute an index of emotional severity, through which annual changes in the mean severity of ‘anxiety’ and ‘depression’ could be evaluated, we followed the procedure developed by Baes et al. [15]. For each corpus, collocates of the centre terms were matched with the Warriner norms dataset, disregarding collocates absent from those norms. Valence and arousal ratings were then summed to generate an index of emotional severity for each collocate, ranging from 2 to 18. Valence ratings were reverse scored (i.e., 1 = happy, 9 = unhappy) while arousal ratings were not (i.e., 1 = calm, 9 = aroused). The summed index assigns low scores for words judged to be emotionally positive and calm, and high scores for those judged to be unpleasant and intense. The severity index was then computed for each corpus by taking the weighted average collocate severity for each year (i.e., weighted by number of repetitions for collocates appearing more than once in the year). This index therefore represents the mean emotional intensity of words collocated with ‘anxiety’ and ‘depression’ in a particular year and would be expected to decline from 1970 to 2018 according to the study hypothesis. Analytic strategy Linear regression was performed to test the hypothesis that the severity index would decline. To investigate patterns of semantic change in greater detail, these analyses were followed up by identifying the most common collocates of ‘anxiety’ and ‘depression’ for each of the five decades, beginning with the 1970s, given the low frequency of most collocates in particular years. By examining changes in the relative frequencies of specific collocates across the decades, drivers of any trends in the severity index could be ascertained. All figures and statistical analyses were produced and processed using R in RStudio [33]. See the Open Science Framework repository for R scripts and associated files used in the current study: https://osf.io/u9mve/?view_only=f877aa015c78454fb62d898e51b47158. Results Psychology corpus Anxiety There were 826,083 anxiety collocates (including repetitions of specific collocates) that matched words in the Warriner norms, representing 80.2% of the anxiety collocates. Their severity scores ranged from 3.4 to 14.8 on the 2 to 18 scale (M = 7.9, SD = 1.7). A linear model with the severity index as the outcome variable and year as the predictor was statistically significant, F(1, 47) = 132.10, p < .001, accounting for 73% of the variance in the severity index. Contrary to hypothesis, the severity index increased over time (see Fig 2), indicating a significant rise in the severity of words occurring in the vicinity of ‘anxiety’, t(47) = 11.49, p < .001, β = 0.86, 95% CI [0.71, 1.01]. 10.1371/journal.pone.0288027.g002 Fig 2 The severity index for “anxiety” in the psychology abstracts corpus from 1970 to 2018. The grey bars around the linear regression line indicate the standard error estimate. The top 10 most frequent anxiety collocates for each decade in the psychology corpus are displayed in Table 1. The words ‘depression’, ‘disorder’, and ‘symptom’ had the highest severity ratings among these collocates, scoring 11.8, 10.1, and 9.8, respectively, all >1 SD above the mean severity scores of the anxiety collocates. These three words were among the most frequent collocates of ‘anxiety’ in every decade except the 1970s and therefore contributed substantially to the rise of the severity index from the 1970s to the 2010s. The relative frequencies of ‘depression’, ‘disorder’, and ‘symptom’ by decade are presented in Fig 3. All illustrate a statistically significant rising trend; for ‘depression’: β = 0.96, 95% CI [0.52, 1.42], p = .006; for ‘disorder’: β = 0.93, 95% CI [0.28, 1.59], p = .020; for ‘symptom’: β = 0.99, 95% CI [0.72, 1.25], p = .001. This suggests that language surrounding ‘anxiety’ in the psychology abstracts became increasingly focused on anxiety’s clinical and pathological aspects. 10.1371/journal.pone.0288027.g003 Fig 3 Relative frequencies of selected anxiety collocates in the psychology abstracts corpus by decade. Relative frequency is the summed repetitions of one lemma within a decade divided by the summed repetitions of all lemmas in the same decade. Larger relative frequency means higher frequency of the lemma in a particular decade. 10.1371/journal.pone.0288027.t001 Table 1 Top 10 anxiety collocates in the psychology abstracts corpus by decade. 1970s 1980s 1990s 2000s 2010s test depression depression depression depression state trait disorder disorder disorder measure state high symptom symptom trait measure trait social social scale high state high high level test symptom child study high self measure measure child group level self level associate score scale patient study level subject report social trait report Depression There were 854,235 Warriner-matched depression collocates in the psychology corpus, 78.1% of all the depression collocates, whose mean severity was 7.9 (SD = 1.7). The linear model was statistically significant, F(1, 47) = 9.47, p = .003, accounting for 15% of the variance in the severity index. Again, contrary to the hypothesis the severity index rose (see Fig 4), rather than fell over time, t(47) = 3.08, p = .003, β = 0.41, 95% CI [0.14, 0.68]. 10.1371/journal.pone.0288027.g004 Fig 4 The severity index for “depression” in the psychology abstracts corpus from 1970 to 2018. The grey bars around the linear regression line indicate the standard error estimate. The top 10 most frequent depression collocates for each decade are listed in Table 2. Collocates with the highest severity ratings were ‘anxiety’, ‘disorder’, and ‘symptom’, scoring 11.4, 10.1, and 9.8, respectively, all well above the mean for all collocates. ‘Anxiety’ and ‘symptom’ were among the most frequent collocates of ‘depression’ in every decade, and ‘disorder’ was frequent in three out of five decades. The relative frequencies of these common and influential collocates are presented by decade in Fig 5. ‘Anxiety’ and ‘symptom’ demonstrate a statistically significant increasing trend over time; for ‘anxiety’: β = 1.00, 95% CI [0.83, 1.15], p < .001; for ‘symptom’: β = 0.94, 95% CI [0.34, 1.55], p = .015; while ‘disorder’ is stable, β = 0.79, 95% CI [-0.35, 1.92], p = .114. 10.1371/journal.pone.0288027.g005 Fig 5 Relative frequencies of selected depression collocates in the psychology abstracts corpus by decade. Relative frequency is the summed repetitions of one lemma within a decade divided by the summed repetitions of all lemmas in the same decade. Larger relative frequency means higher frequency of the lemma in a particular decade. 10.1371/journal.pone.0288027.t002 Table 2 Top 10 depression collocates in psychology abstracts corpus by decade. 1970s 1980s 1990s 2000s 2010s patient anxiety anxiety anxiety anxiety scale patient patient symptom symptom anxiety scale major patient study measure self scale scale associate group measure self study patient symptom inventory symptom major scale self score disorder disorder use control study score associate disorder study symptom measure treatment treatment result major study high high Words were ranked by their relative frequency in each decade, from highest (top row) to lowest. General corpus Anxiety There were 87,163 Warriner-matched anxiety collocates in the general (CoHA/CoCA) corpus, 73.5% of all the anxiety collocates, with an average severity score of 7.9 (SD = 1.9). The linear model was statistically significant F(1, 47) = 7.97, p = .007, accounting for 13% of the variance in the severity index. Contrary to hypothesis, the severity index rose over time, t(47) = 2.82, p = .007, β = 0.38, 95% CI [0.11, 0.65], as Fig 6 shows. 10.1371/journal.pone.0288027.g006 Fig 6 The severity index for “anxiety” in the CoHA/CoCA corpus from 1970 to 2018. The grey bars around the linear regression line indicate the standard error estimate. The top 10 most frequent anxiety collocates for each decade are shown in Table 3. ‘Fear’, ‘depression’, and ‘disorder’ had the highest severity ratings among these collocates, scoring 12.2, 11.8, and 10.1, respectively. They were identified as among the top 10 most frequent collocates in at least two decades, therefore strongly influencing the overall severity index. Relative frequencies of ‘fear’, ‘depression’, and ‘disorder’ across the decades are plotted in Fig 7. The relative frequencies of ‘depression’ and ‘disorder’ both rose steeply with statistical significance; for ‘depression’: β = 0.98, 95% CI [0.58, 1.38], p = .004; for ‘disorder’: β = 0.94, 95% CI [0.31, 1.58], p = .018; whereas ‘fear’ remained relatively stable, β = 0.61, 95% CI [-0.84, 2.07], p = .274. 10.1371/journal.pone.0288027.g007 Fig 7 Relative frequencies of selected anxiety collocates in the CoHA/CoCA corpus by decade. Relative frequency is the summed repetitions of one lemma within a decade divided by the summed repetitions of all lemmas in the same decade. Larger relative frequency means higher frequency of the lemma in a particular decade. 10.1371/journal.pone.0288027.t003 Table 3 Top 10 anxiety collocates in the CoHA/CoCA corpus by decade. 1970s 1980s 1990s 2000s 2010s feel feel feel depression depression man fear know feel people know time fear disorder know time come people know feel fear face like fear like like great think people think ask know depression time fear face day time think disorder think like come like time come leave attack stress come Words were ranked by their relative frequency in each decade, from highest (top row) to lowest. Depression There were 95,476 Warriner-matched depression collocates in the general corpus, 88.7% of all the depression collocates, with an average severity score of 7.9 (SD = 1.9). The linear model was statistically significant, F(1, 47) = 25.95, p < .001, accounting for 34% of the variance in the severity index. Once again, mean severity increased over time (see Fig 8), contrary to hypothesis, t(47) = 5.09, p < .001, β = 0.60, 95% CI [0.36, 0.83]. 10.1371/journal.pone.0288027.g008 Fig 8 The severity index for “depression” in the CoHA/CoCA corpus from 1970 to 2018. The grey bars around the linear regression line indicate the standard error estimate. The top 10 most frequent depression collocates by decade are presented in Table 4. There were several relatively frequent high severity collocates, notably ‘war’, ‘suffer’, ‘problem’, ‘anxiety’, and ‘disorder’. Fig 9 shows that there was no consistent temporal trend for the first two terms; for ‘war’: β = -0.72, 95% CI [-2.00, 0.56], p = .172; for ‘problem’: β = 0.68, 95% CI [-0.66, 2.03], p = .204; whereas ‘suffer’ shows a statistically significant increase, β = 0.94, 95% CI [0.31, 1.57], p = .018. Fig 10 shows a statistically significant rise for the clinical terms ‘anxiety’ (β = 0.94, 95% CI [0.32, 1.56], p = .017) and ‘disorder’ (β = 0.97, 95% CI [0.53, 1.41], p = .006). As in earlier analyses, this pattern indicates an increase in the use of clinical or pathological language in the vicinity of ‘depression’ within the general corpus. 10.1371/journal.pone.0288027.g009 Fig 9 Relative frequencies of selected depression collocates in the CoHA/CoCA corpus by decade. Relative frequency is the summed repetitions of one lemma within a decade divided by the summed repetitions of all lemmas in the same decade. Larger relative frequency means higher frequency of the lemma in a particular decade. 10.1371/journal.pone.0288027.g010 Fig 10 Relative frequencies of selected depression collocates in the CoHA/CoCA corpus by decade. Relative frequency is the summed repetitions of one lemma within a decade divided by the summed repetitions of all lemmas in the same decade. Larger relative frequency means higher frequency of the lemma in a particular decade. 10.1371/journal.pone.0288027.t004 Table 4 Top 10 depression collocates in the CoHA/CoCA corpus by decade. 1970s 1980s 1990s 2000s 2010s time year people people anxiety year people year anxiety know war day know know people day time time year like know like suffer disorder year think woman like like suffer feel think think suffer disorder find help anxiety time time like suffer problem think life life find war problem thing Words were ranked by their relative frequency in each decade, from highest (top row) to lowest. Discussion The present study was the first to systematically examine long-term historical shifts in the meaning and use of ‘depression’ and ‘anxiety’. From the theoretical standpoint of concept creep and using newly developed methods for evaluating changes in the emotional severity of word meanings, the study revealed consistent trends for the two concepts of interest across large text corpora representing academic-professional and general language use. These trends represent linear increases in the emotional severity associated with ‘anxiety’ and ‘depression’ over the past half century. These strong and consistent trends run contrary to the hypothesized direction of change. Based on well-established concerns that some mental health-related concepts have undergone semantic dilution in recent decades, reflected in diagnostic inflation within the mental health professions [6, 10] and colloquial use of clinical terms to reference everyday emotional states [23], we predicted ‘anxiety’ and ‘depression’ would decline in severity. If these concepts had crept vertically, their semantic context, represented by the words collocated with them, should have trended to become less severe over time, as Baes et al. [15] found for ‘trauma’ using an identical methodology. Instead, the average severity of collocates rose from the 1970s to the 2010s. The fact that this rise replicated across concepts and across very different corpora suggests that it is a robust effect rather than one confined to professional or general discourse. Although our primary hypothesis was not supported, follow-up analyses offer some clues to what may have driven the observed rise in severity. Repeatedly, we found that many of the most common collocates of ‘anxiety’ and ‘depression’ were clinical terms and that these clinical collocates became more prevalent over the study period. In particular, the terms ‘disorder’ and ‘symptom’ tended to become more associated with ‘anxiety’ and ‘depression’ in more recent decades. In the psychology corpus, for example, neither term featured in the top 10 collocates in the 1970s and 1980s but ‘disorder’ became the second most common collocate from the 1990s through the 2010s and ‘symptom’ rose to third rank in the 2000s and 2010s. These patterns were almost equally striking in the general corpus. Although the collocates tended to be less clinical overall than in the psychology corpus and ‘symptom’ never featured in the top 10 collocates for either ‘anxiety’ or ‘depression’, ‘disorder’ became a popular collocate in the 2000s and 2010s. By implication, in both the academic and professional discourse of psychology and in the wider culture sampled by the general corpus, ‘anxiety’ and ‘depression’ were increasingly discussed in the context of pathology. That increase appears to have been gradual and lagged between professional and general discourse, and therefore cannot be confidently ascribed to a single historical event, such as the publication of DSM-III in 1980. Whether the rise in pathological language around ‘anxiety’ and ‘depression’ might be associated with historical increases in the prevalence of these conditions (e.g., [34]) remains to be determined. There may or may not be a relationship between the prevalence of a clinical phenomenon and the conditional probability of clinical terminology being used when it is mentioned. Even more striking is the rising trend for the two concepts of interest to co-occur. In the psychology corpus, ‘depression’ became and remained the most common collocate of ‘anxiety’ in the 1980s, after not even entering the top 10 in the 1970s, and ‘anxiety’ became, and remained, the most common collocate for ‘depression’ in the same decade. This co-occurrence mirrors the substantial comorbidity and overlap of anxiety and depressive disorders [16, 18]. A similar convergence was evident in the general corpus, albeit appearing later. ‘Depression’ became the top collocate of ‘anxiety’ in the 2000s and ‘anxiety’ became its top collocate in the 2010s. The two concepts have clearly become a tightly bound pair in both the academic and general discourse. Because the two terms are high in emotional severity (i.e., affectively negative and intense), as are clinical terms like ‘disorder’ and ‘symptom’, at least part of the increase in the severity index over the study period may reflect the rising prominence of these collocations. Although our findings do not support the predicted dilution of the meaning of ‘anxiety’ or ‘depression’, they are consistent with an increased pathologizing of these concepts over recent decades. Whereas ‘anxiety’ and ‘depression’ can refer to ordinary affective states, rather than to clinical conditions, and–judging from the collocates–largely did so in the 1970s and 1980s, the strong trend in both psychological and general discourse has been to place a clinical frame around them. That frame locates them in the context of diagnosis (‘disorder’ and ‘symptom’) rather than normal emotional distress and compounds them as linked pathological entities rather than as distinct experiences. This pathologizing trend appears to be the direct opposite of the normalizing trend (vertical concept creep) we predicted, but it may not be incompatible with it. It is possible that ‘anxiety’ and ‘depression’ are now being used to refer to less severe phenomena than in earlier times, but they are also used in a more clinical idiom than before. A tendency to use pathological or diagnostic language to make sense of everyday distress might be one way in which vertical creep takes place. Even so, the trends we have identified are better described as pathologizing rather than as vertical concept creep (normalizing). The present research inevitably has limitations and weaknesses. Despite the breadth and size of the two text corpora, it is possible that unrelated historical changes in their composition might distort the severity trends we examined in our hypothesis tests (e.g., changes in the proportion of clinical psychology articles in the psychology corpus). It would also be inappropriate to conclude from the patterns observed in the general corpus how everyday people think about ‘anxiety’ and ‘depression’, seeing as the texts that compose the corpus are generated by an unrepresentative group of content producers (e.g., authors, journalists, bloggers). Our severity index, based on published norms of emotional meaning, is a readily automated way to assess the dimension of harm on which vertical concept creep takes place, from intensely negative to mild and innocuous, but it may not fully capture the complexity of harm, which may also have a moral component that is not reducible to affective intensity. It is also possible that some historical shifts have occurred in the connotations of words that might complicate the interpretation of the trends we observed, although we believe it is implausible that there has been a strong generalized tendency for word meanings to have increased in emotional severity over recent decades. Nevertheless, although the methodology has some potential limitations, the magnitude of our data sets and their great historical scope represent some compensating strengths. Future research might refine the methodology, examine additional text corpora including those drawn from social media, to explore media representations of mental illness, and determine whether similar patterns of pathologizing can be observed with other mental health-related concepts. Conclusion Using very large text corpora representing academic psychology and general culture (USA blogs, fiction, magazines, newspapers, spoken language, TV), we found that the concepts of ‘anxiety’ and ‘depression’ have undergone notable shifts in their emotional meaning. Contrary to the hypothesis that they would become increasingly associated with less intense, severe, or harm-related words, the opposite pattern consistently emerged across concepts and corpora. That pattern appeared to reflect, in part, a rising tendency to pathologize ‘anxiety’ and ‘depression’ by locating them in the semantic context of diagnosis, disorder, and symptoms. The cultural implications of this trend, and how it relates to the broad pattern of concept creep, remains to be determined. All corpus pre-processing and data extraction used Spartan, the University of Melbourne’s general purpose hybrid high performance computing system: Lafayette, L., Sauter, G., Vu, L. and Meade, B., 2016. Spartan performance and flexibility: An hpc-cloud chimera. OpenStack Summit, Barcelona, 27. 10.1371/journal.pone.0288027.r001 Decision Letter 0 Ptaszynski Michal Academic Editor © 2023 Michal Ptaszynski 2023 Michal Ptaszynski https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Submission Version0 16 May 2023 PONE-D-23-09235Have the concepts of ‘anxiety’ and ‘depression’ been normalized or pathologized? A corpus study of historical semantic changePLOS ONE Dear Dr. Baes, Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. Please submit your revised manuscript by Jun 30 2023 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file. Please include the following items when submitting your revised manuscript:A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'. An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'. If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter. If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols. We look forward to receiving your revised manuscript. Kind regards, Michal Ptaszynski, PhD Academic Editor PLOS ONE Journal requirements: When submitting your revision, we need you to address these additional requirements. 1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions Comments to the Author 1. Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. Reviewer #1: Yes Reviewer #2: Partly ********** 2. Has the statistical analysis been performed appropriately and rigorously? Reviewer #1: Yes Reviewer #2: No ********** 3. Have the authors made all data underlying the findings in their manuscript fully available? The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified. Reviewer #1: Yes Reviewer #2: Yes ********** 4. Is the manuscript presented in an intelligible fashion and written in standard English? PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here. Reviewer #1: Yes Reviewer #2: Yes ********** 5. Review Comments to the Author Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters) Reviewer #1: Thank you for the opportunity to review this interesting paper for PLoS One. The authors examined the shift in meaning of the words “anxiety” and “depression” by noting the frequency of words with which they were paired in two distinct but large corpora—one academic and one public. Hypothesizing that they would find a more casual use of the words developing from the 1970’s through the 2010’s, they were surprised to find the opposite: more collocation with words of a clinical nature, including with each other. There are just a few minor comments to consider before the paper is ready for publication in this Journal. As a disclaimer, language studies like this are not my area of expertise, but I give my comments as a generally interested clinical psychologist who does research on depression and anxiety. Intro Line 42-43: the bullying example is very illustrative and helpful for someone new to this type of study. Bravo. Line 49-50: Implied is that vertical concept creep can only be “downward,” i.e. moving to include less severe phenomena as in the bullying example. Couldn’t vertical concept creep also include moving “upward” to include more severe phenomena, as the authors sort of found in the present study? Lines 60-64: This argument is made with no citations. It would be stronger if there were some studies that at least hinted at some empirical evidence for the logic presented actually playing out. Line 113-15: Why is it bad if people seek therapy, even if their “depression” is mild? In truth, it could be a public health strain on services etc., or people could end up on antidepressants with side effects that are unnecessary, but without explicating those implications the argument here sounds pretty weak. Methods Lines 181-83: I was confused as to why valence was summed with arousal. Wouldn’t it make more sense to multiply the two? That way a very positive emotion, like overwhelming joy, that causes a lot of arousal would be scored at the opposite end of the spectrum as, say, overwhelming fear. For ease of interpretability one could also make the valence numbers range from -4 to 4, rather than 1 to 9. Just a suggestion. Results No comments. Discussion One partial explanation for the rise in collocation of clinical words for anxiety and depression could be the rise in the actual phenomena of bona fide, DSM-5 depression and anxiety disorders. The authors did not discuss this possibility, but should. It is likely that the authors are correct that non-pathological, every day phenomena are referred to using clinical language more so today than in the 70’s or 80’s, but there are probably also more true clinical phenomena now than in the 70’s or 80’s as well. This should be considered. Lines 326-9: A citation or two would be helpful to illustrate these “well established” concerns. Line 348-51: While I agree with the gradual increase in anxiety’s and depression’s collocation with clinical terms when viewed decade by decade, the increase in severity index scores of the collocated words (figure 1 especially, for anxiety) suggests that the only real shift occurred from 1985-1993 or so. Do the authors have any explanation for this? Line 370: Should probably read, “This pathologizing trend appears to be…” just for style purposes. I sign this review, as is my practice, that my identity will be included with it as long that is acceptable to the journal. Daniel Norton Gordon College Reviewer #2: Review of PONE-D-23-09235 This manuscript describes an analysis of collocates of the terms "anxiety" and "depression" in historical corpora to test the hypothesis that the meaning of these concepts have broadened over time. This is an interesting study, but I have several questions and concerns for the authors to clarify. 1. The terms "emotional intensity" and "semantic severity" seem to be used interchangeably, but I am not sure if they truly reflect the same concept. 2. An key assumption that the authors' methodology relies on is that the valence and arousal of words (as measured by Warriner et al.) do not change over time. Is it reasonable to assume that the valence and arousal value of words are static and constant over such a long time period given that many studies find that word meanings do evolve and change over history? E.g., the concept "gay" has a more positive connotation earlier in history, but one could imagine that it currently has a more negative connotation in modern times. 3. Why was horizontal concept creep not explored in this paper? This was surprisingly after such a detailed description of the two types of concept creep. 4. Many important details about the methodology, materials are not provided. For example: - In the psychology corpus, how are the journals distributed in terms of subfield of psychology? Are these mostly clinical psychology journals? What about psychiatry journals which would have a large coverage of research in depression in anxiety? - How evenly distributed are the corpora in terms of the date range of 1970 to 2018? Do different years provide roughly the same number of words or is it skewed in some years? How frequently do the terms "anxiety" and "depression" occur in the corpus and how is this similar or different across the time period assessed? - There is no information (as far as I can tell) about how many of the collocates *do not* have any valence or arousal norms in the Warriner dataset. What is the level of missingness and would this have any effect on the results? 5. The severity index is an interesting measure but it is simply a summation of the valence and arousal scores of the word and some information loss may occur. For instance, it is possible for two words to have the same summed index score, but for one word this value is driven by negative valence, and for the other word this value is mostly driven by high arousal. How can such an index distinguish the relative contributions of valence and arousal to the computation of "severity"? 6. Some questions about the analysis approach for the authors to consider: - First of all, it would be better if the figures were provided in-text and not on the OSF only. Based on those figures, it seemed like a linear model may not be appropriate given the clearly non-linear trends in the data. Perhaps a Generalized Additive Model (GAM) is more appropriate for capturing non-linearities in the data. - It is not clear how the predictor "year" is represented in the model? Was it mean-centered, contrast coded in some way or just included as a continuous variable (which is not recommended)? The authors may want to look into techniques that can analyze time-series data. 7. Finally, the fact that "anxiety" and "depression" are themselves highly frequent collocates of each other really makes it quite challenging to understand and interpret the results. I find it surprising that these concepts are not "high in emotional severity" (p. 20)... What is the severity for these words? Perhaps a more detailed frequency analysis would be useful to help us understand what the results mean - for instance, could the collocates be analyzed for their /relative change/ in frequency over time to explore what concepts are main drivers of the unexpected rise in emotional severity? ********** 6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy. Reviewer #1: Yes: Daniel J Norton, Ph.D. Reviewer #2: No ********** [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.] While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step. 10.1371/journal.pone.0288027.r002 Author response to Decision Letter 0 Submission Version1 1 Jun 2023 Response to Reviewers We thank the reviewers for their very helpful feedback. We have seriously engaged with all of their concerns and believe that the manuscript has been significantly improved as a result. Please find our responses below in response to each critical comment. We quote each comment in full, excluding remarks that did not request clarification or expressed agreement or positive evaluation of the manuscript. Reviewer 1 1. “Line 49-50: Implied is that vertical concept creep can only be “downward,” i.e. moving to include less severe phenomena as in the bullying example. Couldn’t vertical concept creep also include moving “upward” to include more severe phenomena, as the authors sort of found in the present study?” We now clarify this point with an added sentence which agrees that concepts can indeed shift in the other direction, but notes that such a shift does not constitute concept creep in our sense because by definition concept creep involves an expansion rather than a contraction of concept meaning. Generally, also, the harm-related concepts studied in research on concept creep cannot easily expand vertically because there is an existing severity maximum that does not change (e.g., for ‘trauma’ it is the most traumatic experience possible). The threshold for deciding when the concept applies might move upward (i.e., a more stringent definition of trauma’) but that narrows the meaning of trauma because the maximum does not also rise. 2. “Lines 60-64: This argument is made with no citations. It would be stronger if there were some studies that at least hinted at some empirical evidence for the logic presented actually playing out”. We have now added three citations to back up our argument in lines 60-64. Dakin et al. have empirically documented the trivialization effect, Jones & McNally have shown that broadened trauma concepts can increase negative responses to disturbing events, and Furedi presents qualitative evidence for the rise of vulnerability associated with rising sensitivity to harm. 3. “Line 113-15: Why is it bad if people seek therapy, even if their “depression” is mild? In truth, it could be a public health strain on services etc., or people could end up on antidepressants with side effects that are unnecessary, but without explicating those implications the argument here sounds pretty weak.” We did not mean to suggest that it was inappropriate for people with mild depression to seek treatment, and regret that implication. Our point is simply that if many people were to self-overdiagnose (e.g., seek or demand treatment when they are not clinical depressed), this might in some cases lead to adverse consequences. We have added a sentence that includes two of the consequences the reviewer suggests, along with another that is backed by a new citation (see lines 118-121). 4. “Lines 181-83: I was confused as to why valence was summed with arousal. Wouldn’t it make more sense to multiply the two? That way a very positive emotion, like overwhelming joy, that causes a lot of arousal would be scored at the opposite end of the spectrum as, say, overwhelming fear. For ease of interpretability one could also make the valence numbers range from -4 to 4, rather than 1 to 9. Just a suggestion” We believe summing is the appropriate way of combining the two indices, and it has been carried out in a previous published study using the method. Multiplying the valence and arousal indices would create a very skewed severity measure if we do not re-scale the indices (range 1 to 81 where neutrality is 25) and multiplying the re-scaled (-4 to +4) indices would be inappropriate because a low intensity positive word would get the same score as a high intensity negative one. We believe summing is a simple method that captures emotional severity well by scoring low intensity positive words (e.g., “contented”) lowest, high intensity negative words highest, and low arousal negative and high arousal positive words intermediate. Regarding possibly re-scaling the valence and arousal indices, we would prefer to retain the existing 109 scale, which is based directly on the published norms themselves. 5. “One partial explanation for the rise in collocation of clinical words for anxiety and depression could be the rise in the actual phenomena of bona fide, DSM-5 depression and anxiety disorders. The authors did not discuss this possibility, but should. It is likely that the authors are correct that non-pathological, every day phenomena are referred to using clinical language more so today than in the 70’s or 80’s, but there are probably also more true clinical phenomena now than in the 70’s or 80’s as well. This should be considered.” We agree that there may have been some historical increase in the prevalence of anxiety and depressive disorders over the decades of our study. However, we believe any such rise in prevalence would more likely manifest in a rise in the frequency with which “anxiety” and “depression” appear in our text corpora rather than in the frequency of clinically-related versus other collocates when these words do appear. The fact that there is more clinical anxiety present now than in 1970 should not obviously change the kinds of language that is used around the word “anxiety” now versus then. Clinical anxiety (and probably subclinical anxiety too) may have been less prevalent and talked about then, but it is not clear to us that it was talked about clinical terminology should be less common just because the objective prevalence was lower. Nevertheless, we agree that this is a possibility and have added two sentences to this effect at the end of the third paragraph of the Discussion. Interestingly, our new Figure 1 (see below) shows a substantial rise in the frequency of “anxiety” and “depression” in the psychology corpus. 6. “Lines 326-9: A citation or two would be helpful to illustrate these “well established” concerns.” We have added three relevant citations here, as requested. 7. “Line 348-51: While I agree with the gradual increase in anxiety’s and depression’s collocation with clinical terms when viewed decade by decade, the increase in severity index scores of the collocated words (figure 1 especially, for anxiety) suggests that the only real shift occurred from 1985-1993 or so. Do the authors have any explanation for this?” We do not have a post hoc explanation for the possible difference in the timing of these trends, but note that changes in the overall severity index are driven by a large and diverse assortment of collocates and the specific clinical terms mentioned in the paragraph are likely to be only one modest component of any such changes. We do not claim in the manuscript that shifts in the relative prominence of “disorder” and “symptom” as collocates of “anxiety” and “depression” are primarily responsible for the rising trajectory of the severity indices. 8. “Line 370: Should probably read, “This pathologizing trend appears to be…” just for style purposes.” We have made this change as requested. Reviewer 2 1. “The terms "emotional intensity" and "semantic severity" seem to be used interchangeably, but I am not sure if they truly reflect the same concept.” We were intending to use “emotional intensity” (used 3 times), and “emotional severity” (7 times) and “semantic severity” (9 times) to express the same concept, which our index aims to capture. In essence, we are assessing the extent to which words have meanings that are emotionally negative and high in arousal. “Emotional intensity” is not an entirely satisfactory term because it could in principle refer to positive emotion as well as negative. “Semantic severity” is not perfect because it doesn’t specify the dimension(s) on which severity is being assessed, although severity is a key aspect of the concept. In response to the reviewer’s concern, we therefore now consistently use “emotional severity” throughout the manuscript and refer to the index itself simply as the “severity index”. 2. “A key assumption that the authors' methodology relies on is that the valence and arousal of words (as measured by Warriner et al.) do not change over time. Is it reasonable to assume that the valence and arousal value of words are static and constant over such a long time period given that many studies find that word meanings do evolve and change over history? E.g., the concept "gay" has a more positive connotation earlier in history, but one could imagine that it currently has a more negative connotation in modern times.” We do not believe that our methodology relies on the valence and arousal of words being unchanging. No doubt individual words are subject to some shifts in valence- and arousal-related connotations over periods of decades, although we suspect major changes are likely to be rare. We note that all of our severity trend findings are based on (weighted) average mean severity scores of thousands of unique words, so these trends could only be invalid if changes in the connotations of words are consistently and strongly occurring in one direction. That is, our findings could only occur due to changes in word connotations if words in general are becoming markedly more negative and high arousal over time, a possibility we find implausible. Nevertheless, we have added the following sentence to the penultimate paragraph of the Discussion to acknowledge this issue: “It is also possible that some historical shifts have occurred in the connotations of words that might complicate the interpretation of the trends we observed, although we believe it is implausible that there has been a strong generalized tendency for word meanings to have increased in emotional severity over recent decades.” 3. “Why was horizontal concept creep not explored in this paper? This was surprisingly after such a detailed description of the two types of concept creep.” Horizontal creep was not explored in the current manuscript for two reasons. First, there have been numerous studies of horizontal creep of a range of harm-related concepts using computational linguistic methods (see below for references), but only one previous study of vertical creep of a single concept using the current method. Second, the issue of pathologization that we focus on in the manuscript is primarily to do with vertical rather than horizontal creep, specifically the encroachment of pathology-related language on “normal” (less severe) phenomena. We believe a focused investigation on vertical creep is appropriate for these reasons. Haslam, N., Vylomova, E., Zyphur, M., & Kashima, Y. (2021). The cultural dynamics of concept creep. American Psychologist, 76(6), 1013–1026. Vylomova, E., & Haslam, N. (2021). Semantic changes in harm-related concepts in English. In N. Tahmasebi, L. Borin, A. Jatowt, Y. Xu & S. Hengchen (Eds.), Computational approaches to semantic change (pp. 93-121). Language Science Press. Vylomova, E., Murphy, S., & Haslam, N. (2019). Evaluation of semantic change of harm-related concepts in psychology. In Proceedings of the 1st International Workshop on Computational Approaches to Historical Language Change, 29–34. 4. “Many important details about the methodology, materials are not provided. For example: - In the psychology corpus, how are the journals distributed in terms of subfield of psychology? Are these mostly clinical psychology journals? What about psychiatry journals which would have a large coverage of research in depression in anxiety? - How evenly distributed are the corpora in terms of the date range of 1970 to 2018? Do different years provide roughly the same number of words or is it skewed in some years? How frequently do the terms "anxiety" and "depression" occur in the corpus and how is this similar or different across the time period assessed? - There is no information (as far as I can tell) about how many of the collocates *do not* have any valence or arousal norms in the Warriner dataset. What is the level of missingness and would this have any effect on the results?” a) The 875 journals in the psychology corpus are not mostly clinical psychology journals. These make up 171 (19.5%) of the journal set, which is distributed across all subfields of psychology. The 875 journals were all those listed under “psychology” in the E-Research and PubMed databases and therefore are broad in scope. We have added a short description of the main groupings of journals according to their Scimago classification in the Materials section of the Method. Psychiatry journals were not picked up using this search process unless also tagged as “psychology” in the relevant databases. b) The psychology corpus contains substantially fewer abstracts in the early years than in later ones, reflecting the explosion of psychology publication over the last half century. As noted in the Method, we began the study period in 1970 because prior to then the data were sparse. The general corpus has a much more even distribution of words. We emphasize that the size of the corpus in a particular year should have no systematic relationship with the mean severity score for that year, merely increasing the error around the mean (e.g., observe the greater jaggedness of the plots in the earlier years for the psychology corpus). We have now added a graph (Figure 1) presenting the relative frequency of “anxiety” and “depression” in the two corpora. c) The Warriner norms cover almost 14,000 common English lemmas. Any material in the corpora that did not match to one of these normed lemmas is likely to reflect very infrequent words. Although it is impossible to determine whether these “missing” (i.e., unmatched) lemmas would alter the severity trends because there is no way to evaluate their severity, we believe it is implausible because a large majority of the collocates are non-missing lemmas, representing the most common words in English and in the corpora. We have now added to the results section the proportion of each set of collocates that matched to the Warriner norms (i.e., that were non-missing), and these show that more than 70% of the collocates matched the norms. We thank the reviewer for raising this issue and believe these new data add confidence to our findings. 5. “The severity index is an interesting measure but it is simply a summation of the valence and arousal scores of the word and some information loss may occur. For instance, it is possible for two words to have the same summed index score, but for one word this value is driven by negative valence, and for the other word this value is mostly driven by high arousal. How can such an index distinguish the relative contributions of valence and arousal to the computation of "severity"? The index cannot distinguish the relative contribution of its two components, but that contribution should tend to be approximately equal given that two variables measured on the same scale with roughly equal variability are being summed. Our goal in creating the summed index was to capture the concept of emotional severity, understood as involving connotations of high arousal and negative valence, not to examine each component separately. We aimed to assess collocates on a continuum from “emotionally positive and calm” to “unpleasant and intense”, and the fact that some intermediate scoring collocates might be positive and high arousal (e.g., excited), negative and low arousal (bored), or neutral and average arousal is therefore not problematic. This issue of possible heterogeneity among midrange cases arises for all indices that sum imperfectly correlated variables. We therefore believe it is legitimate to use our summed index for its intended purpose. 6. “Some questions about the analysis approach for the authors to consider: - First of all, it would be better if the figures were provided in-text and not on the OSF only. Based on those figures, it seemed like a linear model may not be appropriate given the clearly non-linear trends in the data. Perhaps a Generalized Additive Model (GAM) is more appropriate for capturing non-linearities in the data. - It is not clear how the predictor "year" is represented in the model? Was it mean-centered, contrast coded in some way or just included as a continuous variable (which is not recommended)? The authors may want to look into techniques that can analyze time-series data.” a) It was our intention for the figures to be displayed in the manuscript. Please accept our apologies, as it seems the editorial system did not display the TIFF files well in the pdf version. We have attempted to rectify this. Having said that, we believe the trend graphs are generally reasonably close to linear (except Figure 2). As the hypotheses we tested in the analyses were simple ones – historical declines in severity index – and our goal was not to model the trends in greater detail, we believe it is appropriate to run the simplest analyses. Despite some possible nonlinearities, we believe the rising trends are unmistakeable in every figure. b) For the same reason, we used year as a continuous variable in the simple regression (essentially a correlation) and believe this is appropriate for the simple analytic purpose of testing for a historical rise or fall. We agree that time series analysis would be appropriate if we were carrying out more complex analyses involving predictors of the trend, exploration of endogenous factors/autocorrelation underlying the trend, forecasting the trend, and so on. 7. “Finally, the fact that "anxiety" and "depression" are themselves highly frequent collocates of each other really makes it quite challenging to understand and interpret the results. I find it surprising that these concepts are not "high in emotional severity" (p. 20)... What is the severity for these words? Perhaps a more detailed frequency analysis would be useful to help us understand what the results mean - for instance, could the collocates be analyzed for their /relative change/ in frequency over time to explore what concepts are main drivers of the unexpected rise in emotional severity?” In fact, “anxiety” and “depression” are both high in emotional severity, rated 11.4 and 11.8 respectively compared to a mean for all collocates of 7.9 (these figures are presented in the Results). The sentence in question was badly punctuated: instead of “Because the two terms are not surprisingly high in emotional severity” it should have read “Because the two terms are, not surprisingly, high in emotional severity”. We apologize for this error and have corrected it. In view of this, we believe our figures showing historical rises in the frequency of several high severity collocates helps to explain the unexpected rises we observed. Again, we thank both reviewers for their very careful attention to our manuscript and believe that our responses have made the work stronger. 10.1371/journal.pone.0288027.r003 Decision Letter 1 Ptaszynski Michal Academic Editor © 2023 Michal Ptaszynski 2023 Michal Ptaszynski https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Submission Version1 19 Jun 2023 Have the concepts of ‘anxiety’ and ‘depression’ been normalized or pathologized? A corpus study of historical semantic change PONE-D-23-09235R1 Dear Dr. Baes, We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements. Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication. An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org. If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org. Kind regards, Michal Ptaszynski, PhD Academic Editor PLOS ONE Additional Editor Comments (optional): Reviewers' comments: Reviewer's Responses to Questions Comments to the Author 1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation. Reviewer #1: All comments have been addressed Reviewer #2: All comments have been addressed ********** 2. Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. Reviewer #1: Yes Reviewer #2: (No Response) ********** 3. Has the statistical analysis been performed appropriately and rigorously? Reviewer #1: Yes Reviewer #2: (No Response) ********** 4. Have the authors made all data underlying the findings in their manuscript fully available? The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified. Reviewer #1: Yes Reviewer #2: (No Response) ********** 5. Is the manuscript presented in an intelligible fashion and written in standard English? PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here. Reviewer #1: Yes Reviewer #2: (No Response) ********** 6. Review Comments to the Author Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters) Reviewer #1: The authors did a good job addressing all of my concerns. I recommend the paper be accepted, and appreciate the opportunity to review this interesting paper. Reviewer #2: (No Response) ********** 7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy. Reviewer #1: No Reviewer #2: No ********** 10.1371/journal.pone.0288027.r004 Acceptance letter Ptaszynski Michal Academic Editor © 2023 Michal Ptaszynski 2023 Michal Ptaszynski https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. 21 Jun 2023 PONE-D-23-09235R1 Have the concepts of ‘anxiety’ and ‘depression’ been normalized or pathologized? A corpus study of historical semantic change Dear Dr. Baes: I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now with our production department. If your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information please contact onepress@plos.org. If we can help with anything else, please email us at plosone@plos.org. Thank you for submitting your work to PLOS ONE and supporting open access. Kind regards, PLOS ONE Editorial Office Staff on behalf of Dr. Michal Ptaszynski Academic Editor PLOS ONE ==== Refs References 1 Haslam N. (2016). Concept creep: Psychology’s expanding concepts of harm and pathology. Psychological Inquiry, 27 (1 ), 1–17. doi: 10.1080/1047840X.2016.1082418 2 Haslam N. , Dakin B. , Fabiano F. , McGrath M. , Rhee J. , & Vylomova E. et al . (2020). Harm inflation: Making sense of concept creep. European Review of Social Psychology, 31 (1 ), 254–286. doi: 10.1080/10463283.2020.1796080 3 Dakin B. C. , Rhee J. , McGrath M. J. , & Haslam N. (2023). Broadened concepts of harm appear less serious. Social Psychological and Personality Science, 14 , 72–83. 4 Furedi F. (2004). Therapy culture: Cultivating vulnerability in an uncertain age. Routledge. 5 Jones P. J. , & McNally R. J. (2022). Does broadening one’s concept of trauma undermine resilience? Psychological Trauma: Theory, Research, Practice, and Policy, 14 (S1 ), S131–S139. doi: 10.1037/tra0001063 34197173 6 Frances A. (2013). Saving normal: An insider revolts against out-of-control psychiatric diagnosis, DSM-5, BigPharma, and the medicalization of ordinary life. William Morrow. 7 Horwitz A. V. , & Wakefield J. C. (2007). The loss of sadness: How psychiatry transformed normal sorrow into depressive disorder. Oxford University Press. 8 Beeker T. , Mills C. , Bhugra D. , Te Meerman S. , Thoma S. , Heinze M. , et al . (2021). Psychiatrization of society: A conceptual framework and call for transdisciplinary research. Frontiers in Psychiatry, 12 , 645556. doi: 10.3389/fpsyt.2021.645556 34149474 9 Haslam N. , Tse J. S. , & De Deyne S. (2021). Concept creep and psychiatrization. Frontiers in Sociology, 6 (806147 ). doi: 10.3389/fsoc.2021.806147 34977230 10 Fabiano F. , & Haslam N. (2020). Diagnostic inflation in the DSM: A meta-analysis of changes in the stringency of psychiatric diagnosis from DSM-III to DSM-5. Clinical Psychology Review, 80 , 101889. doi: 10.1016/j.cpr.2020.101889 32736153 11 Kazda L. , Bell K. , Thomas R. , McGeechan K. , Sims R. , & Barratt A. (2021). Overdiagnosis of Attention-Deficit/Hyperactivity Disorder in children and adolescents: A systematic scoping review. JAMA Network Open, 4 (4 ), e215335. doi: 10.1001/jamanetworkopen.2021.5335 33843998 12 Haslam N. , & McGrath M. J. (2020). The creeping concept of trauma. Social Research: An International Quarterly, 87 (3 ), 509–531. doi: 10.1353/sor.2020.0052 13 Haslam N. , Vylomova E. , Zyphur M. , & Kashima Y. (2021). The cultural dynamics of concept creep. American Psychologist, 76 (6 ), 1013–1026. doi: 10.1037/amp0000847 34914436 14 Vylomova E. , & Haslam N. (2021). Semantic changes in harm-related concepts in English. In Tahmasebi N. , Borin L. , Jatowt A. , Xu Y. & Hengchen S. (Eds.), Computational approaches to semantic change (pp. 93–121). Language Science Press. 15 Baes N. , Vylomova E. , Zyphur M. , Haslam N. (2023). The semantic inflation of ‘trauma’ in psychology. Psychology of Language and Communication, 27 , 23–45. doi: 10.58734/plc-2023-0002 16 Dobson K. (1985). The relationship between anxiety and depression. Clinical Psychology Review, 5 (4 ), 307–324. doi: 10.1016/0272-7358(85)90010-8 17 Gorman J. (1996). Comorbid depression and anxiety spectrum disorders. Depression and Anxiety, 4 (4 ), 160–168. doi: 10.1002/(SICI)1520-6394(1996)4:4<160::AID-DA2>3.0.CO;2-J 9166648 18 Wetzler S. , & Katz M. (1989). Problems with the differentiation of anxiety and depression. Journal of Psychiatric Research, 23 (1 ), 1–12. doi: 10.1016/0022-3956(89)90013-7 2666646 19 Wakefield J. , & First M. (2003). Clarifying the distinction between disorder and nondisorder: Confronting the overdiagnosis (false-positives) problem in DSM-V. In Phillips K. A. , First M. B. , & Pincus H. A. (Eds.), Advancing DSM: Dilemmas in psychiatric diagnosis (pp. 23–55). American Psychiatric Association. 20 Horwitz A. V. , & Wakefield J. C. (2012). All we have to fear: Psychiatry’s transformation of natural anxieties into mental disorders. Oxford University Press. 21 Uher R. , Payne J. L. , Pavlova B. , & Perlis R. H. (2013). Major depressive disorder in DSM-5: Implications for clinical practice and research of changes from DSM-IV. Depression and Anxiety, 31 (6 ), 459–471. doi: 10.1002/da.22217 24272961 22 American Psychiatric Association. (2013). Diagnostic and Statistical Manual of Mental Disorders (5th ed.). American Psychiatric Association. 23 Bröer C. , & Besseling B. (2017). Sadness or depression: Making sense of low mood and the medicalization of everyday life. Social Science & Medicine, 183 , 28–36. doi: 10.1016/j.socscimed.2017.04.025 28458072 24 Brown J. S. , Boardman J. , Elliott S. A. , Howay E. , & Morrison J. (2005). Are self-referrers just the worried well? Social Psychiatry and Psychiatric Epidemiology, 40 (5 ), 396–401. doi: 10.1007/s00127-005-0896-z 15902410 25 Katz S. J. , Kessler R. C. , Frank R. G. , Leaf P. , Lin E. , & Edlund M. (1997). The use of outpatient mental health services in the United States and Ontario: the impact of mental morbidity and perceived need for care. American Journal of Public Health, 87 (7 ), 1136–1143. doi: 10.2105/ajph.87.7.1136 9240103 26 Jackson H. J. , & Haslam N. (2022). Ill-defined: Concepts of mental health and illness are becoming broader, looser and more benign. Australasian Psychiatry, 30 , 490–493. doi: 10.1177/10398562221077898 35156400 27 Warriner A. B. , Kuperman V. , & Brysbaert M. (2013). Norms of valence, arousal, and dominance for 13,915 English lemmas. Behavior Research Methods, 45 (4 ), 1191–1207. doi: 10.3758/s13428-012-0314-x 23404613 28 Vylomova, E., Murphy, S., & Haslam, N. (2019). Evaluation of semantic change of harm-related concepts in psychology. In Proceedings of the 1 st International Workshop on Computational Approaches to Historical Language Change, 29–34. 29 Cleveland D. B. , & Cleveland A. D. (2013). Introduction to indexing and abstracting (4th ed.). Libraries Unlimited. 30 Davies M. (2012). Expanding horizons in historical linguistics with the 400-million word Corpus of Historical American English. Corpora, 7 (2 ), 121–157. doi: 10.3366/cor.2012.0024 31 Davies M. (2010). The Corpus of Contemporary American English as the first reliable monitor corpus of English. Literary and Linguistic Computing, 25 (4 ), 447–464. doi: 10.1093/llc/fqq018 32 Gablasova D. , Brezina V. , & McEnery T. (2017). Collocations in corpus-based language learning research: Identifying, comparing, and interpreting the evidence. Language Learning, 67 (S1 ), 155–179. doi: 10.1111/lang.12225 33 R Core Team (2013). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. http://www.R-project.org/ 34 Moreno-Agostino D. , Wu Y. T. , Daskalopoulou C. , Hasan M. T. , Huisman M. , & Prina M. (2021). Global trends in the prevalence and incidence of depression: a systematic review and meta-analysis. Journal of Affective Disorders, 281 , 235–243. doi: 10.1016/j.jad.2020.12.035 33338841