
==== Front
Hum Brain Mapp
Hum Brain Mapp
10.1002/(ISSN)1097-0193
HBM
Human Brain Mapping
1065-9471
1097-0193
John Wiley & Sons, Inc. Hoboken, USA

10.1002/hbm.70023
HBM70023
Research Article
Research Article
Engagement of the speech motor system in challenging speech perception: Activation likelihood estimation meta‐analyses
Perron et al.
Perron Maxime https://orcid.org/0000-0001-7015-5858
1 2 maxime.perron@mail.utoronto.ca

Vuong Veronica https://orcid.org/0000-0002-9915-3511
1 3 4
Grassi Madison W. 1
Imran Ashna https://orcid.org/0009-0005-1043-6371
1
Alain Claude https://orcid.org/0000-0003-4459-1538
1 2 3 4
1 Rotman Research Institute, Baycrest Academy for Research and Education Toronto Ontario Canada
2 Department of Psychology University of Toronto Toronto Ontario Canada
3 Institute of Medical Sciences, Temerty Faculty of Medicine University of Toronto Toronto Ontario Canada
4 Music and Health Science Research Collaboratory, Faculty of Music University of Toronto Toronto Ontario Canada
* Correspondence
Maxime Perron, Rotman Research Institute, Baycrest Academy for Research and Education, 3560 Bathurst Street, Toronto, ON, Canada, M6A 2E1.
Email: maxime.perron@mail.utoronto.ca

13 9 2024
9 2024
45 13 10.1002/hbm.v45.13 e7002320 8 2024
08 5 2024
29 8 2024
© 2024 The Author(s). Human Brain Mapping published by Wiley Periodicals LLC.
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the terms of the http://creativecommons.org/licenses/by-nc-nd/4.0/ License, which permits use and distribution in any medium, provided the original work is properly cited, the use is non‐commercial and no modifications or adaptations are made.

Abstract

The relationship between speech production and perception is a topic of ongoing debate. Some argue that there is little interaction between the two, while others claim they share representations and processes. One perspective suggests increased recruitment of the speech motor system in demanding listening situations to facilitate perception. However, uncertainties persist regarding the specific regions involved and the listening conditions influencing its engagement. This study used activation likelihood estimation in coordinate‐based meta‐analyses to investigate the neural overlap between speech production and three speech perception conditions: speech‐in‐noise, spectrally degraded speech and linguistically complex speech. Neural overlap was observed in the left frontal, insular and temporal regions. Key nodes included the left frontal operculum (FOC), left posterior lateral part of the inferior frontal gyrus (IFG), left planum temporale (PT), and left pre‐supplementary motor area (pre‐SMA). The left IFG activation was consistently observed during linguistic processing, suggesting sensitivity to the linguistic content of speech. In comparison, the left pre‐SMA activation was observed when processing degraded and noisy signals, indicating sensitivity to signal quality. Activations of the left PT and FOC activation were noted in all conditions, with the posterior FOC area overlapping in all conditions. Our meta‐analysis reveals context‐independent (FOC, PT) and context‐dependent (pre‐SMA, posterior lateral IFG) regions within the speech motor system during challenging speech perception. These regions could contribute to sensorimotor integration and executive cognitive control for perception and production.

This meta‐analysis revealed a common neural network between speech production and challenging speech perception tasks, supporting the role of the motor system in facilitating perception under difficult listening conditions. The extent of neural overlap varied depending on the listening conditions, with the left frontal operculum emerging as a hub for challenging speech perception.

listening effort
meta‐analysis
speech perception
speech production
speech‐in‐noise
Natural Sciences and Engineering Research Council of Canada 10.13039/501100000038 RGPIN‐2021‐02721 Canadian Institutes of Health Research (CIHR) 10.13039/501100000024 Alzheimer Society of Canada 10.13039/501100000143 source-schema-version-number2.0
cover-dateSeptember 2024
details-of-publishers-convertorConverter:WILEY_ML3GV2_TO_JATSPMC version:6.4.8 mode:remove_FC converted:13.09.2024
Perron, M. , Vuong, V. , Grassi, M. W. , Imran, A. , & Alain, C. (2024). Engagement of the speech motor system in challenging speech perception: Activation likelihood estimation meta‐analyses. Human Brain Mapping, 45 (13 ), e70023. 10.1002/hbm.70023
==== Body
pmc Practitioner Points

Meta‐analyses of functional neuroimaging studies revealed a common neural network for speech production and challenging speech perception tasks, supporting the idea that these functions share neural resources.

The degree of neural overlap between speech production and perception varies with listening conditions such as noise, spectral degradation, and linguistic complexity, suggesting a context‐dependent role of the motor system.

Involvement of the left frontal operculum was observed across all conditions, underscoring its pivotal role as the hub of the speech motor system in challenging speech perception.

1 INTRODUCTION

The interplay between speech production and perception was initially introduced by Liberman et al. (1967) as the Motor Theory of Speech Perception (for a review, see McGettigan & Tremblay, 2018). This theory asserts that speech perception depends on articulatory representations in the brain—the motor schemes used to produce speech. Speech perception is thought to automatically activate the motor representations corresponding to each sound in the continuous speech signal from the listener's articulatory repertoire (Liberman & Mattingly, 1985). In the Motor Theory, speech sounds serve as informative cues, facilitating the immediate perception of gestures. This theory gained popularity following the discovery of mirror neurons in the premotor cortex of macaque monkeys, which discharged when they observed hand and mouth movements performed by an experimenter (Gallese et al., 1996; Rizzolatti & Craighero, 2004). Notably, the firing rate of these neurons also increased in response to the sounds accompanying the actions (Kohler et al., 2002). These findings suggest that the motor system is activated simultaneously with the perception of action (whether auditory or visual) and spurred extensive research into the role of the motor network in human speech perception.

However, the evidence from functional magnetic resonance imaging (fMRI) and transcranial magnetic stimulation (TMS) studies supporting the role of the speech motor network, including key regions such as the primary motor cortex (M1), ventral premotor cortex (PMv), and posterior inferior frontal gyrus (IFG), in perception remains equivocal. For instance, using fMRI, Wilson et al. (2004) examined brain activity associated with producing and listening to meaningless syllables. They found that the upper part of the left PMv was activated during both conditions. Watkins and Paus (2004) showed that increased excitability in the motor system during speech perception correlates with activity in the left pIFG, which was taken as evidence that the left pIFG primes the motor system in response to speech sounds. Additionally, Pulvermüller et al. (2006) provided fMRI evidence indicating that activity in the left M1 during speech perception is organized somatotopically, resembling mirror‐like neurons that reflect the specific effector used in speech production. Several TMS studies showed that stimulation of M1 and PMv impacts perception (e.g., Brisson & Tremblay, 2021; D'Ausilio et al., 2009; Meister et al., 2007; Mottonen et al., 2014; Sato et al., 2009; Smalle et al., 2015), providing support for the role of the speech motor system in perception. However, not all fMRI studies show motor area activity during speech perception (e.g., Matchin et al., 2014), and lesions of frontal speech areas do not significantly affect speech perception (Hickok et al., 2011; Stasenko et al., 2013). These findings challenge the prevailing notion that the speech motor system is essential to speech perception, suggesting instead that speech perception can occur independently of speech motor system involvement.

An alternative perspective gaining ground, often described as the “weak” version of the Motor Theory of Speech Perception, posits that the involvement of speech motor network in perception is contingent on the task demands (Wu et al., 2014). This idea is founded on converging evidence from neuroimaging studies showing increased activation in speech motor regions when listening to noisy or degraded speech than when listening to clear speech. For instance, Du et al. (2014) showed that the left PMv activation in young adults covary as a function of signal‐to‐noise ratio (SNR) and correlated with performance in a phonemic identification task. Likewise, Callan et al. (2014) found that reducing the SNR enhanced PMv responses during a vowel‐identification task. Moreover, using dynamic causal modeling for effective connectivity analysis, Osnes et al. (2011) observed top‐down connections from the premotor cortex to posterior temporal regions during the processing of partially intelligible speech, with no such connections during the perception of non‐speech sounds. While these findings are consistent with the proposal that the speech motor system provides top‐down information to resolve signal ambiguity in adverse listening conditions (e.g., Du et al., 2014), the absence of direct comparisons between brain activity related to speech perception and that related to speech production in several studies prevents confirmation of the overlap in brain activity between the two functions.

Activation likelihood estimation (ALE) meta‐analyses provide a means to address this issue because they identify and compare brain activation across fMRI studies, regardless of methodological differences including tasks, sample sizes, and scanner types. Only one meta‐analysis has investigated the neural overlap between speech production and difficult speech perception (Adank, 2012). The author observed overlaps in the bilateral anterior superior temporal sulcus and pre‐supplementary motor area (pre‐SMA), suggesting the existence of a shared network of temporal and superior frontal regions between the two functions. However, these results should be interpreted with caution due to the small number of studies included in the analysis (10 for speech perception and 21 for speech production), which could produce spurious effects induced by a single study. At least 17 experiments per condition should be included in an ALE meta‐analysis to achieve sufficient power (Eickhoff et al., 2016). In addition, the meta‐analysis of speech perception included a wide range of studies involving the manipulation of multiple parameters, including temporal compression, accentuation, pitch shifting, noise, and speech vocoding. The convergence of brain activation from speech perception and production tasks may vary depending on the listening conditions.

In a meta‐analysis by Alain et al. (2018), the engagement of specific brain regions in perception was examined in three challenging listening contexts: speech‐in‐noise (SIN), spectrally degraded speech and linguistically complex speech. While all conditions activated the bilateral superior temporal cortex and the insula, they also involved distinct areas in the frontal, parietal, and temporal lobes, indicating specificity within the speech perception network. Notably, each task recruited a different part of the left IFG. However, as the authors did not include speech production studies in their meta‐analysis, it is unclear whether the identified regions overlap with those from the speech motor system. Based on these findings, if the speech motor system is indeed shared with the perception network, it is likely to occur in a context‐dependent manner, as different listening conditions engage different brain regions.

In this study, we used coordinate‐based ALE meta‐analyses to test whether the increased frontal activity in adverse listening situations overlaps with frontal activation elicited during speech production. We focused on the three listening conditions examined by Alain et al. (2018) (i.e., SIN, spectrally degraded speech and linguistically complex speech) as they are extensively studied in the literature. All three have demonstrated recruitment of frontal regions, which could overlap with the speech motor system. SIN requires effort to filter out background noise, spectrally degraded speech requires effort to compensate for missing or distorted spectral information in the speech signal, and linguistically complex speech requires effort to process and understand content. We hypothesized the existence of shared neural networks between speech production and difficult speech perception, particularly in left frontal areas, supporting the idea that speech motor areas are activated under difficult listening conditions to facilitate perception. We also anticipated variations in this neural overlap as a function of specific listening conditions. If validated, this result could help to reconcile inconsistencies in previous studies by showing that the involvement of motor processes in difficult speech perception conditions is context‐dependent.

2 METHODS

Different databases and search strategies were used to identify challenging (i.e., difficult) speech perception and speech production studies used in the meta‐analyses. For the difficult speech perception meta‐analysis, we adhered to the PRISMA 2020 (Preferred Reporting Items for Systematic Reviews and Meta‐Analyses) guidelines (Page et al., 2021). For the speech production meta‐analysis, we used the BrainMap database. We used this approach because our aim was not to review the speech production network systematically, but to build a neurobiological model of the speech motor system. The list of all coordinates and the activations maps are openly available on the Borealis Dataverse (https://doi.org/10.5683/SP3/KKC0RM).

2.1 Speech perception meta‐analysis

2.1.1 Search strategy

The literature search was conducted on Medline (Pubmed) and PsycINFO (ProQuest) in January 2022, and updated on May 2, 2024. The search was conducted to identify brain regions involved in (1) SIN, (2) spectrally degraded speech, and (3) linguistically complex speech. SIN refers to situations where speech is presented simultaneously with other sounds at different SNRs, causing energetic or informational masking. Energetic masking occurs when noise physically interferes with the speech signal, as in the case of white noise, speech‐shaped noise or multi‐talker babble noise. Informational masking, on the other hand, occurs when noise perceptually interferes with the speech signal, often resulting in overlapping streams that make it difficult to distinguish between the task‐relevant speech sounds and other irrelevant speech sounds. Spectrally degraded speech refers to situations where the speech signal has been modified to reduce its spectral (frequency) resolution using noise vocoding, filtering, spectral modulation, or inversion techniques. Linguistically complex speech refers to situations in which the speech signal remains unchanged, but the linguistic structure is manipulated to increase the processing requirements for comprehension. Studies involving the manipulation of syntactic, lexical, semantic, grammatical, or phonological elements were included.

Three search queries were created on both databases: (1) difficult speech perception, (2) neuroimaging, and (3) difficult speech perception AND neuroimaging. The search queries and the number of results for each are presented in Supplemental Material 1. On Medline, queries contained relevant MeSH terms and combinations of key terms (text only [tw]). On PsycINFO, queries contained relevant index terms identified in the ProQuest thesaurus and key terms (All fields) identical to those in Medline. The search was further constrained to peer‐reviewed articles. In addition, key terms (title/abstract) were included to exclude literature on children and adolescents. Truncated terms were used to ensure that as many word forms as possible were included in the search (e.g., listen* for “listen,” “listening,” “listeners”). Finally, the reference sections of the selected articles were examined for additional articles of interest, and all review articles found in the primary search were manually examined for articles that had not been previously identified.

2.1.2 Screening process

All references were exported to Covidence (www.covidence.org) where the entire selection process took place. Covidence is a web‐based collaboration software platform that streamlines the production of systematic and other literature reviews. The first four authors of this manuscript conducted the selection process. The selection process is described in Figure 1. Two reviewers independently screened the titles and abstracts, followed by the full text to identify studies that met the eligibility criteria. The first author of this manuscript reviewed each article at each stage. Disagreements regarding inclusion were clarified by discussion. When no consensus was reached between the two reviewers, a third reviewer was consulted.

FIGURE 1 Screening process for studies included in the speech perception meta‐analysis.

Articles were included if they met the following criteria: (a) the study was published in a peer‐reviewed journal, (b) the primary methodology was fMRI or Positron Emission Tomography (PET), (c) participants were right‐handed, young or middle‐aged, healthy adults with no hearing problems, psychiatric or neurological disorders, or brain abnormalities, (d) the study reported a speech perception task performed during scanning, (e) the speech material was presented auditorily, (f) the study reported higher activity associated with difficult speech perception, (g) the study included whole‐brain fMRI or PET scan, and (h) the study reported Talairach or Montreal Neurological Institute (MNI) coordinates for activation. To ensure consistency and reduce confounding variables, we focused exclusively on fMRI and PET, excluding other techniques like magnetoencephalography (MEG) and electroencephalography. This choice was made to avoid variations caused by differences in spatial and temporal resolution between these imaging methods.

Increased activity during challenging speech perception tasks has been operationalized by correlations or contrasts between conditions that impose higher processing demands compared to those that impose lower demands. In SIN studies, this involves increased activity when noise intensity increases (i.e., a decrease in SNR) or when noisy conditions are contrasted with quieter ones. Similarly, in speech spectral degradation tasks, it involves increased activity with greater spectral degradation or when highly degraded speech is contrasted with less degraded speech. Regarding linguistic complexity, it involves increased activity when speech exhibits greater linguistic complexity or when the linguistic content presents comprehension challenges (e.g., semantically ambiguous sentences), compared to conditions where such challenges are absent (e.g., unambiguous sentences).

Studies of aging and those that combined data from healthy participants and patients were excluded. Studies contrasting or using only nonwords as speech stimuli were excluded because the perception and production of nonwords may recruit regions other than those generally used for normal communication. In the case of bilingual studies, they were included in the meta‐analysis if at least one contrast of interest was reported for the native language of the participants. Although studies were not excluded based on language, only articles written in English or French were included.

Inter‐rater reliability analysis revealed strong agreement among reviewers (McHugh, 2012), for both abstract and title screening (Cohen's kappa = 0.81, 94% agreement) and full‐text review (Cohen's kappa = 0.80, 93% agreement).

2.2 Speech production meta‐analysis

A different search strategy was used for speech production. We used the BrainMap database, which comprised 16,901 experiments at the time of the study. We conducted the database search using Sleuth 3.0.4 (http://www.brainmap.org/sleuth/). The exact search query can be found in Supplemental Material 1. We only considered activation data from fMRI or PET studies with healthy (i.e., normal mapping), right‐handed adult participants aged 18–60 years. In addition, we selected all behavioral domains that captured speech production (i.e., execution‐speech) without restriction for paradigm type. Finally, we only included experiments with a low‐level control condition. We chose low‐level control conditions to capture all aspects of speech production. High‐level control conditions, such as noun production versus verb production or covert speech versus overt speech, might have excluded important elements of production. To ensure this choice did not affect our results, we compared the production maps from both control conditions. Maps showed similar activations in bilateral superior temporal gyrus (STG), frontal, and insular areas. Our exclusion criteria were applied to the experimental conditions, not the article itself.

2.3 Activation likelihood estimation

Coordinate‐based meta‐analyses were performed using the ALE algorithm on GingerALE software (version 3.0.2) available at https://www.brainmap.org/ale (Eickhoff et al., 2009; Eickhoff et al., 2012). All Talairach coordinates were transformed into MNI space using the icbm2tal algorithm (Lancaster et al., 2007) implemented in the GingerALE toolbox (spatial normalization of the data was done using SPM). The ALE method constructs a brain activation map by calculating a 3D Gaussian density distribution centered on each coordinate within an experiment. Then, for each Gaussian density distribution function, the full width at half maximum (FWHM) is weighted by the sample size included in the experiment. The FWHM is estimated using the Euclidean distance between foci representing the spatial uncertainty around each coordinate, and the activation probabilities for each voxel are then calculated. Finally, voxel‐wise scores are obtained, indicating the convergence of activation in similar voxel locations between studies.

For each ALE calculation, we used the MNI152 reference space, with a more conservative mask size and the Turkeltaub nonadditive method (Turkeltaub et al., 2012), which corrects for within‐experiment effects derived from the proximity of foci reported in experiments. Statistical significance of the ALE scores was determined by a permutation test using cluster‐level inference at p < .05 (FWE), with a cluster‐forming threshold set at p < .001 and 1000 permutations. ALE maps were extracted for speech production, difficult speech perception (all speech perception studies combined), and each listening condition (SIN, spectral degradation, linguistic complexity). The resulting ALE maps from each group were tested for similarity (conjunction) to investigate whether speech perception and production shared common brain regions. Conjunction analyses were performed between (a) difficult speech perception (all studies) and speech production, (b) SIN and speech production, (c) spectrally degraded speech and speech production, and (c) linguistically complex speech and speech production. In GingerALE, the conjunction image is created using the voxel‐wise minimum value between the two thresholded ALE images of interest. All significant clusters were anatomically labeled using the SPM Anatomy toolbox (v3.0) (Eickhoff et al., 2005; Eickhoff et al., 2006). Anatomy toolbox provides anatomical labels for peak coordinates in MNI space based on probabilistic maps. We used cytoarchitectural probability maps within the frontal region to determine if the activation peaks were situated in BA44, BA45 (Amunts et al., 2004), or their corresponding ventrolateral frontal operculum (FOC) areas (Amunts et al., 2010). The Talairach Daemon atlas, as implemented in GingerALE, was used to determine the Brodmann area (BA) for each peak (based on the nearest gray matter within 5 mm) (Lancaster et al., 1997; Lancaster et al., 2000). Finally, we used the Human Motor Area Template to specify motor areas (Mayka et al., 2006). For visualization, ALE‐statistic maps were overlaid on the MNI152 brain template using MRIcroGL (https://www.nitrc.org/projects/mricrogl/).

3 RESULTS

For speech perception, 89 studies met the inclusion criteria, including 26 studies of SIN (424 participants), 19 studies of spectrally degraded speech (348 participants), and 45 studies of linguistically complex speech (822 participants) (Table 1). All studies included only one of the three listening conditions, except for one study that included two conditions (SIN and spectrally degraded speech) (Zekveld et al., 2014). For speech production, 35 studies were included (486 participants) (Supplemental Material 2). Figure 2 shows the individual foci in each condition.

TABLE 1 Studies included in the speech perception meta‐analysis.

Paper	Stimuli and contrast	Task and behavioral findings	Number of participants (sex and age)	Foci per contrast	Coordinate space	Source in paper	
Speech in noise	
Salvi et al. (2002)	Sentences in noise > sentences in quiet	Repeat aloud the last sentence word, not reported	10 (5F; 23–34 years)	7	Talairach	Table 1	
Sekiyama et al. (2003)	Syllables at low SNR minus visual control	McGurk effect for lower SNRs > higher SNRs	8 (1F; 22–46 years)	9	Talairach	Table 1	
Binder et al. (2004)	Phonemes in noise: negative correlation with SNR	Two‐interval forced‐choice procedure, decreased accuracy and increased RT with increasing noise	18 (10F; 24–47 years)	3	Talairach	Table 1	
Scott et al. (2004)	Sentences in noise masker and speech masker, negative correlation with SNR	Listen to the female speaker for meaning, accuracy for higher SNRs > lower SNRs	7 (0F; 35–52 years)	2	MNI	Figure 4	
Gaab et al. (2007)	Words in recorded scanner background noise > words in silence	Indicate whether two of four words presented were the same or not, no difference in accuracy	10 (5F; 18–22 years)	3	MNI	Table 1	
Wong et al. (2008)	Words in multitalker babble: SNR 20 dB > quiet, SNR −5 dB > quiet, SNR −5 dB > 20 dB	Word identification, longer RT and lower accuracy for SNR −5 dB than SNR 20 dB	11 (7F; 20–34 years)	20, 20, 2	Talairach	Table 1	
Dos Santos Sequeira et al. (2010)	Consonant‐vowel syllables in babble noise > babble noise	Dichotic listening task, accuracy in quiet > babble conditions	21 (8F; 19–35 years)	19	MNI	Table 4	
Peelle et al. (2010)	Sentences in continuous scanning EPI > quiet EPI	Respond to a probe word on a screen, no difference in behavioral performance	6 (3F; 20–26 years)	8	MNI	Table 4	
Adank et al. (2012)	Sentences in speech‐shaped noise > sentences in quiet	Indicate if a sentence was true or false, RT for speech in noise > clear speech	26 (20F; 18–28 years)	8	MNI	Table 2	
Zekveld et al. (2012)	Sentences in long‐term average speech spectrum > speech in quiet	Repeat sentence aloud, accuracy for speech in quiet > in noise	18 (9F; 20–29 years)	12	MNI	Table 2	
Buchweitz et al. (2012)	Dual sentences > right ear only, dual sentences > left ear only	Indicate whether statements are true or false in one sentence or two concurrent sentences; RT for concurrent sentences > single sentence.	12 (8F; 18–28 years)	8, 8	MNI	Table 2	
Manan et al. (2012)	Verbal working memory in noise > working memory in quiet	Repeat backward aloud all the words that were presented, accuracy for working memory in noise > working memory in quiet	15 (0F; 20–29 years)	7	MNI	Table 4	
Golestani et al. (2013)	Word in noise; negative correlation with SNR‐modulated accuracy	Choose between two visual probes, increase RT with decreasing SNR	9 (6F, not provided)	6	MIN	Table 2	
Vaden Jr. et al. (2013)	CVC syllables in multitalker babble: SNR +3 dB > +10 dB	Repeat each word aloud, hit rate for +10 dB > +3 dB SNR	18 (10F; 20–38 years)	5	MNI	Table 1	
Yusoff et al. (2014)	Arithmetic addition in noise > addition in quiet	Perform arithmetic addition, accuracy not reported	18 (0F; 20–28 years)	4	MNI	Table 1	
Du et al. (2014)	Phonemes in broadband noise: negative correlation with SNR‐modulated accuracy	Phoneme identification task, decreased accuracy and increased RT with decreasing SNR	16 (8F; 21–34 years)	8	Talairach	SI, Table 2	
Hervais‐Adelman et al. (2014)	Word in noise, negative correlation with SNR	Choose between two visual probes, increased RT with decreasing SNR	9 (6F, not provided)	6	MNI	SI, Table 1	
Zekveld et al. (2014)	Sentences in single‐talker masker > clear sentences, sentences in fluctuating noise > clear sentences	Indicate whether a visual probe was part of the sentence, accuracy for clear > sentences in noise, no difference between masker type	17 (9F; 19–33 years)	6, 8	MNI	Table 4	
Evans et al. (2016)	Two talker setting > single talker	Listen to a target speaker, hit rate for clear speech > masked speech	20 (10F; 19–36 years)	8	MNI	Table 1	
Guerreiro et al. (2016)	Words in low SNR > high SNR	One‐back task, accuracy for high SNR > low SNR	9 (7F; 19–56 years)	9	Talairach	Table 3	
Vaden Jr. et al. (2017)	Words in low SNR > high SNR	Repeat the word aloud with a delayed recognition memory task, word identification for low SNR < high SNR	20 (10F; mean age 29.8 ± 5.9 years	4	MNI	SI, Table 1	
Rammell et al. (2019)	Sentences in noise > sentences in quiet (native language contrast)	Repeat aloud sentences, no behavioral data provided	18 (13F; not provided; undergraduate students)	2	MNI	Table 1	
Planton et al. (2019)	Sentences in noise > sentences in quiet	Identify whether a sentence was repeated twice in a row or detect false statement, hit rate for sentences in quiet > sentences in noise	24 (11F; 20–32 years)	25	MNI	SI, Table 2	
Olano et al. (2020)	Words in low SNR > words in quiet, word in medium SNR > words in quiet	Indicate if they recognize the words, accuracy for quiet > medium SNR > low SNR	25 (16F; not provided; university students)	2, 2	MNI	Table 2	
Holmes and Johnsrude (2021)	Sentence masked by another sentence > sentence alone	Indicate whether a visual probe sentence was the same as the target sentence, sensitivity for masked sentence < alone sentence	27 (18F; 19–68 years)	21	MNI	Table 1	
Agmon et al. (2022)	Target in multiple speakers > single speaker (distributed > selective)	Listen to a single (selective attention) or multiple speakers (distributed attention) and respond to a target word, reduced accuracy as the number of speakers increased	32 (17F; 18.5–35 years)	8, 8	MNI	Table 3	
Spectrally degraded speech	
Meyer et al. (2002)	Prosodic filtered (PURR‐filtering) sentences > unfiltered sentences	Indicate if the sentence was normal or has prosody, accuracy for unfiltered > filtered speech	14 (6F; mean age 25.2 ± 6 6 years)	8	Talairach	Table 3	
Meyer et al. (2004)

	Distorted sentences > clear sentences	Prosody comparison task, no behavioral data provided	14 (8F; 18–27 years)	10	Talairach	Table 1	
Hesling et al. (2005)	Expressive filtered sentences > unfiltered sentences	Shadowing of normal speech > expressive filtered speech	12 (0F; 24–38 years)	4	Talairach	Table 5	
Allen et al. (2005)	Distorted words > clear words	Indicate whether the speech they heard was their own or not, greater misattribution when processing their own distorted speech	11 (0F; 24–36 years)	3	Talairach	Table 3	
Obleser et al. (2007)	Spectrally rotated analogues > intelligible consonants	Listen passively	13 (7F, 23.5 ± 7 years)	1	MNI	Table 1	
Sabri et al. (2008)	Spectrally rotated > clear words and pseudowords	Attend to either auditory or visual and performed one‐back matching task; sensitivity for rotated speech < clear words/pseudowords	28 (13F; mean age 26.5 ± 6.9 years)	11	Talairach	Table 1	
Eckert et al. (2009)	Filtered words, parametric increased in activity with decreasing intelligibility	Repeat words aloud, word recognition (accuracy) decreases with decreasing high‐frequency information	11 (5F; mean age 32.8 6 ± 10.9 years)	7	MNI	SI, Table 1	
Sharp et al. (2010)	Noise‐vocoded words > clear words	Indicate whether two words were related or not, RT for noise‐vocoded speech > clear speech	12 (5F; 35–61 years)	2	MNI	Table 2	
Takeichi et al. (2010)

	Modulated sentences > non‐modulated sentences	Listen to a narrative, accuracy for non‐modulated > modulated speech	23 (12F; 20–38 years)	2	MNI	Table 1	
Hervais‐Adelman et al. (2012)	Noise‐vocoded words > clear words	Respond to target sounds, accuracy for clear > noise‐vocoded speech	15 (10F; 18–35 years)	4	MNI	Table 2	
Wild, Davis, and Johnsrude (2012)	Noise‐vocoded sentences vs. clear sentences, main effect of intelligibility, noise‐elevated response	Repeat sentences aloud, accuracy for clear > noise‐vocoded speech	21 (13F; 19–27 years)	4	MNI	Table 1	
Wild, Yusuf, et al. (2012)	Noise‐vocoded sentences > clear speech and rotated noise‐vocoded sentences	Listened passively	19 (11F; 18–26 years)	1	MNI	Table 1	
Erb et al. (2013)	Noise‐vocoded sentence > clear sentence	Repeat sentence aloud, accuracy for clear > noise‐vocoded speech	30 (15F; 21–31 years)	5	MNI	Table 1	
Kyong et al. (2014)	Noise‐vocoded sentences: negative correlates of sentence intelligibility	Try to understand what was being said. Clear speech > noise vocoded speech	19 (18–40 years)	6	MNI	Table 5	
Zekveld et al. (2014)	Noise‐vocoded sentences > clear sentences	Indicate whether a visual probe was part of the sentence, accuracy for clear > noise‐vocoded speech	17 (9F; 19–33 years)	4	MNI	Table 4	
Evans and Davis (2015)	Noise‐vocoded phoneme > clear phoneme	One back task, sensitivity for clear > noise‐vocoded speech	17 (12F; 18–40 years)	3	MNI	Table 1	
Lee et al. (2016)	Noise‐vocoded sentences > clear sentences	Indicate gender of the character performing the action, accuracy for clear speech marginally higher than noise vocoded speech	26 (12F; 20–34 years)	7	MNI	Table 4	
Tuennerhoff and Noppeney (2016)	Filtered sentences > unfiltered sentences	Listen passively	20 (10F; median age 24.05 years)	12	MNI	Table 1	
Lin et al. (2022)	Noise‐vocoded two‐word terms > clear two‐word terms (pre‐training)	Pay attention to the played sounds and rate its intelligibility, no behavioral data provided	26 (14F; 22.6 ± 3.56 years)	3	MNI	Appendix, Tables 1–3	
Linguistic complexity	
Mellet et al. (1998)	Abstract > concrete definitions	Listen to and understand abstract and concrete words and their definition with implicit recall task, recall for abstract < concrete words	8 (0F; 20–25 years)	4	Talairach	Table 2	
Caplan et al. (1999)	Semantically implausible sentences > plausible sentences	Plausibility judgment task, RT for implausible > plausible sentences	16 (8F; 22–34 years)	3	Talairach	Table 2	
Kotz et al. (2002)	Semantically unrelated pairs > related pairs	Decide whether the target stimulus of word or pseudoword pairs is a word or a pseudoword, RT and errors for unrelated pairs > related pairs	13 (7F; mean age 23.5 years)	9	Talairach	Table 4	
Rissman et al. (2003)	Semantically unrelated words > related words	Lexical decision task, longer RT, and lower accuracy for unrelated than related words	15 (8F; 18–44 years)	5	Talairach	Table 2	
Peelle et al. (2004)	Object‐relative sentences > subject‐relative sentences	Indicate the gender of the character performing the action, longer RT, and lower accuracy for object related than subject‐related sentences	8 (4F; 19–27 years)	2	Talairach	Table 1	
Wartenburger et al. (2004)	Grammaticality incorrect sentences > correct sentences	Judge the grammaticality of sentences, no difference in performance between grammaticality incorrect and correct sentences	13 (6F; mean age 25.8 ± 4.9)	2	MNI	Table 2	
Rodd et al. (2005)	High semantic ambiguity sentences > low ambiguity sentences	Relatedness judgment task, no difference in performance between high and low ambiguity	15 (10F; 18–40 years)	4	MNI	Table 4	
Tyler et al. (2005)	Regular verb vs. irregular verb	Same or different judgment task, RT regular > irregular verb, no difference in error rate	18 (8F, mean age 24 ± 7)	9	MNI	Table 2	
Ferstl et al. (2005)	Inconsistent narratives > coherent narratives	Screen for chronological and emotional inconsistencies of stories; error rates for inconsistent > consistent, RT: significant type × consistency interaction	20 (12F; 21–34 years)	1	Talairach	Table 3	
Rüschemeyer et al. (2005)	Syntactically incorrect sentences > correct sentences, semantically incorrect sentences > correct sentences (native speakers)	Judgment of sentence semantic and syntactic correctness; RT for incorrect sentences > correct sentences	18 (7; 23–30 years)	2, 3	Talairach	Table 4	
Davis et al. (2007)	High ambiguity sentences > low ambiguity sentences	Listen passively	12 (3F; 29–42 years)	7	MNI	SI, Table 3	
Siebörger et al. (2007)	Pragmatically distantly related sentences > closely related sentences	Rate the perceived strength of coherence on a 4‐point scale.	14 (6F, 20–32 years)	1	Talairach	Table 3	
Bilenko et al. (2008)

	Semantically ambiguous words > unambiguous words	Speeded lexical decision task, slower RT for ambiguous than unambiguous words	16 (9F; 18–21 years)	2	Talairach	Table 3	
Grindrod et al. (2008)

	Sequences of semantically discordant words > neutral words	Lexical decision task, RT for discordant > concordant or neutral words	15 (8F; 19–29 years)	1	Talairach	Table 3	
Ruff et al. (2008)	Semantic unrelated words > related words	Relatedness judgment task, RT for unrelated > related words	15 (8F; 19–31 years)	9	Talairach	Table 2	
Shetreet et al. (2009)	Sentences with embedding > sentences without embedding	Indicate whether the event in a sentence described was more likely to happen at home or not; performance >85% for all participants	19 (8F; 23–38 years)	12	Talairach	Table 2	
Hillert and Buračas (2009)

	Ambiguous idiomatic > literal sentences	Judge whether sentences are meaningful or not, no difference in performance between sentence types	10 (7F; 21–31 years)	6	MNI	Table 2	
Friederici et al. (2010)	Syntactically incorrect sentences > correct sentences	Listen attentively	17 (9F; 20–30 years)	5	MNI	Table 1	
Meltzer et al. (2010)	Main effect of syntactic sentence complexity	Match heard spoken sentence with visual probes, RT for reversible > irreversible sentence	24 (12F; 22–37 years)	1	Talairach	Table 3	
Rodd et al. (2010)	High semantic ambiguity words > low semantic ambiguity words	Indicate whether a visual probe word is related or not to the sentence meaning	14 (19–37 years)	1	MNI	Table 3	
Raettig et al. (2010)	Morphosyntactically incorrect sentences > correct sentences, incorrect sentences > correct verb‐argument structure	Grammatical judgment task, high accuracy in all conditions	15 (8F; 21–29 years)	1, 1	Talairach	Table 3	
Obleser and Kotz (2010)	Sentences incorporating a verb low in cloze probability > verb high in cloze probability (main effect)	Listen passively	16 (9F; 22–32 years)	3	MNI	Table 1	
Bekinschtein et al. (2011)	Semantically ambiguous sentence > unambiguous sentence	Rate sentences as funny or not, no significant difference in performance between ambiguous and unambiguous sentences	12 (not provided; young adults)	2	MNI	Table 1	
Obleser et al. (2011)	Positive correlation with syntactic complexity	Listen attentively, no difference in accuracy	14 (8F; 23.4 ± 2.1 years)	1	MNI	Table 2	
Wright et al. (2011)	Complex words > simple words, main effect of complexity	Lexical decision task, RT for non‐words > real words, RT for complex > simple words	14 (sex not provided; 19–34 years)	1	MNI	Table 3	
Zhuang et al. (2011)	High cohort competition > low cohort competition	Indicate whether the incoming sound is a word, or a non‐word, RT for high competition > low competition	14 (7F; 19–33 years)	5	MNI	Table 3	
Yu et al. (2011)

	Factually incorrect sentences > correct sentences	Listen passively	18 (7F; 21–41 years)	7	MNI	Table 1	
Herrmann et al. (2012)	Syntactically incorrect sentences > correct sentences	Grammatical judgment task, accuracy for correct > incorrect grammar	25 (12F; 22–32 years)	5	MNI	Table 2	
Meyer et al. (2012)	Sentences with a long argument–verb distance > short argument–verb distance	Answer comprehension questions via button press, no difference in performance between conditions	24 (14F; 27.1 ± 3.2 years)	3	MNI	Table 1	
Guediche et al. (2013)	Semantically ambiguous sentences > unambiguous sentences	Press a button to target word, longer RT, and lower accuracy in ambiguous than unambiguous target	17 (8F; 24.5 ± 3.6 years)	3	Talairach	Table 3	
Rothermich and Kotz (2013)	Semantic incongruent > congruent (global effect in the semantic task), semantic incongruent > congruent (global effect in the metric task)	Determine the metrically or semantically correctness of sentences, no difference in performance between semantic congruent and incongruent sentences	16 (8F; mean 26 ± 3.8 years)	1, 3	MNI	Table 3	
Nagels et al. (2013)	Similes > non‐figurative control sentences	Listen passively	16 (0F; mean age 27 ± 6.65 years)	2	MNI	Table 1	
Buchweitz et al. (2014)	Unfamiliar sentences > familiar sentences	True–false question that probed comprehension, RT for unfamiliar > familiar passages	9 (3F; 18–25 years)	6	MNI	Table 3	
Vitello et al. (2014)	Semantic ambiguous sentences > unambiguous semantic sentences	Indicate whether a probe was related or unrelated to the sentence they just heard	20 (11F; 18–35 years)	3	MNI	Table 4	
Kristensen et al. (2014)	Main effect of context appropriateness (inappropriate context > appropriate context)	Listen to and occasionally answer a visually presented question, main effect of context appropriateness for percentage of correct responses.	32 (11F; 18–38 years)	2	MNI	Table 2	
Deschamps and Tremblay (2014)

	Syllabic complexity (complex > simple), supra‐syllabic complexity (complex > simple)	Listen passively	15 (9F; 21–34 years)	2, 2	MNI	Table 4	
Lopes et al. (2016)	Semantically complex > easy semantic decision task	Semantic decision task, accuracy for easy > complex semantic decision	24 (15F; 20–31 years)	6	MNI	Table 2	
Lyu et al. (2016)	Unexpected phrases > expected phrases	Indicate the gender of the speakers, RT for unexpected > expected phrases, no difference in accuracy	30 (15F; 21–28 years)	2	MNI	Table 1	
Tune et al. (2016)	Easy to detect anomalies > control sentences, borderline anomalies to detect > control sentences	Indicate whether sentences contain semantic anomalies, detection rate for easy to detect anomalies > borderline anomalies.	22 (11F; 20–30 years)	12, 9	MNI	Table 4	
Yang et al. (2017)	Long phrase > short phrase	Listen to and occasionally decide whether a visually presented word could be meaningfully related to the previously heard stimulus, no behavioral data provided	18 (8F; 19–33 years)	7	MNI	Table 4	
Lee et al. (2018)	Object‐relative sentences > subject‐relative sentences	Indicate the gender of the character performing the action; accuracy high across all conditions	42 (20F; 18–41 years)	3	MNI	Table 2	
Kousaie et al. (2019)	Low predictable sentences > high predictable sentences	Repeat aloud the last word of the sentences presented in their native or non‐native language, error rates for low predictability > high predictability sentences	30 (23F; mean age per group 23.1, 24.8, 26.6 years)	1	MNI	Table 3	
Agmon et al. (2021)	Negative quantifiers > positive quantifiers	True‐false decision whether a sentence described a picture, RT for positive > negative quantifiers	30 (17F; 19–36 years)	11	MNI	Table 1	
Rysop et al. (2021)	Low predictable sentences > high predictable sentences	Repeat sentences varying in intelligibility and predictability, proportion correct for low predictability < high predictability	26 (15F; 19–29 years)	8	MNI	Table 1	
Mechtenberg et al. (2021)	Semantically non‐predictive sentences > highly predictive sentences	Listen to and occasionally press a button when a visual word is presented to indicate whether the word was present in the sentence	23 (8F; 21–36 years)	1	Talairach	Table 1	

FIGURE 2 Foci from the speech‐in‐noise, spectrally degraded speech, linguistic complexity, and speech production studies.

3.1 Overlap between speech perception and production, regardless of listening condition

We first identified the convergence of brain activation during difficult speech perception by analyzing data from all 89 speech perception studies, with a total of 567 foci. Six clusters were found (Figure 3), with peak activations centered on the left pars opercularis of the IFG, bilateral planum temporale (PT), right frontal orbital cortex, and left pre‐SMA (Table 2a). Supplemental Material 3 shows that the activity broadly extended to bilateral areas of the STG, including its posterior and anterior portions, and Heschl's gyrus (i.e., primary auditory cortex). The activity also extended to the left insula and the bilateral FOC.

FIGURE 3 Overlay of ALE‐statistic maps for difficult speech perception (all studies, red clusters), speech production (green clusters), and the overlap between the two conditions (blue clusters).

TABLE 2 Meta‐analysis results showing brain areas consistently activated during (a) difficult speech perception (all listening conditions combined) and (b) speech production, and (c) overlapping brain regions between difficult speech perception and speech production.

Brain region, probability (cytoarchitecture assignment) a	Motor region b	BA c	MNI coordinates				
x	y	z	ALE (peak)	Cluster size (mm3)	# Experiments d	
a. Speech perception (all studies)	
Left inferior frontal gyrus, pars opercularis, 23% (BA44, 19%; OP9, 8%; OP8, 8%; BA45, 6%)	‐	45	−42	24	−2	0.05	16,856	52	
Left planum temporale, 19%	‐	41	−56	−20	4	0.05	12,552	35	
Right planum temporale, 24% (OP4, 4%)	‐	41	60	−18	2	0.04	6568	25	
Right frontal orbital cortex, 25% (OP9, 8%; OP8, 4%)	‐	*	32	24	0	0.06	5424	26	
Left paracingulate gyrus, 46%	Pre‐SMA	6	0	18	50	0.03	2928	15	
b. Speech production	
Left precentral gyrus, 15%	PMd/M1	4	−48	−10	40	0.07	44,272	54	
Right precentral gyrus, 20.3% (OP4, 6%)	PMd/M1	4	52	−8	38	0.06	18,816	26	
Left supplementary motor cortex, 30%	Pre‐SMA	32	−2	14	44	0.06	12,352	46	
Left fusiform gyrus, 35%	‐	37	−48	−56	−16	0.03	3120	14	
Right cerebellum VI, 91%	‐	‐	16	−62	−22	0.04	3056	14	
Left cerebellum VI, 88%	‐	‐	−12	−64	−16	0.04	2872	12	
Right thalamus, 61%	‐	‐	14	−18	2	0.04	1968	10	
Right putamen, 69%	‐	‐	22	0	6	0.04	1752	8	
c. Speech perception ∧ speech production	
Left planum temporale, 25%	‐	41	−56	−20	4	0.05	9192	64	
Left insula, 34% (OP8, 12%, OP9, 5%)	‐	13	−34	24	4	0.04	4552	46	
Right superior temporal gyrus, posterior division, 24% (OP4, 6%)	‐	41	62	−18	4	0.04	4472	35	
Left paracingulate gyrus, 46%	Pre‐SMA	6	0	18	50	0.03	2208	24	
Left inferior frontal gyrus, pars opercularis, 58%	‐	44	−54	8	8	0.02	80	0	
Left precentral gyrus, 54%	PMv/M1	9	−48	2	26	0.02	80	0	
Left inferior frontal gyrus, pars opercularis, 68%	‐	44	−52	12	6	0.02	40	0	
Left inferior frontal gyrus, pars triangularis, 71%	‐	44	−52	16	4	0.02	32	0	
Left paracingulate gyrus, 83%	Pre‐SMA	32	−4	26	38	0.02	8	0	
Abbreviations: M1, primary motor cortex; PMv, ventral premotor cortex; PMd, dorsal premotor cortex; SMA, supplementary motor area.

a Labels were assigned using the Anatomy toolbox based on the macroanatomy maximum likelihood map. The percentage value represents the probability of belonging to the region. Information in brackets represents the probability of cytoarchitecture for Brodmann areas (BA) 44 and 45, as well as for each part of the frontal operculum.

b Labels were assigned using the Human Motor Area Template. Cases with “‐” represent non‐motor areas.

c Brodmann area according to Talairach Daemon. Cases with “*” represent areas where no Brodmann area was found. Cases with “‐” represent areas with no Brodmann area.

d Number of experiments contributing to the clusters.

Subsequently, we identified brain areas during speech production by combining data from all 35 speech production studies into a single analysis. The analysis included a total of 1309 foci. Eight clusters were found (Figure 3), with peak activation centered in the bilateral premotor cortex and M1, left pre‐SMA, left fusiform gyrus, bilateral cerebellum VI, and right thalamus and putamen (Table 2b). As shown in Supplemental Material 3, activation peaks spread towards the left PMv, bilateral posterior temporal regions, the left basal ganglia, the right cerebellum, and the left insula and central opercular cortex.

The blue clusters in Figure 3 show overlap in brain activation between speech perception and production. A total of nine clusters were identified (Table 2c), with peak activations centered in the left PT, left insula, right posterior STG, left pre‐SMA, and, to a minor extent (<100 mm3), the left PMv, as well as the left pars opercularis and pars triangularis of the IFG. Activations also extensively extended to the bilateral superior temporal regions (Supplemental Material 3). Notably, the peak activation originating from the left insula further extended to the pars opercularis of the IFG and the FOC (likely OP8).

3.2 Overlap between speech perception and production, depending on the listening condition

Here, we examined whether the degree of overlap between speech perception and production differed between SIN, spectrally degraded speech, and linguistically complex speech.

3.2.1 SIN

SIN manipulations yielded a total of 268 foci across 26 experiments. Processing SIN results in peak activations centered in the bilateral PT, the left FOC (likely OP8), the right insula, the left pars opercularis of the IFG, and the right pre‐SMA (Figure 4, Table 3a, Supplemental Material 4). The conjunction analysis between SIN and speech production revealed areas of overlap in the bilateral PT, the right pre‐SMA, the left FOC (likely OP8), and to a very limited extent (<10 mm3) within the left PMv and M1 (Figure 4, Table 4a, Supplemental Material 5).

FIGURE 4 Overlay of ALE‐statistic maps for speech perception and speech production (green clusters), and the overlap between the conditions (blue clusters), depending on the listening manipulations: (a) speech‐in‐noise, (b) spectrally degraded speech, and (c) linguistically complex speech.

TABLE 3 Meta‐analysis results showing brain areas consistently activated during listening to (a) speech‐in‐noise (SIN), (b) spectrally degraded speech, and (c) linguistically complex speech.

Brain region, probability (cytoarchitecture assignment) a	Motor region b	BA c	MNI coordinates	ALE (peak)	Cluster size (mm3)	# Experiments d	
x	y	z	
a. SIN	
Left planum temporale, 33%	‐	41	−56	−20	4	0.03	4200	12	
Left frontal operculum cortex, 37% (OP8, 15%; OP9, 11%)	‐	13	−32	24	8	0.03	3232	14	
Right insula, 32% (OP9, 10%; OP8, 9%)	‐	13	34	24	4	0.03	2520	10	
Right planum temporale, 38% (OP4, 9%)	‐	41	58	−18	2	0.02	2168	9	
Left inferior frontal gyrus, pars opercularis, 54% (BA44, 50%)	‐	9	−48	8	26	0.02	1312	7	
Right paracingulate gyrus, 56%	Pre‐SMA	6	2	18	50	0.02	944	5	
b. Spectrally degraded speech	
Left insula, 46% (OP8, 10%)	‐	*	−32	18	−6	0.02	2456	8	
Right insula, 54%	‐	*	34	20	−4	0.03	2328	8	
Right planum temporale, 58%	‐	41	62	−20	12	0.01	1296	6	
Left cingulate gyrus, posterior division, 44%	M1	23	2	−22	28	0.02	1040	5	
Left paracingulate gyrus, 70%	Pre‐SMA	8	−8	22	44	0.02	968	4	
Left planum temporale, 34%	‐	41	−42	−28	8	0.01	888	4	
c. Linguistic complexity	
Left inferior frontal gyrus, pars opercularis, 30% (BA44, 21%; BA45, 11%; OP9, 11%; OP8, 5%)	‐	9	−54	20	20	0.03	11,096	29	
Left middle temporal gyrus, 15%	‐	22	−54	−40	4	0.03	6864	18	
Right superior temporal gyrus, posterior division, 35%	‐	22	60	−16	2	0.03	2472	8	
Right frontal orbital cortex, 41% (OP9, 5%)	‐	*	32	28	−2	0.02	1176	5	
Abbreviations: M1, primary motor cortex; PMv, ventral premotor cortex; PMd, dorsal premotor cortex; SMA, supplementary motor area.

a Labels were assigned using the Anatomy toolbox based on the macroanatomy maximum likelihood map. The percentage value represents the probability of belonging to the region. Information in brackets represents the probability of cytoarchitecture for Brodmann areas (BA) 44 and 45, as well as for each part of the frontal operculum.

b Labels were assigned using the Human Motor Area Template. Cases with “‐” represent non‐motor areas.

c Brodmann area according to Talairach Daemon. Cases with “*” represent areas where no Brodmann area was found. Cases with “‐” represent areas with no Brodmann area.

d Number of experiments contributing to the clusters.

TABLE 4 Meta‐analysis results showing brain area overlap between speech production and listening to (a) speech in noise (SIN), (b) spectrally degraded speech, and (c) linguistically complex speech.

Brain region, probability (cytoarchitecture assignment) a	Motor region b	BA c	MNI coordinates				
x	y	z	ALE (peak)	Cluster size (mm3)	# Experiments d	
a. SIN ∧ speech production	
Left planum temporale, 34%	‐	41	−56	−20	4	0.03	4088	33	
Left frontal operculum cortex, 33% (OP8, 13%; OP9, 8%)	‐	13	−32	24	8	0.03	2752	31	
Right planum temporale, 26% (OP4, 11%)	‐	41	58	−18	2	0.02	1720	14	
Left paracingulate gyrus, 70%	Pre‐SMA	6	2	18	50	0.02	680	7	
Left precentral gyrus, 38% (BA44, 13%)	PMv/M1	9	−48	2	26	0.01	8	0	
b. Spectrally degraded speech ∧ speech production	
Left insula, 56% (OP8, 14%)	‐	47	−34	18	−6	0.02	1720	17	
Left paracingulate gyrus, 95%	Pre‐SMA	8	−8	22	44	0.02	544	4	
Left planum temporale, 44%	‐	41	−42	−28	8	0.01	496	5	
Right planum temporale, 49%	‐	41	62	−20	10	0.01	248	2	
Left planum temporale, 44.2%	‐	41	−46	−34	8	0.01	8	0	
c. Linguistic complexity ∧ speech production	
Left planum temporale, 18%	‐	22	−54	−40	4	0.03	4192	39	
Right superior temporal gyrus, posterior division, 34%	‐	22	60	−16	2	0.03	2464	21	
Left frontal orbital cortex, 59% (OP9, 10%; OP8, 7%)	‐	47	−44	24	−8	0.02	1496	12	
Left inferior frontal gyrus, pars opercularis, 62% (BA44, 37%)	‐	44	−54	8	8	0.02	224	0	
Left supramarginal gyrus, posterior division, 40%	‐	21	−52	−46	12	0.01	8	0	
Abbreviations: M1, primary motor cortex; PMv, ventral premotor cortex; PMd, dorsal premotor cortex; SMA, supplementary motor area.

a Labels were assigned using the Anatomy toolbox based on the macroanatomy maximum likelihood map. The percentage value represents the probability of belonging to the region. Information in brackets represents the probability of cytoarchitecture for Brodmann areas (BA) 44 and 45, as well as for each part of the frontal operculum.

b Labels were assigned using the Human Motor Area Template. Cases with “‐” represent non‐motor areas.

c Brodmann area according to Talairach Daemon. Cases with “*” represent areas where no Brodmann area was found. Cases with “‐” represent areas with no Brodmann area.

d Number of experiments contributing to the clusters.

3.2.2 Spectrally degraded speech

Spectral manipulations yielded a total of 97 foci within 19 experiments. Processing spectrally degraded speech results in peak activations centered in the bilateral insula (extending to FOC), bilateral PT, left posterior cingulate gyrus, and left pre‐SMA (Figure 4, Table 3b, Supplemental Material 4). The conjunction analysis revealed overlap in the bilateral PT, left insula (extending to FOC, likely OP8), and left pre‐SMA (Figure 4, Table 4b, Supplemental Material 5).

3.2.3 Linguistic complexity

Linguistic manipulations yielded a total of 202 foci across 45 experiments. Processing linguistically complex speech results in peak activations centered in the left pars opercularis of the IFG, extending notably to the pars triangularis and the FOC (likely OP9). Additional peak activations were identified in the left middle temporal gyrus, right posterior STG, and right frontal orbital cortex (Figure 4, Table 3c, Supplemental Material 4). Conjunction analysis showed overlap in the left PT, right posterior STG, left pars opercularis of the IFG (extending to FOC), left frontal orbital cortex, and, to a very small extent (<10 mm3), the left supramarginal gyrus (Figure 4, Table 4c, Supplemental Material 5).

3.3 Common region in overlaps

Additional analysis was conducted to determine if there was a common region in the overlap between speech production and the three listening conditions. To do this, the three ALE conjunction maps were overlaid in Mango Viewer (https://mangoviewer.com) and the “Find overlay clusters” option was used. The analysis revealed a cluster (size = 6994 mm) with peak activation at x = −40, y = 22, z = 0 (Figure 5). The identified cluster was input into the SPM Anatomy toolbox, which identified that this cluster predominantly corresponds to the left FOC (63%), most likely within its OP8 (36%) areas that is adjacent to the pars opercularis of the IFG (Amunts et al., 2010).

FIGURE 5 Brain regions common to speech production and the three conditions of speech perception (speech‐in‐noise, spectrally degraded speech, linguistically complex speech).

4 DISCUSSION

The present study tested the hypothesis that the speech motor network engagement during perception depends on task demands (Wu et al., 2014). Using ALE meta‐analyses, we found a significant overlap in brain activation between difficult speech perception and speech production tasks, thereby supporting this hypothesis. We observed that this overlap is predominantly left‐lateralized, which aligns with the well‐established idea that multiple language functions exhibit a left‐lateralized pattern (for a review, see Van der Haegen & Cai, 2019). This finding also aligns with dual‐models of speech processing, suggesting left‐lateralized auditory‐motor integration in the dorsal speech stream (Hickok & Poeppel, 2007; Rauschecker & Scott, 2009).

Specifically, we identified key nodes in the left frontal, insular, and temporal regions—such as the FOC, PT, pre‐SMA, and posterior lateral IFG—as crucial for speech perception and production processes. The extent of this overlap differed depending on the listening conditions. The left posterior lateral IFG showed overlap only during challenging linguistic processing, the left pre‐SMA during the perception of degraded and noisy signals, while the left PT and FOC were engaged in all conditions. Notably, a distinct area in the left posterior FOC systematically overlapped in all conditions. These results highlight the context‐independent (FOC, PT) and context‐dependent (pre‐SMA, posterior lateral IFG) roles of the speech motor system in difficult speech perception. These regions likely contribute to sensorimotor integration and cognitive control for speech production and perception.

Another important finding of this meta‐analysis is the limited, almost nonexistent overlap in the left M1 and PMv areas. Although their role in perception cannot be completely ruled out, these areas do not appear to be central components of the network involved in difficult speech perception. Rather, our results suggest the context‐dependent and context‐independent involvement of other regions of the speech motor system in difficult speech perception.

4.1 Context‐independent overlaps

4.1.1 The role of the frontal operculum

A key result of this meta‐analysis is that the left posterior FOC, associated with speech production, is consistently activated in challenging listening situations such as when individuals are presented with SIN, spectrally degraded speech, and linguistically complex speech stimuli. This finding is consistent with a previous meta‐analysis from Alain et al. (2018), who observed a peak of activity in the left FOC that was shared between SIN and spectrally degraded speech (Talairach coordinates: x = −36, y = 18, z = 8), although this region was labeled as the left insula (the atlas implemented in GingerALE does not include the FOC). Here, we additionally found that the FOC region adjacent to the pars opercularis of the IFG overlaps in all conditions, indicating its potential central role in perception and production.

Anatomically, the FOC can be defined as the extension of the lateral prefrontal cortex that forms the lid of the insular cortex. It begins at the anterior branch of the lateral fissure and extends to the lower parts of the precentral gyrus, encompassing the pars triangularis and opercularis. While traditionally considered part of the “Broca area,” its cytoarchitectonic structure suggests potential subdivisions into distinct regions, exhibiting structural differences from the lateral portion of the IFG (Amunts et al., 2010; Unger et al., 2023). The FOC may be a relatively recent evolutionary change of the human brain that has evolved in response to the growing need for speech‐related cognitive and motor functions (Amiez et al., 2023). Anwander et al. (2006) showed that FOC connects to the anterior superior temporal regions via the uncinate fasciculus, while the lateral part of the IFG is predominantly connected to the posterior STG via the arcuate fasciculus. In a study of cortico‐cortical‐evoked potentials in epileptic patients, Mălîia et al. (2018) showed that the left FOC projects to M1 and sensory areas, the rolandic operculum and inferior prefrontal cortex. The FOC's most significant afferent connections were with the dorsolateral prefrontal cortex, middle cingulate gyrus, SMA, posterior insula, and precentral gyrus. This structural and connectivity profile supports that the left FOC plays a crucial role as a hub for both speech production and perception processes.

Research on speech production highlights the involvement of the FOC in various aspects of language processing. For instance, Bohland and Guenther (2006) showed that the production of complex sequences of syllables (phonological complexity) led to bilateral FOC activation compared to simple sequences of syllables. Mălîia et al. (2018) found that in epileptic patients, FOC stimulation is associated with expressive language‐related effects, with patients reporting a lack of internal word generation or intrusive feelings disrupting verbal sequences. However, the role of the left FOC in speech perception remains somewhat uncertain. Some argue for a general role in syntax and grammatical processing (Friederici, 2011; Friederici et al., 2006), particularly in detecting errors and violations. This aligns with our observation of its engagement in linguistic complexity. Others posit that the FOC operates within the broader cingulo‐operculum network responsible for monitoring and adjusting performance across various tasks through top‐down attentional control (e.g., Seeley et al., 2007). This is consistent with studies suggesting that activation in cingulo‐opercular regions predicts performance in SIN tasks (Vaden Jr. et al., 2013; Vaden Jr. et al., 2017). Evidence also suggests the FOC plays a crucial role in task monitoring, particularly in evaluating both internal (action error) and external feedback (Amiez et al., 2016). In this context, the left FOC might contribute to monitoring performance in perception and production tasks.

In summary, our study underscores the importance of the FOC in speech perception and production, shedding light on a region seldom explored in the literature. The FOC may play a role in housing speech representations crucial to these functions or contributing to a broader executive cognitive network. Our findings emphasize the necessity of further investigating the involvement of the FOC in speech processing.

4.1.2 The role of the PT

Like the left FOC, the left PT activation overlapped between speech production and the three listening conditions. However, unlike the FOC, there was no consistent region shared across all conditions, suggesting specificity in the overlap. This observation aligns with evidence proposing that the PT comprises multiple cortical fields with distinct functional roles (e.g., Isenberg et al., 2012; Tremblay et al., 2013).

The PT is a triangular region of the superior temporal plane located posterior to the primary auditory area (Economo & Horn, 1930). According to Galaburda and Sanides (1980), the anterior region of the PT is designated as the belt and parabelt regions, representing secondary and tertiary auditory areas. Numerous neuroimaging studies showed activation in PT during listening to a variety of sounds, including speech (e.g., Giraud & Price, 2001; Vouloumanos et al., 2001), tones (e.g., Binder et al., 1996), and melodies (e.g., Zatorre et al., 1994). Griffiths and Warren (2002) proposed that the PT acts as a computational hub, untangling complex sounds by isolating acoustic characteristics, and matching them with stored templates.

The PT's rearmost segment, called the temporoparietal area (Galaburda & Sanides, 1980), houses the Sylvian temporal parietal area (Spt). Neuroimaging studies revealed activation of the Spt when listening to speech and producing covert speech (e.g., Hickok et al., 2003; Isenberg et al., 2012; Okada & Hickok, 2006), which is thought to index sensorimotor integration. This integration is likely facilitated by the arcuate fasciculus, which establishes connections between the PT and frontal motor regions (Catani et al., 2002). Such sensorimotor integration, especially in challenging speech perception, aligns with the revised version of the Motor Theory of Speech Perception, where motor representations are proposed to enhance the perception of degraded and noisy speech signals. It is therefore possible that the PT houses two distinct mechanisms for processing complex auditory speech signals: the anterior part could be involved in general acoustic processing, while the Spt region could specialize in sensory‐motor integration.

Here, we found that spectrally degraded speech and speech production share common areas in the PT near the belt and parabelt regions. SIN activation overlaps with speech production activation, extending from the belt and parabelt regions to the Spt region, up to the ventral bank of the cortex. Linguistic complexity and speech production activation overlap from the Spt region to the ventral bank of the STG, with no apparent overlap in the belt and parabelt regions. These observations suggest that SIN and linguistic processing may rely on sensorimotor integrations. Conversely, spectrally degraded speech showed little overlap with speech production in the Spt region, which could indicate a greater reliance on acoustic processing cues to identify the speech sound. Finally, we also observed that the overlap for linguistic processing was concentrated in the ventral bank of the STG. This is consistent with previous research showing that language processing occurs mainly in the ventral part of the STG and general auditory processing in the PT (Binder et al., 1996).

4.2 Context‐dependent overlap

4.2.1 The role of the pre‐SMA

Compared to the previous two regions, the left pre‐SMA shows a context‐dependent overlap. The pre‐SMA was consistently activated during listening tasks involving SIN and spectrally degraded speech, while linguistically complex speech listening did not engage the pre‐SMA. This finding suggests that the pre‐SMA activation is primarily related to signal quality rather than linguistic aspects. Our results align with a meta‐analysis by Adank et al. (2012) showing an overlap between speech production and degraded, accented, and noisy speech perception in the bilateral pre‐SMA. However, our study underscores the context‐dependent nature of pre‐SMA activation.

Located in the medial superior frontal gyrus, the pre‐SMA is positioned upstream of the SMA‐proper (Picard & Strick, 2001). Conventionally, these two brain regions are not considered a part of the primary language network and are absent from most speech processing models (Hickok & Poeppel, 2007; Rauschecker & Scott, 2009). However, distinct functions have been attributed to the pre‐SMA and the SMA based on their connectivity profiles. The SMA‐proper is directly linked to the motor system, including the M1 cortex (Luppino et al., 1993), and is thought to play a role in the execution of movements. In contrast, the pre‐SMA lacks connections to the motor system but exhibits dense connectivity with the prefrontal cortex, including the IFG (Luppino et al., 1993). The pre‐SMA is believed to be involved in higher‐order aspects of actions, such as response selection (Tremblay & Gracco, 2009) and exerting control over voluntary actions (Nachev et al., 2007). For speech production, the pre‐SMA has been shown to exhibit activation across diverse tasks, including syllable sequence production (Bohland & Guenther, 2006), word generation (Tremblay & Gracco, 2006), and sentence production (Tremblay & Small, 2011b). More recently, Cummine et al. (2017) showed stronger pre‐SMA activation during reading tasks that require higher cognitive resources. For speech perception, the pre‐SMA exhibits responsiveness to syllables (e.g., Lee et al., 2012), words (e.g., Binder et al., 2008), and sentences (e.g., Tremblay & Small, 2011a). Notably, the engagement of the pre‐SMA during perception is influenced by the level of difficulty associated with comprehending speech. Its activity was found to increase with background noise (e.g., Du et al., 2014; Scott et al., 2004) and temporal compression of speech (Vagharchakian et al., 2012). Lima et al. (2016) proposed that pre‐SMA is part of the network linking auditory information to the corresponding motor programs, like the proposed role of the PT. This sensorimotor engagement may be influenced by processes controlled within a broader network involving prefrontal and temporal regions. Building on evidence indicating increased pre‐SMA activation when participants subjectively perceive physically interrupted words as continuous (Shahin et al., 2009), Lima et al. (2016) proposed that the sensorimotor engagement of the pre‐SMA might contribute to repairing degraded auditory information. In summary, the pre‐SMA, alongside the PT, could play a role in a sensorimotor mechanism that enhances speech comprehension, particularly in degraded and noisy conditions.

4.2.2 The role of the lateral portion of the inferior frontal gyrus

The second region showing context‐dependent overlap is the left pars opercularis of the IFG. The left IFG did not exhibit consistent activation in response to SIN and spectrally degraded speech. Increased activation was observed specifically in the presence of linguistically complex stimuli. In contrast to pre‐SMA, this suggests that the left IFG is more sensitive to linguistic variations than to signal quality. Moreover, it appears that the activation observed in the FOC during SIN and spectrally degraded speech extends to the left IFG when speech becomes linguistically complex. This extension suggests the involvement of additional resources in linguistic processing. This finding is consistent with numerous neuroimaging studies associating the left IFG with semantic, syntactic, and phonological processing in perceptual and production contexts (Caplan et al., 1998; Haller et al., 2005; Ishkhanyan et al., 2020), as well as with the perception and production of prosody (Aziz‐Zadeh et al., 2010). The left IFG could also be involved in response selection and resolving linguistic ambiguities in speech perception and production. This hypothesis is based on data showing the importance of the left IFG in selecting semantic information from competing alternatives during perception (Grindrod et al., 2008). In speech production, the left IFG may help overcome the interference of semantically‐related alternatives during word selection (Riès et al., 2015; Thompson‐Schill et al., 1997). A recent study has also showed the role of the left IFG in generating grammatically appropriate verbal responses at the sentence level (Ishkhanyan et al., 2020). Collectively, these results suggest that the left lateral part of the IFG could be essential for selecting the appropriate response or resolving conflicts between different speech units during speaking and listening. We propose that the activations of the lateral part of the IFG observed in studies of speech perception are due to the linguistic aspects of the task rather than to the difficulty of perceiving a degraded or noisy speech signal, the latter two being more likely to recruit the adjacent FOC.

4.3 The absence of primary motor and premotor areas for speech perception

One of the surprising results of this meta‐analysis is the minimal (almost nonexistent) contribution of M1 and PMv to difficult speech perception. This calls into question the idea that TMS‐induced modulation of M1 and PMv activity affects SIN performance. Although we cannot rule out the possibility that these regions play a role in perception, our results indicate that they are not an integral part of the network involved in the perception of difficult speech. It is possible that the specific contrast of interest in our meta‐analysis (difficult vs. clear speech) masked any contribution from M1 and PMv, as these regions could have been similarly recruited under normal and difficult conditions. Although previous studies have reported increased activity in M1 and PMv in response to noise, other data suggest that their involvement may be independent of noise levels (e.g., Glanz et al., 2018; Panouillères et al., 2018).

Another explanation could be the type of stimuli used in many of the studies included in the meta‐analysis. Most of the studies included employed tasks with lexical‐level stimuli, such as words and sentences. In contrast, studies that observed increased involvement of the speech motor system under difficult listening conditions often used tasks requiring phoneme discrimination or categorization. The involvement of M1 and PMv could therefore depend on the demand for phonological processing rather than the demand for listening. This interpretation would be consistent with the proposed role of the left PMv in segmenting the speech stream into constituent phonemes (e.g., Sato et al., 2009).

Another potential explanation for our findings relates to the temporal resolution limitations of fMRI and PET imaging. The temporal resolution of fMRI, for instance, ranges from approximately 1 to 4 seconds, which may not capture rapid neural dynamics effectively. Increasing evidence suggests that auditory‐motor integration occurs rapidly during speech perception (for a review, see Liebenthal & Möttönen, 2018), with the M1 and PMv being activated within a few hundred milliseconds of speech onset. Consequently, our meta‐analysis of fMRI and PET studies may have missed the overlap between speech production and perception in M1/PMv, given these temporal constraints. In contrast, MEG offers superior temporal resolution and could provide a more detailed view of these dynamics. Previous MEG studies have observed activity in M1 and PMv during both oro‐facial movement observation (e.g., Muthukumaraswamy et al., 2004) and syllable listening (e.g., Alho et al., 2012), even when no visible movement is present. However, it remains unclear whether activity in these regions increases under challenging listening conditions. For instance, Alho et al. (2012) found speech‐evoked activity in PMv activity 100 ms after syllable onset, which was stronger when syllables were embedded in noise than when there was no background noise. In contrast to these findings, our recent MEG results show that M1 and pSTG are activated simultaneously following syllable onset, with this activation remaining consistent regardless of SNRs (Perron et al., 2024).

5 LIMITATIONS AND FUTURE DIRECTIONS

Despite methodological variations between studies, including differences in tasks, control conditions, sample sizes, equipment and scanner types, meta‐analyses of neuroimaging studies are effective tools for understanding the brain regions that support task performance. Our analysis was performed using ALE software using the most recent and recommended meta‐analysis guidelines. Although we have carefully reviewed the literature, some studies may have been missed. In addition, the distribution of studies across the three listening conditions is uneven. This disparity could lead to lower statistical power and less reliable ALE scores in conditions with fewer studies, while conditions with more studies could exhibit higher activation due to larger data sets, potentially biasing comparisons. However, we are confident that each condition includes the recommended number of studies to ensure robust results with adequate statistical power (Eickhoff et al., 2016). It is also important to consider that although the same brain area may appear to respond to two tasks at the gross anatomical scale visible from fMRI, the neural tissue involved at a micro‐scale could be entirely distinct for each task (e.g., see Braga & Buckner, 2017).

Another limitation of this meta‐analysis is that we focused on only three conditions of difficult speech perception (SIN, spectrally degraded speech, and linguistically complex speech). The speech motor system may show differential overlaps with other types of listening manipulations, such as speech rate and accented speech; however, neuroimaging studies including these manipulations remain limited, making coordinate‐based meta‐analysis challenging. In addition, the networks involved in speech production and perception are also likely to overlap in different ways depending on the linguistic unit (e.g., phonemes, syllables, words, and sentences). Further research is needed to explore how these two processes interact as a function of the linguistic unit and task, incorporating considerations such as stimulus type and listening manipulation.

Another question for future studies is the influence of age on the overlap between speech perception and production. It is plausible that this overlap decreases with age, which could explain why older adults often experience difficulties perceiving speech, especially in noisy environments. Conversely, it is possible that this overlap increases with age, with activation of the motor system serving as a compensatory mechanism for age‐related SIN difficulties (Du et al., 2016). Our analyses focused on studies with young and middle‐aged adults. If the overlap decreases with age, the inclusion of middle‐aged adults may have reduced the observed overlap. However, the number of studies involving middle‐aged adults included in the meta‐analysis is small, which is unlikely to have influenced our results.

6 CONCLUSION

Our meta‐analysis shows that the speech motor network is recruited during difficult speech perception in context‐dependent and context‐independent ways. This overlap demonstrates the sharing of representations and resources between the two functions. Traditionally, the left motor and premotor cortices are central to this theory. However, our results challenge this notion by highlighting the role of other regions of the speech motor system in difficult speech perception. The left FOC, PT, pre‐SMA and lateral posterior IFG are crucial nodes for both processes. It should be noted that the left FOC stands out as the hub of the speech motor system for difficult speech perception, its posterior part being consistently activated under all listening conditions. We call for more precise labeling of brain regions in future studies to elucidate the distinct role of the FOC in language processing in relation to the neighboring lateral cortex and insula.

FUNDING INFORMATION

This work was supported by a grant to CA from the Natural Sciences and Engineering Research Council of Canada (grant number RGPIN‐2021‐02721). MP was funded by a Canadian Institutes of Health Research (CIHR) graduate scholarship. VV was funded by the Alzheimer Society of Canada.

CONFLICT OF INTEREST STATEMENT

The authors have no known conflict of interest to declare.

Supporting information

Data S1: Supporting Information.

ACKNOWLEDGMENT

We would like to thank all the individuals and organizations who indirectly contributed to the research and realization of this article. We also thank the reviewers and editors for their helpful comments. This work was supported by the Natural Sciences and Engineering Research Council of Canada.

DATA AVAILABILITY STATEMENT

The list of all coordinates as well as activations maps have been made publicly available at the Borealis Dataverse and can be accessed at https://doi.org/10.5683/SP3/KKC0RM. This meta‐analysis was not preregistered.
==== Refs
REFERENCES

* indicates studies included in the meta‐analysis.

Adank, P. (2012). The neural bases of difficult speech comprehension and speech production: Two activation likelihood estimation (ALE) meta‐analyses. Brain and Language, 122 (1 ), 42–54. 10.1016/j.bandl.2012.04.014 22633697
* Adank, P. , Davis, M. H. , & Hagoort, P. (2012). Neural dissociation in processing noise and accent in spoken language comprehension. Neuropsychologia, 50 (1 ), 77–84. 10.1016/j.neuropsychologia.2011.10.024 22085863
* Agmon, G. , Bain, J. S. , & Deschamps, I. (2021). Negative polarity in quantifiers evokes greater activation in language‐related regions compared to negative polarity in adjectives. Experimental Brain Research, 239 (5 ), 1427–1438. 10.1007/s00221-021-06067-y 33682044
* Agmon, G. , Yahav, P. H. , Ben‐Shachar, M. , & Golumbic, E. Z. (2022). Attention to speech: Mapping distributed and selective attention systems. Cerebral Cortex, 32 (17 ), 3763–3776. 10.1093/cercor/bhab446 34875678
Alain, C. , Du, Y. , Bernstein, L. J. , Barten, T. , & Banai, K. (2018). Listening under difficult conditions: An activation likelihood estimation meta‐analysis. Human Brain Mapping, 39 (7 ), 2695–2709. 10.1002/hbm.24031 29536592
Alho, J. , Sato, M. , Sams, M. , Schwartz, J. L. , Tiitinen, H. , & Jääskeläinen, I. P. (2012). Enhanced early‐latency electromagnetic activity in the left premotor cortex is associated with successful phonetic categorization. NeuroImage, 60 (4 ), 1937–1946. 10.1016/j.neuroimage.2012.02.011 22361165
* Allen, P. P. , Amaro, E. , Fu, C. H. , Williams, S. C. , Brammer, M. , Johns, L. C. , & McGuire, P. K. (2005). Neural correlates of the misattribution of self‐generated speech. Human Brain Mapping, 26 (1 ), 44–53. 10.1002/hbm.20120 15884023
Amiez, C. , Verstraete, C. , Sallet, J. , Hadj‐Bouziane, F. , Ben Hamed, S. , Meguerditchian, A. , Procyk, E. , Wilson, C. R. E. , Petrides, M. , Sherwood, C. C. , & Hopkins, W. D. (2023). The relevance of the unique anatomy of the human prefrontal operculum to the emergence of speech. Communications Biology, 6 (1 ), 693. 10.1038/s42003-023-05066-9 37407769
Amiez, C. , Wutte, M. G. , Faillenot, I. , Petrides, M. , Burle, B. , & Procyk, E. (2016). Single subject analyses reveal consistent recruitment of frontal operculum in performance monitoring. NeuroImage, 133 , 266–278. 10.1016/j.neuroimage.2016.03.003 26973171
Amunts, K. , Lenzen, M. , Friederici, A. D. , Schleicher, A. , Morosan, P. , Palomero‐Gallagher, N. , & Zilles, K. (2010). Broca's region: Novel organizational principles and multiple receptor mapping. PLoS Biology, 8 (9 ), e1000489. 10.1371/journal.pbio.1000489 20877713
Amunts, K. , Weiss, P. H. , Mohlberg, H. , Pieperhoff, P. , Eickhoff, S. , Gurd, J. M. , Marshall, J. C. , Shah, N. J. , Fink, G. R. , & Zilles, K. (2004). Analysis of neural mechanisms underlying verbal fluency in cytoarchitectonically defined stereotaxic space—The roles of Brodmann areas 44 and 45. NeuroImage, 22 (1 ), 42–56. 10.1016/j.neuroimage.2003.12.031 15109996
Anwander, A. , Tittgemeyer, M. , von Cramon, D. , Friederici, A. , & Knösche, T. (2006). Connectivity‐based parcellation of Broca's area. Cerebral Cortex, 17 (4 ), 816–825. 10.1093/cercor/bhk034 16707738
Aziz‐Zadeh, L. , Sheng, T. , & Gheytanchi, A. (2010). Common premotor regions for the perception and production of prosody and correlations with empathy and prosodic ability. PLoS One, 5 (1 ), e8759. 10.1371/journal.pone.0008759 20098696
* Bekinschtein, T. A. , Davis, M. H. , Rodd, J. M. , & Owen, A. M. (2011). Why clowns taste funny: The relationship between humor and semantic ambiguity. The Journal of Neuroscience, 31 (26 ), 9665–9671. 10.1523/JNEUROSCI.5058-10.2011 21715632
* Bilenko, N. Y. , Grindrod, C. M. , Myers, E. B. , & Blumstein, S. E. (2008). Neural correlates of semantic competition during processing of ambiguous words. Journal of Cognitive Neuroscience, 21 (5 ), 960–975. 10.1162/jocn.2009.21073
Binder, J. R. , Frost, J. A. , Hammeke, T. A. , Rao, S. M. , & Cox, R. W. (1996). Function of the left planum temporale in auditory and linguistic processing. Brain, 119 (4 ), 1239–1247. 10.1093/brain/119.4.1239 8813286
* Binder, J. R. , Liebenthal, E. , Possing, E. T. , Medler, D. A. , & Ward, B. D. (2004). Neural correlates of sensory and decision processes in auditory object identification. Nature Neuroscience, 7 (3 ), 295–301. 10.1038/nn1198 14966525
Binder, J. R. , Swanson, S. J. , Hammeke, T. A. , & Sabsevitz, D. S. (2008). A comparison of five fMRI protocols for mapping speech comprehension systems. Epilepsia, 49 (12 ), 1980–1997. 10.1111/j.1528-1167.2008.01683.x 18513352
Bohland, J. W. , & Guenther, F. H. (2006). An fMRI investigation of syllable sequence production. NeuroImage, 32 , 821–841. Retrieved from http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&list_uids=16730195&dopt=Abstract 16730195
Braga, R. M. , & Buckner, R. L. (2017). Parallel interdigitated distributed networks within the individual estimated by intrinsic functional connectivity. Neuron, 95 (2 ), 457–471.e455. 10.1016/j.neuron.2017.06.038 28728026
Brisson, V. , & Tremblay, P. (2021). Improving speech perception in noise in young and older adults using transcranial magnetic stimulation. Brain and Language, 222 , 105009. 10.1016/j.bandl.2021.105009 34425411
* Buchweitz, A. , Keller, T. A. , Meyler, A. , & Just, M. A. (2012). Brain activation for language dual‐tasking: Listening to two people speak at the same time and a change in network timing. Human Brain Mapping, 33 (8 ), 1868–1882. 10.1002/hbm.21327 21618666
* Buchweitz, A. , Mason, R. A. , Meschyan, G. , Keller, T. A. , & Just, M. A. (2014). Modulation of cortical activity during comprehension of familiar and unfamiliar text topics in speed reading and speed listening. Brain and Language, 139 , 49–57. 10.1016/j.bandl.2014.09.010 25463816
Callan, D. E. , Jones, J. A. , & Callan, A. (2014). Multisensory and modality specific processing of visual speech in different regions of the premotor cortex. Frontiers in Psychology, 5 , 389. 10.3389/fpsyg.2014.00389 24860526
Caplan, D. , Alpert, N. , & Waters, G. (1998). Effects of syntactic structure and propositional number on patterns of regional cerebral blood flow. Journal of Cognitive Neuroscience, 10 (4 ), 541–552. 10.1162/089892998562843 9712683
* Caplan, D. , Alpert, N. , & Waters, G. (1999). PET studies of syntactic processing with auditory sentence presentation. NeuroImage, 9 (3 ), 343–351. 10.1006/nimg.1998.0412 10075904
Catani, M. , Howard, R. J. , Pajevic, S. , & Jones, D. K. (2002). Virtual in vivo interactive dissection of white matter fasciculi in the human brain. NeuroImage, 17 (1 ), 77–94. 10.1006/nimg.2002.1136 12482069
Cummine, J. , Hanif, W. , Dymouriak‐Tymashov, I. , Anchuri, K. , Chiu, S. , & Boliek, C. A. (2017). The role of the supplementary motor region in overt reading: Evidence for differential processing in SMA‐proper and pre‐SMA as a function of task demands. Brain Topography, 30 (5 ), 579–591. 10.1007/s10548-017-0553-3 28260167
D'Ausilio, A. , Pulvermuller, F. , Salmas, P. , Bufalari, I. , Begliomini, C. , & Fadiga, L. (2009). The motor somatotopy of speech perception. Current Biology, 19 (5 ), 381–385. 10.1016/j.cub.2009.01.017 19217297
* Davis, M. H. , Coleman, M. R. , Absalom, A. R. , Rodd, J. M. , Johnsrude, I. S. , Matta, B. F. , Owen, A. M. , & Menon, D. K. (2007). Dissociating speech perception and comprehension at reduced levels of awareness. Proceedings of the National Academy of Sciences of the United States of America, 104 (41 ), 16032–16037. 10.1073/pnas.0701309104 17938125
* Deschamps, I. , & Tremblay, P. (2014). Sequencing at the syllabic and supra‐syllabic levels during speech perception: An fMRI study. Frontiers in Human Neuroscience, 8 , 14. 10.3389/fnhum.2014.00492 24478680
* Dos Santos Sequeira, S. , Specht, K. , Moosmann, M. , Westerhausen, R. , & Hugdahl, K. (2010). The effects of background noise on dichotic listening to consonant‐vowel syllables: An fMRI study. Laterality, 15 (6 ), 577–596. 10.1080/13576500903045082 19626537
* Du, Y. , Buchsbaum, B. R. , Grady, C. L. , & Alain, C. (2014). Noise differentially impacts phoneme representations in the auditory and speech motor systems. Proceedings of the National Academy of Sciences of the United States of America, 111 (19 ), 7126–7131. 10.1073/pnas.1318738111 24778251
Du, Y. , Buchsbaum, B. R. , Grady, C. L. , & Alain, C. (2016). Increased activity in frontal motor cortex compensates impaired speech perception in older adults. Nature Communications, 7 , 12241. 10.1038/ncomms12241
* Eckert, M. A. , Menon, V. , Walczak, A. , Ahlstrom, J. , Denslow, S. , Horwitz, A. , & Dubno, J. R. (2009). At the heart of the ventral attention system: The right anterior insula. Human Brain Mapping, 30 (8 ), 2530–2541. 10.1002/hbm.20688 19072895
Economo, C. V. , & Horn, L. (1930). Über Windungsrelief, Maße und Rindenarchitektonik der Supratemporalfläche, ihre individuellen und ihre Seitenunterschiede. Zeitschrift für Die Gesamte Neurologie Und Psychiatrie, 130 (1 ), 678–757. 10.1007/BF02865945
Eickhoff, S. B. , Bzdok, D. , Laird, A. R. , Kurth, F. , & Fox, P. T. (2012). Activation likelihood estimation meta‐analysis revisited. NeuroImage, 59 (3 ), 2349–2361. 10.1016/j.neuroimage.2011.09.017 21963913
Eickhoff, S. B. , Heim, S. , Zilles, K. , & Amunts, K. (2006). Testing anatomically specified hypotheses in functional imaging using cytoarchitectonic maps. NeuroImage, 32 (2 ), 570–582. 10.1016/j.neuroimage.2006.04.204 16781166
Eickhoff, S. B. , Laird, A. R. , Grefkes, C. , Wang, L. E. , Zilles, K. , & Fox, P. T. (2009). Coordinate‐based activation likelihood estimation meta‐analysis of neuroimaging data: A random‐effects approach based on empirical estimates of spatial uncertainty. Human Brain Mapping, 30 (9 ), 2907–2926. 10.1002/hbm.20718 19172646
Eickhoff, S. B. , Nichols, T. E. , Laird, A. R. , Hoffstaedter, F. , Amunts, K. , Fox, P. T. , Bzdok, D. , & Eickhoff, C. R. (2016). Behavior, sensitivity, and power of activation likelihood estimation characterized by massive empirical simulation. NeuroImage, 137 , 70–85. 10.1016/j.neuroimage.2016.04.072 27179606
Eickhoff, S. B. , Stephan, K. E. , Mohlberg, H. , Grefkes, C. , Fink, G. R. , Amunts, K. , & Zilles, K. (2005). A new SPM toolbox for combining probabilistic cytoarchitectonic maps and functional imaging data. NeuroImage, 25 (4 ), 1325–1335. 10.1016/j.neuroimage.2004.12.034 15850749
* Erb, J. , Henry, M. J. , Eisner, F. , & Obleser, J. (2013). The brain dynamics of rapid perceptual adaptation to adverse listening conditions. The Journal of Neuroscience, 33 (26 ), 10688–10697. 10.1523/JNEUROSCI.4596-12.2013 23804092
* Evans, S. , & Davis, M. H. (2015). Hierarchical organization of auditory and motor representations in speech perception: Evidence from searchlight similarity analysis. Cerebral Cortex, 25 (12 ), 4772–4788. 10.1093/cercor/bhv136 26157026
* Evans, S. , McGettigan, C. , Agnew, Z. K. , Rosen, S. , & Scott, S. K. (2016). Getting the cocktail party started: Masking effects in speech perception. Journal of Cognitive Neuroscience, 28 (3 ), 483–500. 10.1162/jocn_a_00913 26696297
* Ferstl, E. C. , Rinck, M. , & von Cramon, D. Y. (2005). Emotional and temporal aspects of situation model processing during text comprehension: An event‐related fMRI study. Journal of Cognitive Neuroscience, 17 (5 ), 724–739. 10.1162/0898929053747658 15904540
Friederici, A. D. (2011). The brain basis of language processing: From structure to function. Physiological Reviews, 91 (4 ), 1357–1392. 10.1152/physrev.00006.2011 22013214
Friederici, A. D. , Fiebach, C. J. , Schlesewsky, M. , Bornkessel, I. D. , & von Cramon, D. Y. (2006). Processing linguistic complexity and grammaticality in the left frontal cortex. Cerebral Cortex, 16 (12 ), 1709–1717. 10.1093/cercor/bhj106 16400163
* Friederici, A. D. , Kotz, S. A. , Scott, S. K. , & Obleser, J. (2010). Disentangling syntax and intelligibility in auditory language comprehension. Human Brain Mapping, 31 (3 ), 448–457. 10.1002/hbm.20878 19718654
* Gaab, N. , Gabrieli, J. D. , & Glover, G. H. (2007). Assessing the influence of scanner background noise on auditory processing. II. An fMRI study comparing auditory processing in the absence and presence of recorded scanner noise using a sparse design. Human Brain Mapping, 28 (8 ), 721–732. 10.1002/hbm.20299 17089376
Galaburda, A. , & Sanides, F. (1980). Cytoarchitectonic organization of the human auditory cortex. Journal of Comparative Neurology, 190 (3 ), 597–610. 10.1002/cne.901900312 6771305
Gallese, V. , Fadiga, L. , Fogassi, L. , & Rizzolatti, G. (1996). Action recognition in the premotor cortex. Brain, 119 (Pt 2 ), 593–609. 10.1093/brain/119.2.593 8800951
Giraud, A. L. , & Price, C. J. (2001). The constraints functional neuroimaging places on classical models of auditory word processing. Journal of Cognitive Neuroscience, 13 (6 ), 754–765. 10.1162/08989290152541421 11564320
Glanz, O. , Derix, J. , Kaur, R. , Schulze‐Bonhage, A. , Auer, P. , Aertsen, A. , & Ball, T. (2018). Real‐life speech production and perception have a shared premotor‐cortical substrate. Scientific Reports, 8 (1 ), 8898. 10.1038/s41598-018-26801-x 29891885
* Golestani, N. , Hervais‐Adelman, A. , Obleser, J. , & Scott, S. K. (2013). Semantic versus perceptual interactions in neural processing of speech‐in‐noise. NeuroImage, 79 , 52–61. 10.1016/j.neuroimage.2013.04.049 23624171
Griffiths, T. D. , & Warren, J. D. (2002). The planum temporale as a computational hub. Trends in Neurosciences, 25 (7 ), 348–353. 10.1016/S0166-2236(02)02191-4 12079762
* Grindrod, C. M. , Bilenko, N. Y. , Myers, E. B. , & Blumstein, S. E. (2008). The role of the left inferior frontal gyrus in implicit semantic competition and selection: An event‐related fMRI study. Brain Research, 1229 , 167–178. 10.1016/j.brainres.2008.07.017 18656462
* Guediche, S. , Salvata, C. , & Blumstein, S. E. (2013). Temporal cortex reflects effects of sentence context on phonetic processing. Journal of Cognitive Neuroscience, 25 (5 ), 706–718. 10.1162/jocn_a_00351 23281778
* Guerreiro, M. J. , Putzar, L. , & Röder, B. (2016). The effect of early visual deprivation on the neural bases of auditory processing. The Journal of Neuroscience, 36 (5 ), 1620–1630. 10.1523/jneurosci.2559-15.2016 26843643
Haller, S. , Radue, E. W. , Erb, M. , Grodd, W. , & Kircher, T. (2005). Overt sentence production in event‐related fMRI. Neuropsychologia, 43 (5 ), 807–814. 10.1016/j.neuropsychologia.2004.09.007 15721193
* Herrmann, B. , Obleser, J. , Kalberlah, C. , Haynes, J. D. , & Friederici, A. D. (2012). Dissociable neural imprints of perception and grammar in auditory functional imaging. Human Brain Mapping, 33 (3 ), 584–595. 10.1002/hbm.21235 21391281
* Hervais‐Adelman, A. , Pefkou, M. , & Golestani, N. (2014). Bilingual speech‐in‐noise: Neural bases of semantic context use in the native language. Brain and Language, 132 , 1–6. 10.1016/j.bandl.2014.01.009 24594855
* Hervais‐Adelman, A. G. , Carlyon, R. P. , Johnsrude, I. S. , & Davis, M. H. (2012). Brain regions recruited for the effortful comprehension of noise‐vocoded words. Language and Cognitive Processes, 27 (7–8 ), 1145–1166. 10.1080/01690965.2012.662280
* Hesling, I. , Clément, S. , Bordessoules, M. , & Allard, M. (2005). Cerebral mechanisms of prosodic integration: Evidence from connected speech. NeuroImage, 24 (4 ), 937–947. 10.1016/j.neuroimage.2004.11.003 15670670
Hickok, G. , Buchsbaum, B. , Humphries, C. , & Muftuler, T. (2003). Auditory‐motor interaction revealed by fMRI: Speech, music, and working memory in area Spt. Journal of Cognitive Neuroscience, 15 (5 ), 673–682. 10.1162/089892903322307393 12965041
Hickok, G. , Costanzo, M. , Capasso, R. , & Miceli, G. (2011). The role of Broca's area in speech perception: Evidence from aphasia revisited. Brain and Language, 119 (3 ), 214–220. 10.1016/j.bandl.2011.08.001 21920592
Hickok, G. , & Poeppel, D. (2007). The cortical organization of speech processing. Nature Reviews. Neuroscience, 8 (5 ), 393–402. 10.1038/nrn2113 17431404
* Hillert, D. G. , & Buračas, G. T. (2009). The neural substrates of spoken idiom comprehension. Language and Cognitive Processes, 24 (9 ), 1370–1391. 10.1080/01690960903057006
* Holmes, E. , & Johnsrude, I. S. (2021). Speech‐evoked brain activity is more robust to competing speech when it is spoken by someone familiar. NeuroImage, 237 , 118107. 10.1016/j.neuroimage.2021.118107 33933598
Isenberg, A. L. , Vaden, K. I., Jr. , Saberi, K. , Muftuler, L. T. , & Hickok, G. (2012). Functionally distinct regions for spatial processing and sensory motor integration in the planum temporale. Human Brain Mapping, 33 (10 ), 2453–2463. 10.1002/hbm.21373 21932266
Ishkhanyan, B. , Michel Lange, V. , Boye, K. , Mogensen, J. , Karabanov, A. , Hartwigsen, G. , & Siebner, H. R. (2020). Anterior and posterior left inferior frontal gyrus contribute to the implementation of grammatical determiners during language production [original research]. Frontiers in Psychology, 11 , 1–13. 10.3389/fpsyg.2020.00685
Kohler, E. , Keysers, C. , Umiltà, M. A. , Fogassi, L. , Gallese, V. , & Rizzolatti, G. (2002). Hearing sounds, understanding actions: Action representation in mirror neurons. Science, 297 (5582 ), 846–848. 10.1126/science.1070311 12161656
* Kotz, S. A. , Cappa, S. F. , von Cramon, D. Y. , & Friederici, A. D. (2002). Modulation of the lexical‐semantic network by auditory semantic priming: An event‐related functional MRI study. NeuroImage, 17 (4 ), 1761–1772. 10.1006/nimg.2002.1316 12498750
* Kousaie, S. , Baum, S. , Phillips, N. A. , Gracco, V. , Titone, D. , Chen, J. K. , Chai, X. J. , & Klein, D. (2019). Language learning experience and mastering the challenges of perceiving speech in noise. Brain and Language, 196 , 104645. 10.1016/j.bandl.2019.104645 31284145
* Kristensen, L. B. , Engberg‐Pedersen, E. , & Wallentin, M. (2014). Context predicts word order processing in Broca's region. Journal of Cognitive Neuroscience, 26 (12 ), 2762–2777. 10.1162/jocn_a_00681 25000525
* Kyong, J. S. , Scott, S. K. , Rosen, S. , Howe, T. B. , Agnew, Z. K. , & McGettigan, C. (2014). Exploring the roles of spectral detail and intonation contour in speech intelligibility: An FMRI study. Journal of Cognitive Neuroscience, 26 (8 ), 1748–1763. 10.1162/jocn_a_00583 24568205
Lancaster, J. L. , Rainey, L. H. , Summerlin, J. L. , Freitas, C. S. , Fox, P. T. , Evans, A. C. , Toga, A. W. , & Mazziotta, J. C. (1997). Automated labeling of the human brain: A preliminary report on the development and evaluation of a forward‐transform method. Human Brain Mapping, 5 (4 ), 238–242. 10.1002/(sici)1097-0193(1997)5:4<238::aid-hbm6>3.0.co;2-4 20408222
Lancaster, J. L. , Tordesillas‐Gutiérrez, D. , Martinez, M. , Salinas, F. , Evans, A. , Zilles, K. , Mazziotta, J. C. , & Fox, P. T. (2007). Bias between MNI and Talairach coordinates analyzed using the ICBM‐152 brain template. Human Brain Mapping, 28 (11 ), 1194–1205. 10.1002/hbm.20345 17266101
Lancaster, J. L. , Woldorff, M. G. , Parsons, L. M. , Liotti, M. , Freitas, C. S. , Rainey, L. , Kochunov, P. V. , Nickerson, D. , Mikiten, S. A. , & Fox, P. T. (2000). Automated Talairach atlas labels for functional brain mapping. Human Brain Mapping, 10 (3 ), 120–131. 10.1002/1097-0193(200007)10:3<120::aid-hbm30>3.0.co;2-8 10912591
* Lee, Y. S. , Min, N. E. , Wingfield, A. , Grossman, M. , & Peelle, J. E. (2016). Acoustic richness modulates the neural networks supporting intelligible speech processing. Hearing Research, 333 , 108–117. 10.1016/j.heares.2015.12.008 26723103
Lee, Y.‐S. , Turkeltaub, P. , Granger, R. , & Raizada, R. D. S. (2012). Categorical speech processing in Broca's area: An fMRI study using multivariate pattern‐based analysis. The Journal of Neuroscience, 32 (11 ), 3942–3948. 10.1523/jneurosci.3814-11.2012 22423114
* Lee, Y. S. , Wingfield, A. , Min, N. E. , Kotloff, E. , Grossman, M. , & Peelle, J. E. (2018). Differences in hearing acuity among “normal‐hearing” young adults modulate the neural basis for speech comprehension. eNeuro, 5 (3 ), 1–12. 10.1523/ENEURO.0263-17.2018
Liberman, A. M. , Cooper, F. S. , Shankweiler, D. P. , & Studdert‐Kennedy, M. (1967). Perception of the speech code. Psychological Review, 74 (6 ), 431–461. 10.1037/h0020279 4170865
Liberman, A. M. , & Mattingly, I. G. (1985). The motor theory of speech perception revised. Cognition, 21 (1 ), 1–36. 10.1016/0010-0277(85)90021-6 4075760
Liebenthal, E. , & Möttönen, R. (2018). An interactive model of auditory‐motor speech perception. Brain and Language, 187 , 33–40. 10.1016/j.bandl.2017.12.004 29268943
Lima, C. F. , Krishnan, S. , & Scott, S. K. (2016). Roles of supplementary motor areas in auditory processing and auditory imagery. Trends in Neurosciences, 39 (8 ), 527–542. 10.1016/j.tins.2016.06.003 27381836
* Lin, Y. , Tsao, Y. , & Hsieh, P. J. (2022). Neural correlates of individual differences in predicting ambiguous sounds comprehension level. NeuroImage, 251 , 119012. 10.1016/j.neuroimage.2022.119012 35183745
* Lopes, T. M. , Yasuda, C. L. , Campos, B. M. , Balthazar, M. L. F. , Binder, J. R. , & Cendes, F. (2016). Effects of task complexity on activation of language areas in a semantic decision fMRI protocol. Neuropsychologia, 81 , 140–148. 10.1016/j.neuropsychologia.2015.12.020 26721760
Luppino, G. , Matelli, M. , Camarda, R. , & Rizzolatti, G. (1993). Corticocortical connections of area F3 (SMA‐proper) and area F6 (pre‐SMA) in the macaque monkey. The Journal of Comparative Neurology, 338 (1 ), 114–140. 10.1002/cne.903380109 7507940
* Lyu, B. , Ge, J. , Niu, Z. , Tan, L. H. , & Gao, J. H. (2016). Predictive brain mechanisms in sound‐to‐meaning mapping during speech processing. The Journal of Neuroscience, 36 (42 ), 10813–10822. 10.1523/JNEUROSCI.0583-16.2016 27798136
Mălîia, M. D. , Donos, C. , Barborica, A. , Popa, I. , Ciurea, J. , Cinatti, S. , & Mîndruţă, I. (2018). Functional mapping and effective connectivity of the human operculum. Cortex, 109 , 303–321. 10.1016/j.cortex.2018.08.024 30414541
* Manan, H. , Franz, L. , Yusoff, A. , & Mukari, S. (2012). Hippocampal‐cerebellar involvement in enhancement of performance in word‐based BRT with the presence of background noise: An initial fMRI study. Psychology & Neuroscience, 5 , 247–256. 10.3922/j.psns.2012.2.16
Matchin, W. , Groulx, K. , & Hickok, G. (2014). Audiovisual speech integration does not rely on the motor system: Evidence from articulatory suppression, the McGurk effect, and fMRI. Journal of Cognitive Neuroscience, 26 (3 ), 606–620. 10.1162/jocn_a_00515 24236768
Mayka, M. A. , Corcos, D. M. , Leurgans, S. E. , & Vaillancourt, D. E. (2006). Three‐dimensional locations and boundaries of motor and premotor cortices as defined by functional brain imaging: A meta‐analysis. NeuroImage, 31 (4 ), 1453–1474. 10.1016/j.neuroimage.2006.02.004 16571375
McGettigan, C. , & Tremblay, P. (2018). In S.‐A. Rueschemeyer & M. G. Gaskell (Eds.), Links between perception and production examining the roles of motor and premotor cortices in understanding speech. Oxford University Press. 10.1093/oxfordhb/9780198786825.013.14
McHugh, M. L. (2012). Interrater reliability: The kappa statistic. Biochemia Medicine, 22 (3 ), 276–282.
* Mechtenberg, H. , Xie, X. , & Myers, E. B. (2021). Sentence predictability modulates cortical response to phonetic ambiguity. Brain and Language, 218 , 104959. 10.1016/j.bandl.2021.104959 33930722
Meister, I. G. , Wilson, S. M. , Deblieck, C. , Wu, A. D. , & Iacoboni, M. (2007). The essential role of premotor cortex in speech perception. Current Biology, 17 (19 ), 1692–1696. 10.1016/j.cub.2007.08.064 17900904
* Mellet, E. , Tzourio, N. , Denis, M. , & Mazoyer, B. (1998). Cortical anatomy of mental imagery of concrete nouns based on their dictionary definition. Neuroreport, 9 (5 ), 803–808. 10.1097/00001756-199803300-00007 9579669
* Meltzer, J. A. , McArdle, J. J. , Schafer, R. J. , & Braun, A. R. (2010). Neural aspects of sentence comprehension: Syntactic complexity, reversibility, and reanalysis. Cerebral Cortex, 20 (8 ), 1853–1864. 10.1093/cercor/bhp249 19920058
* Meyer, L. , Obleser, J. , Anwander, A. , & Friederici, A. D. (2012). Linking ordering in Broca's area to storage in left temporo‐parietal regions: The case of sentence processing. NeuroImage, 62 (3 ), 1987–1998. 10.1016/j.neuroimage.2012.05.052 22634860
* Meyer, M. , Alter, K. , Friederici, A. D. , Lohmann, G. , & von Cramon, D. Y. (2002). fMRI reveals brain regions mediating slow prosodic modulations in spoken sentences. Human Brain Mapping, 17 (2 ), 73–88. 10.1002/hbm.10042 12353242
* Meyer, M. , Steinhauer, K. , Alter, K. , Friederici, A. D. , & von Cramon, D. Y. (2004). Brain activity varies with modulation of dynamic pitch variance in sentence melody. Brain and Language, 89 (2 ), 277–289. 10.1016/S0093-934X(03)00350-X 15068910
Mottonen, R. , Rogers, J. , & Watkins, K. E. (2014). Stimulating the lip motor cortex with transcranial magnetic stimulation. Journal of Visualized Experiments, 88 , e51665. 10.3791/51665
Muthukumaraswamy, S. D. , Johnson, B. , Gaetz, W. C. , & Cheyne, D. O. (2004). Neuromagnetic imaging reveals primary motor cortex activation during the observation of oro‐facial movements. Neurology and Clinical Neurophysiology, 2 , 5–8.
Nachev, P. , Wydell, H. , O'Neill, K. , Husain, M. , & Kennard, C. (2007). The role of the pre‐supplementary motor area in the control of action. NeuroImage, 36 (Suppl 2 ), T155–T163. 10.1016/j.neuroimage.2007.03.034 17499162
* Nagels, A. , Kauschke, C. , Schrauf, J. , Whitney, C. , Straube, B. , & Kircher, T. (2013). Neural substrates of figurative language during natural speech perception: An fMRI study. Frontiers in Behavioral Neuroscience, 7 , 8. 10.3389/fnbeh.2013.00121 23407621
* Obleser, J. , & Kotz, S. A. (2010). Expectancy constraints in degraded speech modulate the language comprehension network. Cerebral Cortex, 20 (3 ), 633–640. 10.1093/cercor/bhp128 19561061
* Obleser, J. , Meyer, L. , & Friederici, A. D. (2011). Dynamic assignment of neural resources in auditory comprehension of complex sentences. NeuroImage, 56 (4 ), 2310–2320. 10.1016/j.neuroimage.2011.03.035 21421059
* Obleser, J. , Zimmermann, J. , Van Meter, J. , & Rauschecker, J. P. (2007). Multiple stages of auditory speech perception reflected in event‐related fMRI. Cerebral Cortex, 17 (10 ), 2251–2257. 10.1093/cercor/bhl133 17150986
Okada, K. , & Hickok, G. (2006). Left posterior auditory‐related cortices participate both in speech perception and speech production: Neural overlap revealed by fMRI. Brain and Language, 98 (1 ), 112–117. 10.1016/j.bandl.2006.04.006 16716388
* Olano, M. A. , Elizalde Acevedo, B. , Chambeaud, N. , Acuña, A. , Marcó, M. , Kochen, S. , & Alba‐Ferrara, L. (2020). Emotional salience enhances intelligibility in adverse acoustic conditions. Neuropsychologia, 147 , 107580. 10.1016/j.neuropsychologia.2020.107580 32827539
Osnes, B. , Hugdahl, K. , & Specht, K. (2011). Effective connectivity analysis demonstrates involvement of premotor cortex during speech perception. NeuroImage, 54 (3 ), 2437–2445. 10.1016/j.neuroimage.2010.09.078 20932914
Page, M. J. , McKenzie, J. E. , Bossuyt, P. M. , Boutron, I. , Hoffmann, T. C. , Mulrow, C. D. , Shamseer, L. , Tetzlaff, J. M. , Akl, E. A. , Brennan, S. E. , Chou, R. , Glanville, J. , Grimshaw, J. M. , Hróbjartsson, A. , Lalu, M. M. , Li, T. , Loder, E. W. , Mayo‐Wilson, E. , McDonald, S. , McGuinness, L. A , … Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372 , n71. 10.1136/bmj.n71 33782057
Panouillères, M. T. N. , Boyles, R. , Chesters, J. , Watkins, K. E. , & Möttönen, R. (2018). Facilitation of motor excitability during listening to spoken sentences is not modulated by noise or semantic coherence. Cortex, 103 , 44–54. 10.1016/j.cortex.2018.02.007 29554541
* Peelle, J. E. , Eason, R. J. , Schmitter, S. , Schwarzbauer, C. , & Davis, M. H. (2010). Evaluating an acoustically quiet EPI sequence for use in fMRI studies of speech and auditory processing. NeuroImage, 52 (4 ), 1410–1419. 10.1016/j.neuroimage.2010.05.015 20483377
* Peelle, J. E. , McMillan, C. , Moore, P. , Grossman, M. , & Wingfield, A. (2004). Dissociable patterns of brain activity during comprehension of rapid and syntactically complex speech: Evidence from fMRI. Brain and Language, 91 (3 ), 315–325. 10.1016/j.bandl.2004.05.007 15533557
Perron, M. , Ross, B. , & Alain, C. (2024). Left motor cortex contributes to auditory phonological discrimination. Cerebral Cortex. 10.1093/cercor/bhae369
Picard, N. , & Strick, P. L. (2001). Imaging the premotor areas. Current Opinion in Neurobiology, 11 (6 ), 663–672. 10.1016/s0959-4388(01)00266-5 11741015
* Planton, S. , Chanoine, V. , Sein, J. , Anton, J. L. , Nazarian, B. , Pallier, C. , & Pattamadilok, C. (2019). Top‐down activation of the visuo‐orthographic system during spoken sentence processing. NeuroImage, 202 , 116135. 10.1016/j.neuroimage.2019.116135 31470125
Pulvermüller, F. , Huss, M. , Kherif, F. , del Prado, M. , Martin, F. , Hauk, O. , & Shtyrov, Y. (2006). Motor cortex maps articulatory features of speech sounds. Proceedings of the National Academy of Sciences of the United States of America, 103 (20 ), 7865–7870. 10.1073/pnas.0509989103 16682637
* Raettig, T. , Frisch, S. , Friederici, A. D. , & Kotz, S. A. (2010). Neural correlates of morphosyntactic and verb‐argument structure processing: An EfMRI study. Cortex, 46 (5 ), 613–620. 10.1016/j.cortex.2009.06.003 19664766
* Rammell, C. S. , Cheng, H. , Pisoni, D. B. , & Newman, S. D. (2019). L2 speech perception in noise: An fMRI study of advanced Spanish learners. Brain Research, 1720 , 146316. 10.1016/j.brainres.2019.146316 31278936
Rauschecker, J. P. , & Scott, S. K. (2009). Maps and streams in the auditory cortex: Nonhuman primates illuminate human speech processing. Nature Neuroscience, 12 (6 ), 718–724. 10.1038/nn.2331 19471271
Riès, S. K. , Karzmark, C. R. , Navarrete, E. , Knight, R. T. , & Dronkers, N. F. (2015). Specifying the role of the left prefrontal cortex in word selection. Brain and Language, 149 , 135–147. 10.1016/j.bandl.2015.07.007 26291289
* Rissman, J. , Eliassen, J. C. , & Blumstein, S. E. (2003). An event‐related FMRI investigation of implicit semantic priming. Journal of Cognitive Neuroscience, 15 (8 ), 1160–1175. 10.1162/089892903322598120 14709234
Rizzolatti, G. , & Craighero, L. (2004). The mirror‐neuron system. Annual Review of Neuroscience, 27 , 169–192. 10.1146/annurev.neuro.27.070203.144230
* Rodd, J. M. , Davis, M. H. , & Johnsrude, I. S. (2005). The neural mechanisms of speech comprehension: fMRI studies of semantic ambiguity. Cerebral Cortex, 15 (8 ), 1261–1269. 10.1093/cercor/bhi009 15635062
* Rodd, J. M. , Longe, O. A. , Randall, B. , & Tyler, L. K. (2010). The functional organisation of the fronto‐temporal language system: Evidence from syntactic and semantic ambiguity. Neuropsychologia, 48 (5 ), 1324–1335. 10.1016/j.neuropsychologia.2009.12.035 20038434
* Rothermich, K. , & Kotz, S. A. (2013). Predictions in speech comprehension: fMRI evidence on the meter‐semantic interface. NeuroImage, 70 , 89–100. 10.1016/j.neuroimage.2012.12.013 23291188
* Ruff, I. , Blumstein, S. E. , Myers, E. B. , & Hutchison, E. (2008). Recruitment of anterior and posterior structures in lexical‐semantic processing: An fMRI study comparing implicit and explicit tasks. Brain and Language, 105 (1 ), 41–49. 10.1016/j.bandl.2008.01.003 18279947
* Rüschemeyer, S. A. , Fiebach, C. J. , Kempe, V. , & Friederici, A. D. (2005). Processing lexical semantic and syntactic information in first and second language: fMRI evidence from German and Russian. Human Brain Mapping, 25 (2 ), 266–286. 10.1002/hbm.20098 15849713
* Rysop, A. U. , Schmitt, L. M. , Obleser, J. , & Hartwigsen, G. (2021). Neural modelling of the semantic predictability gain under challenging listening conditions. Human Brain Mapping, 42 (1 ), 110–127. 10.1002/hbm.25208 32959939
* Sabri, M. , Binder, J. R. , Desai, R. , Medler, D. A. , Leitl, M. D. , & Liebenthal, E. (2008). Attentional and linguistic interactions in speech perception. NeuroImage, 39 (3 ), 1444–1456. 10.1016/j.neuroimage.2007.09.052 17996463
* Salvi, R. J. , Lockwood, A. H. , Frisina, R. D. , Coad, M. L. , Wack, D. S. , & Frisina, D. R. (2002). PET imaging of the normal human auditory system: Responses to speech in quiet and in background noise. Hearing Research, 170 (1–2 ), 96–106. 10.1016/s0378-5955(02)00386-6 12208544
Sato, M. , Tremblay, P. , & Gracco, V. L. (2009). A mediating role of the premotor cortex in phoneme segmentation. Brain and Language, 111 (1 ), 1–7. 10.1016/j.bandl.2009.03.002 19362734
* Scott, S. K. , Rosen, S. , Wickham, L. , & Wise, R. J. (2004). A positron emission tomography study of the neural basis of informational and energetic masking effects in speech perception. The Journal of the Acoustical Society of America, 115 (2 ), 813–821. 10.1121/1.1639336 15000192
Seeley, W. W. , Menon, V. , Schatzberg, A. F. , Keller, J. , Glover, G. H. , Kenna, H. , Reiss, A. L., & Greicius, M. D. (2007). Dissociable intrinsic connectivity networks for salience processing and executive control. The Journal of Neuroscience, 27 (9 ), 2349–2356. 10.1523/jneurosci.5587-06.2007 17329432
* Sekiyama, K. , Kanno, I. , Miura, S. , & Sugita, Y. (2003). Auditory‐visual speech perception examined by fMRI and PET. Neuroscience Research, 47 (3 ), 277–287. 10.1016/s0168-0102(03)00214-1 14568109
Shahin, A. J. , Bishop, C. W. , & Miller, L. M. (2009). Neural mechanisms for illusory filling‐in of degraded speech. NeuroImage, 44 (3 ), 1133–1143. 10.1016/j.neuroimage.2008.09.045 18977448
* Sharp, D. J. , Awad, M. , Warren, J. E. , Wise, R. J. , Vigliocco, G. , & Scott, S. K. (2010). The neural response to changing semantic and perceptual complexity during language processing. Human Brain Mapping, 31 (3 ), 365–377. 10.1002/hbm.20871 19777554
* Shetreet, E. , Friedmann, N. , & Hadar, U. (2009). An fMRI study of syntactic layers: Sentential and lexical aspects of embedding. NeuroImage, 48 (4 ), 707–716. 10.1016/j.neuroimage.2009.07.001 19595775
* Siebörger, F. T. , Ferstl, E. C. , & von Cramon, D. Y. (2007). Making sense of nonsense: An fMRI study of task induced inference processes during discourse comprehension. Brain Research, 1166 , 77–91. 10.1016/j.brainres.2007.05.079 17655831
Smalle, E. H. M. , Rogers, J. , & Möttönen, R. (2015). Dissociating contributions of the motor cortex to speech perception and response bias by using transcranial magnetic stimulation. Cerebral Cortex, 25 (10 ), 3690–3698. 10.1093/cercor/bhu218 25274987
Stasenko, A. , Garcea, F. E. , & Mahon, B. Z. (2013). What happens to the motor theory of perception when the motor system is damaged? Language and Cognition, 5 (2–3 ), 225–238. 10.1515/langcog-2013-0016 26823687
* Takeichi, H. , Koyama, S. , Terao, A. , Takeuchi, F. , Toyosawa, Y. , & Murohashi, H. (2010). Comprehension of degraded speech sounds with m‐sequence modulation: An fMRI study. NeuroImage, 49 (3 ), 2697–2706. 10.1016/j.neuroimage.2009.10.063 19878726
Thompson‐Schill, S. L. , D'Esposito, M. , Aguirre, G. K. , & Farah, M. J. (1997). Role of left inferior prefrontal cortex in retrieval of semantic knowledge: A reevaluation. Proceedings of the National Academy of Sciences of the United States of America, 94 (26 ), 14792–14797. 10.1073/pnas.94.26.14792 9405692
Tremblay, P. , Deschamps, I. , & Gracco, V. L. (2013). Regional heterogeneity in the processing and the production of speech in the human planum temporale. Cortex, 49 (1 ), 143–157. 10.1016/j.cortex.2011.09.004 22019203
Tremblay, P. , & Gracco, V. L. (2006). Contribution of the frontal lobe to externally and internally specified verbal responses: fMRI evidence. NeuroImage, 33 (3 ), 947–957. 10.1016/j.neuroimage.2006.07.041 16990015
Tremblay, P. , & Gracco, V. L. (2009). Contribution of the pre‐SMA to the production of words and non‐speech oral motor gestures, as revealed by repetitive transcranial magnetic stimulation (rTMS). Brain Research, 1268 , 112–124. 10.1016/j.brainres.2009.02.076 19285972
Tremblay, P. , & Small, S. L. (2011a). From language comprehension to action understanding and back again. Cerebral Cortex, 21 (5 ), 1166–1177. 10.1093/cercor/bhq189 20940222
Tremblay, P. , & Small, S. L. (2011b). Motor response selection in overt sentence production: A functional MRI study. Frontiers in Psychology, 2 , 253. 10.3389/fpsyg.2011.00253 21994500
* Tuennerhoff, J. , & Noppeney, U. (2016). When sentences live up to your expectations. NeuroImage, 124 (Pt A ), 641–653. 10.1016/j.neuroimage.2015.09.004 26363344
* Tune, S. , Schlesewsky, M. , Nagels, A. , Small, S. L. , & Bornkessel‐Schlesewsky, I. (2016). Sentence understanding depends on contextual use of semantic and real world knowledge. NeuroImage, 136 , 10–25. 10.1016/j.neuroimage.2016.05.020 27177762
Turkeltaub, P. E. , Eickhoff, S. B. , Laird, A. R. , Fox, M. , Wiener, M. , & Fox, P. (2012). Minimizing within‐experiment and within‐group effects in activation likelihood estimation meta‐analyses. Human Brain Mapping, 33 (1 ), 1–13. 10.1002/hbm.21186 21305667
* Tyler, L. K. , Stamatakis, E. A. , Post, B. , Randall, B. , & Marslen‐Wilson, W. (2005). Temporal and frontal systems in speech comprehension: An fMRI study of past tense processing. Neuropsychologia, 43 (13 ), 1963–1974. 10.1016/j.neuropsychologia.2005.03.008 16168736
Unger, N. , Haeck, M. , Eickhoff, S. B. , Camilleri, J. A. , Dickscheid, T. , Mohlberg, H. , Bludau, S., Caspers, S., & Amunts, K. (2023). Cytoarchitectonic mapping of the human frontal operculum‐new correlates for a variety of brain functions. Frontiers in Human Neuroscience, 17 , 1087026. 10.3389/fnhum.2023.1087026 37448625
* Vaden, K. I., Jr. , Kuchinsky, S. E. , Cute, S. L. , Ahlstrom, J. B. , Dubno, J. R. , & Eckert, M. A. (2013). The cingulo‐opercular network provides word‐recognition benefit. The Journal of Neuroscience, 33 (48 ), 18979–18986. 10.1523/JNEUROSCI.1417-13.2013 24285902
* Vaden, K. I., Jr. , Teubner‐Rhodes, S. , Ahlstrom, J. B. , Dubno, J. R. , & Eckert, M. A. (2017). Cingulo‐opercular activity affects incidental memory encoding for speech in noise. NeuroImage, 157 , 381–387. 10.1016/j.neuroimage.2017.06.028 28624645
Vagharchakian, L. , Dehaene‐Lambertz, G. , Pallier, C. , & Dehaene, S. (2012). A temporal bottleneck in the language comprehension network. The Journal of Neuroscience, 32 (26 ), 9089–9102. 10.1523/jneurosci.5685-11.2012 22745508
Van der Haegen, L. , & Cai, Q. (2019). Lateralization of language. In G. I. de Zubicaray & N. O. Schiller (Eds.), The Oxford handbook of neurolinguistics. Oxford University Press. 10.1093/oxfordhb/9780190672027.013.34
* Vitello, S. , Warren, J. E. , Devlin, J. T. , & Rodd, J. M. (2014). Roles of frontal and temporal regions in reinterpreting semantically ambiguous sentences. Frontiers in Human Neuroscience, 8 , 14. 10.3389/fnhum.2014.00530 24478680
Vouloumanos, A. , Kiehl, K. A. , Werker, J. F. , & Liddle, P. F. (2001). Detection of sounds in the auditory stream: Event‐related fMRI evidence for differential activation to speech and nonspeech. Journal of Cognitive Neuroscience, 13 (7 ), 994–1005. 10.1162/089892901753165890 11595101
* Wartenburger, I. , Heekeren, H. R. , Burchert, F. , Heinemann, S. , De Bleser, R. , & Villringer, A. (2004). Neural correlates of syntactic transformations. Human Brain Mapping, 22 (1 ), 72–81. 10.1002/hbm.20021 15083528
Watkins, K. , & Paus, T. (2004). Modulation of motor excitability during speech perception: The role of Broca's area. Journal of Cognitive Neuroscience, 16 (6 ), 978–987. 10.1162/0898929041502616 15298785
* Wild, C. J. , Davis, M. H. , & Johnsrude, I. S. (2012). Human auditory cortex is sensitive to the perceived clarity of speech. NeuroImage, 60 (2 ), 1490–1502. 10.1016/j.neuroimage.2012.01.035 22248574
* Wild, C. J. , Yusuf, A. , Wilson, D. E. , Peelle, J. E. , Davis, M. H. , & Johnsrude, I. S. (2012). Effortful listening: The processing of degraded speech depends critically on attention. The Journal of Neuroscience, 32 (40 ), 14010–14021. 10.1523/JNEUROSCI.1528-12.2012 23035108
Wilson, S. M. , Saygin, A. P. , Sereno, M. I. , & Iacoboni, M. (2004). Listening to speech activates motor areas involved in speech production. Nature Neuroscience, 7 (7 ), 701–702. 10.1038/nn1263 15184903
* Wong, P. C. , Uppunda, A. K. , Parrish, T. B. , & Dhar, S. (2008). Cortical mechanisms of speech perception in noise. Journal of Speech, Language, and Hearing Research, 51 (4 ), 1026–1041. 10.1044/1092-4388(2008/075)
* Wright, P. , Randall, B. , Marslen‐Wilson, W. D. , & Tyler, L. K. (2011). Dissociating linguistic and task‐related activity in the left inferior frontal gyrus. Journal of Cognitive Neuroscience, 23 (2 ), 404–413. 10.1162/jocn.2010.21450 20201631
Wu, Z. M. , Chen, M. L. , Wu, X. H. , & Li, L. (2014). Interaction between auditory and motor systems in speech perception. Neuroscience Bulletin, 30 (3 ), 490–496. 10.1007/s12264-013-1428-6 24604634
* Yang, Y. H. , Marslen‐Wilson, W. D. , & Bozic, M. (2017). Syntactic complexity and frequency in the neurocognitive language system. Journal of Cognitive Neuroscience, 29 (9 ), 1605–1620. 10.1162/jocn_a_01137 28430044
* Yu, T. , Lang, S. , Birbaumer, N. , & Kotchoubey, B. (2011). Listening to factually incorrect sentences activates classical language areas and thalamus. Neuroreport, 22 (17 ), 865–869. 10.1097/WNR.0b013e32834b6fc6 21968321
* Yusoff, A. , Ng, S. , Teng, X. L. , & abd hamid, A. (2014). Investigating brain activation and neural efficacy during simple arithmetic addition task in quiet and in noise: An fMRI study. Jurnal Sains Kesihatan Malaysia, 12 , 23–33. 10.17576/jskm-1201-2014-04
Zatorre, R. J. , Evans, A. C. , & Meyer, E. (1994). Neural mechanisms underlying melodic perception and memory for pitch. The Journal of Neuroscience, 14 (4 ), 1908–1919. 10.1523/JNEUROSCI.14-04-01908.1994 8158246
* Zekveld, A. A. , Heslenfeld, D. J. , Johnsrude, I. S. , Versfeld, N. J. , & Kramer, S. E. (2014). The eye as a window to the listening brain: Neural correlates of pupil size as a measure of cognitive listening load. NeuroImage, 101 , 76–86. 10.1016/j.neuroimage.2014.06.069 24999040
* Zekveld, A. A. , Rudner, M. , Johnsrude, I. S. , Heslenfeld, D. J. , & Rönnberg, J. (2012). Behavioral and fMRI evidence that cognitive ability modulates the effect of semantic context on speech intelligibility. Brain and Language, 122 (2 ), 103–113. 10.1016/j.bandl.2012.05.006 22728131
* Zhuang, J. , Randall, B. , Stamatakis, E. A. , Marslen‐Wilson, W. D. , & Tyler, L. K. (2011). The interaction of lexical semantics and cohort competition in spoken word recognition: An fMRI study. Journal of Cognitive Neuroscience, 23 (12 ), 3778–3790. 10.1162/jocn_a_00046 21563885
