
==== Front
bioRxiv
BIORXIV
bioRxiv
2692-8205
Cold Spring Harbor Laboratory

39282266
10.1101/2024.09.03.610973
preprint
1
Article
DIFFERENTIAL CONTRIBUTIONS OF THE SPECTRO-TEMPORAL AND VOCAL CHARACTERISTICS OF AUDITORY PSEUDOWORDS TO MULTIPLE SOUND-SYMBOLIC MAPPINGS
Lacey Simon 123
Matthews Kaitlyn L. 45
Hoffmann A.M. 4
Sathian K. 123
Nygaard Lynne C. 4
1 Department of Neurology, Penn State Health Milton S. Hershey Medical Center & Penn State College of Medicine, Hershey, PA 17033, USA
2 Department of Neural & Behavioral Sciences, Penn State Health Milton S. Hershey Medical Center & Penn State College of Medicine, Hershey, PA 17033, USA
3 Department of Psychology, Penn State Health Milton S. Hershey Medical Center & Penn State College of Medicine, Hershey, PA 17033, USA
4 Department of Psychology Emory University, Atlanta, GA 30322, USA
5 Present address: Department of Psychological & Brain Sciences, Washington University in St. Louis, St. Louis MO 63130
Author contributions: SL, KS and LCN designed research; KLM collected data; KLM, AMH and SL analyzed data; and SL, KLM, AMH, KS and LCN wrote the paper.

Corresponding authors: Lynne C. Nygaard, Department of Psychology, Emory University, College of Arts and Sciences, Atlanta, GA 30322, USA, Fax: 717-531-0384, lnygaar@emory.edu, K. Sathian, Department of Neurology, Penn State Health Milton S. Hershey Medical Center, Penn State College of Medicine, Hershey, PA 17033-0859, USA, Fax: 717-531-0384, ksathian@pennstatehealth.psu.edu
05 9 2024
2024.09.03.610973https://creativecommons.org/licenses/by-nc-nd/4.0/ This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which allows reusers to copy and distribute the material in any medium or format in unadapted form only, for noncommercial purposes only, and only so long as attribution is given to the creator.
nihpp-2024.09.03.610973.pdf
Sound symbolism, the idea that the sound of a word alone can convey its meaning, is often studied using auditory pseudowords. For example, people reliably assign the auditory pseudowords “bouba” and “kiki” to rounded and pointed shapes, respectively. Previously we showed that representational dissimilarity matrices (RDMs) of the shape ratings of auditory pseudowords correlated significantly with RDMs of acoustic parameters reflecting spectro-temporal variations; the ratings also correlated significantly with voice quality features. Here, participants rated auditory pseudowords on scales representing categorical opposites across seven meaning domains, including shape. Examination of the relationships of the perceptual ratings to spectro-temporal and vocal parameters of the pseudowords essentially replicated our previous findings for shape while varying patterns emerged for the other domains. Thus, the spectro-temporal and vocal properties of spoken pseudowords contribute differentially to sound-symbolic mapping depending on the meaning domain.

Sound symbolism
Language
Representational similarity analysis
Multisensory
Acoustic
==== Body
pmcINTRODUCTION

Sound symbolism refers to the idea that the sound of a word can convey its meaning (Nuckolls, 1999; Svantesson, 2017); for example, ‘balloon’ and ‘spike’ not only sound rounded and pointed, respectively, but also refer to objects that are generally rounded and pointed (Sučević et al., 2015). The sound-meaning relationships of real words are generally taken to be arbitrary, following de Saussure (1916/2009), and more recently Hockett (1960), with occurrences of sound symbolism being considered too sporadic to imply a functional role in language (Lev-Ari & McKay, 2023). Contrary to this view, sound-symbolic mappings are widespread, occurring in many different languages and across different linguistic lineages (Blasi et al., 2016), suggesting that sound symbolism is sufficiently common to be functionally meaningful.

Sound symbolism is often investigated using auditory pseudowords like ‘maluma’ or ‘bouba’, which are regularly assigned to rounded shapes, and ‘takete’ or ‘kiki’, which are assigned to pointed shapes (Köhler, 1929, 1947; Ramachandran & Hubbard, 2001). Using pseudowords avoids the confound of participants’ prior knowledge of the meaning of real words: as noted above, ‘balloon’ and ‘spike’ both sound like, and mean, something rounded and pointed, respectively. The bulk of sound symbolism studies have focused on this rounded-pointed dimension of shape, but other sound-symbolic mappings include size (e.g., Sapir, 1929), weight (e.g., Walker & Parameswaran, 2019), and visual brightness (e.g., Hirata et al., 2011), and also extend to non-sensory domains such as valence (e.g., Uno et al., 2020).

The phonetic features of such sound-symbolic mappings have been extensively investigated (see the companion paper, Lacey et al., 2024), but whether and how acoustic properties, of either real words or pseudowords, contribute to their sound-symbolic mappings is largely unexplored (Knoeferle et al., 2017). Acoustic properties may be important because they relate to the listener’s end of the ‘speech chain’ (Denes & Pinson, 1993) with the spectro-temporal characteristics of speech mapping to phonetic features. For example, intracranial recordings show that, during listening to natural speech, the superior temporal gyrus (STG) shows phonetic selectivity arising from neuronal populations tuned to the specific spectro-temporal profiles of each phonetic feature (Mesgarani et al., 2014; see also Oganian et al., 2023; Hamilton et al., 2020). Accordingly, cortical responses to phonemic features of speech could be decoded from low-level acoustic parameters (Daube et al., 2019).

However, investigation of the role of such acoustic properties in sound symbolism has been intermittent, largely confined to pseudowords and for a limited range of acoustic features and sound-symbolic mappings. For example, in the study of Knoeferle et al. (2017), participants rated auditory pseudowords containing one of five vowels (back rounded / / and /u/, back unrounded / /, and front unrounded /ε/ and /i/): shape ratings were related to the frequencies of the vowel formants1 F2 and F3, while size ratings were related to F1 and F2. When participants vocalized a single vowel sound (back unrounded / /) in response to visual shapes (dodecagon vs. triangle), the frequency of F3 was higher for triangles, the more pointed of the two (Parise & Pavani, 2011). However, Parise & Pavani (2011) did not find any associations between the fundamental frequency or formant frequencies of vocalizations in response to large and small visual stimuli, perhaps as a result of the restriction to a single vowel. Nonetheless, vocalizations were louder for complex compared to simple shapes (dodecagon vs. triangle) and were also louder for brighter than darker stimuli (Parise & Pavani, 2011). This latter result is consistent with a later finding that speakers produced pseudowords referring to bright colors with greater amplitude than those for darker colors (Tzeng et al., 2018). Speakers also produced pseudowords for brighter colors with higher fundamental frequency, and shorter duration, than those for darker colors, and listeners could use these prosodic cues to reliably assign pseudowords to their target color (Tzeng et al., 2018).

In the first such systematic exploration of its kind, we previously investigated the contributions of a range of spectro-temporal and vocal parameters of a large set of pseudowords to sound-symbolic mapping for the rounded-pointed dimension of shape (Lacey et al., 2020). In a novel application of representational similarity analysis (RSA: Kriegeskorte et al., 2008), we showed that representational dissimilarity matrices (RDMs) for ratings of pseudowords as rounded or pointed were significantly correlated with RDMs for three spectro-temporal parameters: spectral tilt, the fast Fourier transform (FFT) and the speech envelope (Lacey et al., 2020). Spectral tilt, i.e. the slope of the power spectrum over the frequency range, was steeper for rounded pseudowords with power concentrated at lower frequencies, but flatter for pointed pseudowords as spectral power shifted to higher frequencies. The FFT, representing the power spectrum of the pseudoword across its duration, showed smoother distributions of spectral power over time, mainly at lower frequencies, for rounded pseudowords, compared to more irregular distributions spreading to higher frequencies for pointed pseudowords. The speech envelope, measuring changes in the amplitude profile over time, was more continuous and smoother for rounded words, but more uneven and discontinuous for pointed words. The vocal parameters could be broadly divided into those that reflect the relative periodicity of the speech signal, which largely results from the balance between voiced and unvoiced segments, (i.e., the fraction of unvoiced frames [FUF], the mean autocorrelation [MAC], the harmonics-to-noise ratio [HNR], and pulse number), and those that reflect vocal variability (i.e., jitter [variability in frequency], shimmer [variability in amplitude], and pitch standard deviation [PSD: variability in the fundamental frequency of the speech signal]). Conventional correlational analyses showed that periodicity (FUF, MAC, HNR, and pulse number) decreased, and vocal variability (shimmer jitter) increased as ratings of pseudowords transitioned from rounded to pointed (Lacey et al., 2020). However, PSD, while also a measure of vocal variability, was unrelated to shape ratings (Lacey et al., 2020). A full description of all the acoustic parameters used in Lacey et al. (2020) and in the current work is provided in the Methods section below.

Subsequently, in the shape domain, pulse phonation (in which the vocal folds open and close – i.e., pulse – less frequently than normal, resulting in a ‘creaky voice’ [Ishi et al., 2008]) has been associated with pointedness (Akita, 2021), replicating our finding that lower pulse numbers indicated pointedness (Lacey et al., 2020). In addition, pseudowords rated as pointed were associated with increased vocal variability, defined as rapid variation in amplitude (Villegas et al., 2023), consistent with our finding that shimmer increased as ratings changed from rounded to pointed (Lacey et al., 2020). Roundedness has also been associated with higher pitch (Villegas et al., 2023), and with falsetto phonation, which is characterized by a higher fundamental frequency than the speaker’s normal voice (Akita, 2021). However, both of these findings seem to contrast with the crossmodal correspondence between low/high pitch and obtuse/acute, i.e., less/more pointed, angles respectively (Parise & Spence, 2012). In a different domain, pulse phonation was associated with large size (Akita, 2021), and pseudowords rated as bigger were associated with increasing loudness while those rated as smaller were associated with increasing sharpness2 (Villegas et al., 2023).

Here, we provide further evidence for the importance of the acoustic properties in sound-symbolic mapping by replicating our previous work in the shape (rounded/pointed) domain and, more critically, extending it to six other sound-symbolic domains: size (small/big), texture (hard/soft), weight (light/heavy), brightness (bright/dark), arousal (calming/exciting), and valence (good/bad). Participants listened to a large set of spoken pseudowords that sampled from the phonetic space of American English, and rated how rounded, big, hard, etc., each item sounded on a seven-point Likert scale. We then compared these ratings to the spectro-temporal and vocal parameters of the pseudowords, using RSA for the spectro-temporal parameters and conventional correlation analysis for the vocal parameters, as in our previous study (Lacey et al., 2020). This allowed us to distinguish between two alternative hypotheses: whether all sound-symbolic mappings can be traced to a single underlying factor (a domain-general account) or whether the underlying factor(s) are unique to each domain (a domain-specific account). The domain-general account proposes that all sound-symbolic mappings reflect associations with a single, abstract dimension such as arousal, magnitude or valence (Aryani et al., 2020; Sidhu & Pexman, 2018; Spence, 2011). According to this hypothesis, the acoustic profile (i.e., the pattern of the relative importance of the acoustic parameters to the putative underlying abstract domain) should be similar across all domains. For example, if the underlying domain were arousal and HNR were the most important parameter, then HNR would be the most important parameter across all domains. Alternatively, in a domain-specific account, the acoustic profile could vary depending on the meaning domain, regardless of the relationship with a putative abstract domain, with parameters varying in importance for each domain as in, for example, the differing combinations of vowel formant values in associations with shape and size meanings (Knoeferle et al., 2017) referred to above. (See also Lacey et al. [2024] for an examination of domain-general vs -specific accounts of sound symbolism in relation to ratings and phonetic features).

METHODS

Participants

Participants were recruited, and compensated for their time, online via the Prolific participant pool (https://prolific.ac: and see Peer et al., 2017). For full details of the inclusion and exclusion criteria, please refer to Lacey et al. (2024). The final sample comprised 389 participants (171 male, 208 female, 6 non-binary, 2 agender, 1 gender fluid, and 1 who declined to state gender; mean age 27 years, 1 month [SD 8 months]). The rating studies were hosted on the Gorilla platform (https://gorilla.sc: Anwyl-Irvine et al., 2020) where participants also gave informed consent. All procedures were approved by the Emory University Institutional Review Board.

Auditory Pseudoword Stimuli

We used the set of 537 two-syllable CVCV (i.e., consonant/vowel/consonant/vowel) pseudowords created by McCormick et al. (2015), comprising only phonemes and combinations of phonemes that occur in American English. Importantly, although the inventory of phonemes sampled the acoustic-phonetic space in terms of manner and place of articulation as well as voicing characteristics (see Table 1), these pseudowords were constructed using phonemes for which a sound-symbolic mapping to the rounded-pointed dimension of shape had previously been established, i.e., the pseudowords were designed to investigate the shape domain (McCormick et al., 2015). The pseudowords were recorded by a female native speaker of American English, in random order and with neutral intonation, and digitized at a 44.1 kHz sampling rate. Each pseudoword was then down-sampled at 22.05 kHz, which is standard for speech, and amplitude-normalized using PRAAT speech analysis software (Boersma & Weenink, 2012); the mean duration of the pseudowords was 457 ± 62 ms. For full details of the recording procedures and the phonetic content of the pseudowords, please refer to McCormick et al. (2015) and Lacey et al. (2020, 2024).

General procedures

Participants recruited via Prolific followed a link that took them to the experiment on the Gorilla platform. At the landing page, they could read the consent form and a short description of the task, before clicking on ‘yes’ to take part or ‘no’ to exit. Having consented, and before taking part, participants completed a headphone check designed to ensure compliance with headphone use for web-based experiments (Woods et al., 2017). Participants who failed the headphone check could still take part, in order to avoid discriminating against those who did not have access to headphones, but their data were excluded from analysis. The randomizer function in Gorilla then assigned participants to one of the two rating scales for the relevant domain. Once the rating task had been completed, participants were directed to a series of short questionnaires that asked about any task strategies that had been used, demographic information, and language experience and ability (Lacey et al., 2024). Participants could then read a debriefing statement and exit the experiment.

Perceptual rating tasks

Participants were randomly assigned to one of two 7-point Likert-type scales (e.g., not rounded to rounded, or not pointed to pointed) representing categorical opposites across seven different meaning domains: shape (rounded-pointed, N =30/30 respectively), size (small-big, N = 32/31), texture (hard-soft, N = 27/27), weight (light-heavy, N = 29/26), brightness (bright-dark, N = 24/26), arousal (calming-exciting, N = 30/28), and valence (good-bad, N = 26/23). Participants rated all 537 pseudowords on the scale to which they were assigned. For further details of the ratings task, see Lacey et al. (2024).

Acoustic Parameters

Spectro-temporal parameters

As in Lacey et al. (2020), we chose to measure the speech envelope, spectral tilt, and the FFT. For the detailed calculation of each parameter, please see the Supplementary Material.

Speech envelope

The speech envelope measures changes in the amplitude profile over time, largely corresponding to changes in phonemic properties and syllabic transitions (Aiken & Picton, 2008). To the extent that these transitions are abrupt, reflecting stops, affricates, or fricatives3 (see Table 1), the speech envelope is discontinuous and uneven (Figure 1a, left panel); but where they are more gradual, reflecting sonorants, the envelope appears more continuous and smoother (Figure 1a, right panel).

Spectral tilt

Spectral tilt reflects differences in power across frequencies and is an estimate of the overall slope of the power spectrum, with sampling across the complete utterance. When high frequencies have less power than low frequencies, the power spectrum slopes steeply downward from low to high frequencies (Figure 1b, right panel), flattening out when power is more concentrated in the high frequencies (Figure 1b, left panel).

Fast Fourier Transform (FFT):

The FFT derives the frequency components of the speech signal and the variation in their energy over time, thus reflecting the power spectrum of the frequency composition across the duration of the spoken pseudowords. The FFT is illustrated by the spectrogram which shows how power is distributed across frequencies over time; for example, obstruents tend to be reflected in abrupt changes in power as a function of frequency (Figure 1c, left panel) while for sonorants, power varies more gradually with frequency (Figure 1c, right panel).

Vocal Parameters

We chose the same voice parameters as in our previous study (Lacey et al., 2020): the fraction of unvoiced frames (FUF), the mean autocorrelation (MAC), the mean harmonics-to-noise ratio (HNR), and pulse number (these parameters reflect the relative amount of periodicity of the speech signal), together with jitter and shimmer (which reflect the variability of voicing or voice quality during the production of each pseudoword), and the standard deviation of the voice pitch. In the present study, we also included the mean pitch and the duration of the pseudoword. Note that, although extreme values for some of these parameters can indicate vocal pathology (e.g., Brockmann et al., 2011; Ferrand, 2002; Teixeira & Fernandes, 2014), they also vary naturally in a healthy voice as employed here (see Brockmann et al., 2011).

Mean harmonics-to-noise ratio (HNR)

This is the ratio between the periodic, or harmonic, portions of the speech signal and the aperiodic, or noise, portions. The mean HNR is thus an estimate of the overall periodicity of the sound expressed in dB (Teixeira & Fernandes, 2014). The noise element arises from turbulent airflow at the glottis when the vocal cords do not close properly (Ferrand, 2002). As noise increases, and therefore, mean HNR decreases, the voice becomes increasingly hoarse or quavery and the speech pattern becomes progressively more uneven (Ferrand, 2002).

Mean autocorrelation (MAC)

This is a measure of the periodicity of a signal (Boersma & Weenink, 2012). Periodicity should be high for a long vowel like ‘ooo’ or consonant like ‘mmm’, and each successive segment should sound very similar to the one before, i.e. they should be highly correlated. Higher autocorrelation values indicate a smoother voice pattern and/or more voiced segments, while lower values indicate an uneven pattern and/or fewer voiced or periodic segments.

Pulse number

This is the number of glottal pulses, i.e. opening and closing of the vocal folds, during production of vowels or voiced consonants measured across the whole utterance (Boersma & Weenink, 2012). An extreme form of phonation, known as pulse register phonation, will help to understand how the pulse number manifests in the voice. In pulse register phonation, rapid glottal pulses are followed by a long, closed phase (Hollien et al., 1977; Whitehead et al., 1984). This results in an audibly uneven speech pattern described as a ‘creaky voice’ (Ishi et al., 2008) or – onomatopoeically – as a ‘glottal rattle’ (Hornibrook et al., 2018). A lower pulse number indicates a more uneven voice pattern, and/or fewer voiced segments, while higher pulse numbers indicate a smoother voice pattern, and/or more voiced segments.

Fraction of unvoiced frames

The FUF represents the number of unvoiced elements, expressed as the percentage of measurement windows that do not engage the vocal folds (Boersma & Weenink, 2012). The FUF depends on the phonemic content, increasing for those pseudowords that include unvoiced elements, like obstruents, and decreasing for those containing voiced (i.e., periodic) elements, typically long vowels.

Shimmer

Shimmer indexes peak-to-peak variation in the amplitude of the glottal waveform (Brockmann et al., 2011). Shimmer reflects vocal instability: low shimmer results in a smooth speech pattern whereas high shimmer results in an uneven speech pattern and manifests as a hoarse voice.

Jitter

Jitter is defined as the frequency variation between consecutive periods and is a measure of voice quality in that it measures variation in the vibration of the vocal cords (Teixeira & Fernandes, 2014). Vocally, high values of jitter manifest as a ‘breaking’ or rough voice. Jitter is typically measured for long vowel sounds, where little frequency variation would be expected. In the production of the pseudowords, increasing jitter reflects increased vocal instability or variation, and perceived vocal roughness.

Mean pitch & pitch standard deviation (PSD)

Mean pitch is the mean fundamental frequency, F0, of the speech object while PSD indicates the variation in the fundamental frequency present in the speech signal (Boersma & Weenink, 2012). PSD is a measure of vocal inflection, with low PSD resulting in a flat, monotone voice and high PSD in a ‘lively’ voice (Kliper et al., 2016).

Duration

This is the duration of the speech signal in milliseconds (ms) and is likely influenced by differences between long and short vowels (Kluender et al., 1988; Hillenbrand et al., 1995), phonetic context (Kluender et al., 1988), and speaking rate (Miller & Volaitis, 1989).

Domain-specific predictions

Our predictions for the acoustic characteristics associated with each domain flow from a phonetic analysis of the manner and place of articulation, as well as voicing, for the consonants and height, backness, and rounding for the vowels used in the current pseudoword set in relation to their ratings for the same set of domains (see Lacey et al., 2024). Since the spectro-temporal parameters take multiple samples across the waveform of each pseudoword, these are the most easily connected to the phonetic features of the pseudowords. This is less easily done for the vocal parameters, for which we derived only one value per pseudoword: while the FUF, mean HNR, MAC, and pulse number all reflect the relative voicing/periodicity of each item – and therefore connect to the phonetic features in a broad sense – it is less easy to pinpoint the phonetic features contributing to vocal variability.

Consonants characterized as stops, af/fricatives, and sonorants for example, each have a different manner of articulation and thus different acoustic characteristics (see Table 1 for the consonants and vowels used in creating the pseudowords and their phonetic and articulatory features). Stops involve a constriction in the vocal tract that temporarily blocks the flow of air, followed by a release as airflow resumes (Ladefoged & Johnson, 2011). The blockage results in a short period of low energy as the constriction is formed, followed by a ‘burst’ of energy (see Chodroff & Wilson, 2014) as the blockage is released; these can be seen in the spectrogram and speech envelope as abrupt variations in power and amplitude, respectively (see Supplementary Figures 1 and 2 for unvoiced and voiced stops in /tike/ and /gobo/, respectively). Fricatives (/f/, /v/, /s/, and /z/ in Table 1) involve a partial obstruction of the vocal tract that results in a turbulent airflow, and also involve higher frequencies than almost any other phoneme (Ladefoged & Johnson, 2011; see also Jongman et al., 2000). Pseudowords including fricatives should therefore have a relatively flatter spectral tilt since there will be more power at the higher frequencies (see Supplementary Figure 7 for /vu o/). Affricative consonants (/ / and / / in Table 1) combine a stop and a fricative (Ladefoged & Johnson, 2011) and their acoustic consequences therefore manifest as a combination of abrupt variations in power and amplitude, as demonstrated in the spectrogram and speech envelope, respectively, together with concentrations of power at the high frequencies, reflected in both the spectrogram and spectral tilt (see Supplementary Figures 4 and 7 for / i e/ and /vu o/ respectively). Sonorants do not involve an obstruction of the airflow: although there is a constriction in the vocal tract, the airflow continues through the mouth for /l/, and through the nose for /m/ and /n/ (Ladefoged & Johnson, 2011). Since the airflow is also unobstructed for vowels, the transitions between sonorants and vowels in the CVCV pseudowords should result in more gradual variations in power in the spectrogram and a smoother speech envelope (compare, for instance, /m mo/ to /tike/ in Supplementary Figure 1).

Different places of articulation also result in varying acoustic consequences because they alter the length of the vocal tract depending on where the constriction that produces the sound is located (Table 1). Bilabial consonants like /b/, for example, involve resonances of the whole vocal tract because the constriction is at the lips, and therefore are associated with high energy at the lower frequencies (Reetz & Jongman, 2020). By contrast, alveolar consonants like /d/ produce most energy at high frequencies because the constriction is produced by the tongue against the alveolar ridge, just behind the teeth, and therefore the vocal tract in front of the constriction is short (Reetz & Jongman, 2020). Velar consonants like /g/ produce energy in the middle range of frequencies, the constriction being produced by the tongue against the velum, or soft palate, at the back of the mouth, and an intermediate anterior vocal tract length (Reetz & Jongman, 2020). These considerations enable us to make predictions about spectral tilt and the FFT in particular: Pseudowords consisting of bilabial consonants will tend to show a steeper spectral tilt than those consisting of alveolar consonants because of the concentration of energy at low and high frequencies respectively. For the FFT, there will be differences in the distribution of energy at each frequency band over the duration of the pseudoword; these will be apparent from the spectrogram.

Shape

In the shape domain, we expected to replicate, in an independent sample of participants, our previous findings for spectro-temporal and vocal parameters of the auditory speech signal (Lacey et al., 2020), as summarized here. In this earlier study, we found spectral tilt to be steeper for rounded pseudowords where power is concentrated in the low-frequency bands (Lacey et al., 2020), reflecting the sonorants and back rounded vowels that are associated with roundedness (McCormick et al., 2015; Lacey et al., 2024; see also Table 1). However, spectral tilt flattened out for pointed pseudowords as power migrates to the higher frequencies associated with the stops and af/fricatives (obstruents) and/or front unrounded vowels that these pseudowords contain (Lacey et al., 2020). This also reflects differences in the place of articulation with power being concentrated at lower frequencies for the bilabial/labiodental consonants associated with roundedness, but dispersing to the higher frequencies for the alveolar and post-alveolar/velar consonants associated with pointedness (Lacey et al., 2024). We also found that the FFT reflected power variations that were gradual for rounded pseudowords and abrupt for pointed pseudowords (Lacey et al., 2020), again reflecting the presence of sonorants and obstruents, and rounded/unrounded vowels respectively (McCormick et al., 2015; Lacey et al., 2024). The speech envelope was smoother and more continuous for rounded, compared to pointed, pseudowords (Lacey et al., 2020).

For the previously studied vocal parameters, mean HNR decreased as the speech pattern became progressively less smooth and more uneven, reflecting the change from rounded to pointed (Lacey et al., 2020). The MAC decreased in the same way: higher autocorrelation values indicate a smoother voice pattern and/or more voiced segments, associated with roundedness ratings, while lower values indicate an uneven pattern and/or fewer voiced or periodic segments, associated with ratings of pointedness (Lacey et al., 2020). Similarly, higher pulse numbers indicate a smoother voice pattern, and/or more voiced segments, associated with rounded pseudowords, while lower pulse numbers indicate a more uneven voice pattern, and/or fewer voiced segments, associated with pointed pseudowords (Lacey et al., 2020). In keeping with this, the pulse number decreased from the rounded to the pointed pseudowords (Lacey et al., 2020); note that this effect has been independently replicated (Akita, 2021). The FUF increased as ratings of pseudowords transitioned from rounded to pointed (Lacey et al., 2020), because auditory roundedness and pointedness are more associated with voiced and unvoiced elements, respectively (McCormick et al., 2015; Lacey et al., 2024).

Variation in amplitude and frequency as measured by shimmer and jitter, respectively, also increased, reflecting increasing vocal instability and unevenness in the speech pattern, as pseudoword ratings progressed from rounded to pointed (Lacey et al., 2020). Similarly, lesser and greater pitch variability as measured by PSD indicated roundedness and pointedness, respectively (Lacey et al., 2020).

The present study also introduced two vocal parameters that were not included in the prior study of Lacey et al. (2020): mean pitch and pseudoword duration. We expect that mean pitch will increase as ratings move from rounded to pointed (Parise & Pavani, 2011; Knoeferle et al., 2017). Finally, an intuitive prediction is that roundedness and pointedness should be associated with longer and shorter pseudoword duration, respectively.

Size

For the size domain, sound-symbolic smallness ratings were associated with unvoiced, bilabial/labiodental stops but not af/fricatives or vowels, while large size was associated with voiced, post-alveolar/velar stops and af/fricatives, and back rounded vowels (Lacey et al., 2024). We expected that spectral tilt and the spectrogram would be broadly similar for small and big pseudowords since both might involve low and high frequencies. For small pseudowords, although unvoiced consonants are generally produced with higher frequency than voiced consonants, the bilabial/labiodental place of articulation results in lower frequencies. For big pseudowords, fricatives exhibit high frequencies, whereas the post-alveolar/velar place of articulation and back rounded vowels should involve lower frequencies. These concentrations of energy at both low and high frequencies should be apparent from the frequency distributions shown in the spectral tilt and the spectrogram for both small and big pseudowords. Since small and big ratings were both associated with stops, we would expect the speech envelope for both to show a discontinuous profile but, given the difference in voicing, the envelope for voiced, big pseudowords might exhibit greater amplitude compared to the unvoiced, small pseudowords.

For the vocal parameters, we can make predictions based on the well-established crossmodal correspondence in which high and low pitch are associated with small and big size, respectively (reviewed by Spence, 2011). Pseudowords with high/low mean pitch should therefore be rated as small/big respectively. Although pulse phonation, or ‘creaky voice’ has recently been associated with large size (Akita, 2021), the speaker in that study deliberately employed pulse phonation whereas ours spoke in their typical voice. Nonetheless, it is likely that measures of periodicity will increase (HNR, MAC, pulse number) and decrease (FUF) following the change from unvoiced to voiced consonants for small/big pseudowords, respectively (Lacey et al., 2024). However, it is unclear how measures of vocal variability (jitter, shimmer, PSD) relate to size sound symbolism. Pseudowords with shorter/longer durations should be rated as smaller/bigger respectively (Knoeferle et al., 2017).

Texture

Sound-symbolic hardness ratings were associated with obstruents at the alveolar and post-alveolar/velar places of articulation, with little influence of vowels, while softness was associated with sonorants and the bilabial/labiodental place of articulation, and with back rounded vowels (Lacey et al., 2024). We therefore expected that there would be more energy at higher frequencies for hard, compared to soft, pseudowords, which should be reflected in a flatter spectral tilt for the hard pseudowords. The spectrograms should similarly show that hard pseudowords have more energy at the higher frequencies and more abrupt variations in energy, related to the presence of obstruents, compared to the more gradual variations involved in the sonorants that are associated with softness. These differences should also be reflected in the speech envelope, with clear discontinuities for the hard pseudowords but a smoother, more continuous envelope for the soft pseudowords.

For the vocal parameters, we would expect that softness would be associated with pseudowords that sound smoother, i.e. that have greater periodicity and less variability, while those that indicate hardness would involve less periodicity and more variability, i.e. relatively greater vocal roughness. While we are cognizant of describing one aspect of texture in terms of another (hard/soft and rough/smooth are independent dimensions of tactile texture [Hollins et al., 2000]), this is supported by the fact that Lacey et al. (2024) found that hardness was associated with af/fricatives (which are inherently noisy and aperiodic [Reetz & Jongman, 2020]), while softness was associated with sonorants. Thus, we would expect that the mean HNR, MAC, and pulse number would all increase, and that FUF would decrease, as increasing periodicity accompanies the transition of ratings from hard to soft. This transition would also be marked by a reduction in vocal variability and so jitter, shimmer, and PSD would reduce from hard to soft pseudowords. We did not make specific predictions for mean pitch or duration.

Weight

Phonetic analysis of the weight domain showed that lightness ratings were associated with sonorants and the bilabial/labiodental place of articulation, together with unrounded vowels; heaviness ratings were associated with voiced obstruents and the post-alveolar/velar place of articulation, together with rounded vowels (Lacey et al., 2024). We do not expect a major association of ratings with spectral tilt because both light- and heavy-rated pseudowords involve energy across the frequency range: sonorants and bilabial/labiodental articulation involve low frequencies while unrounded vowels are high frequency; by contrast, obstruents with post-alveolar/velar articulation involve somewhat higher frequencies while rounded vowels are generally of lower frequency. Pseudowords rated as light or heavy should, however, be distinguishable by their spectrograms which should show more gradual variations in energy for the light pseudowords and their associated sonorants, compared to more abrupt variations for the heavy pseudowords and their associated obstruents. The difference between sonorants and obstruents for light and heavy pseudowords, respectively, should also be apparent in more discontinuous speech envelopes for the latter compared to the former.

Since there is a real-world relationship in which size and weight are generally positively correlated, some predictions about vocal parameters can be derived by comparison to the size domain. High and low pitch are associated with small and big size, respectively (reviewed by Spence, 2011) and so pseudowords rated as light/heavy should also exhibit high/low mean pitch respectively. We expect that measures of periodicity – pulse number, mean HNR, MAC, and FUF – will show that periodicity decreases as ratings change from light to heavy, tracking the change from sonorants and other voiced consonants, in pseudowords rated as light, to unvoiced consonants in those rated as heavy (Lacey et al., 2024). But it is an open question whether vocal variability – as measured by jitter, shimmer, and PSD – is related to weight. By analogy to size, however, we would predict that shorter/longer duration would reflect light/heavy pseudowords, respectively.

Brightness

Our phonetic analysis of this domain showed that consonant voicing and vowel rounding were the main predictors of sound-symbolic brightness ratings (Lacey et al., 2024). Thus, we expected that pseudowords rated as dark would have energy concentrated in the lower-frequency bands, given their association with voiced consonants and back rounded vowels (Lacey et al., 2024), and thus would show a steeper spectral tilt than those rated as bright, which are associated with unvoiced consonants and front unrounded vowels (Lacey et al., 2024) involving energy at higher frequencies. The spectrogram should therefore show more power at high frequencies for the bright, compared to the dark, pseudowords. The crossmodal correspondence in which louder/quieter sounds are associated with bright/dark stimuli, respectively (reviewed in Spence, 2011; see also Tzeng et al., 2018), suggests that the speech envelope would reveal greater amplitude for the bright, compared to the dark, pseudowords. Note that place of articulation was not a particular predictor of bright/dark ratings and thus we offer no related acoustic hypotheses.

For the vocal parameters, there is also a crossmodal correspondence between high/low pitch and bright/dark stimuli, respectively (reviewed in Spence, 2011; see also Tzeng et al., 2018). We could therefore expect pseudowords reflecting the brightness/darkness dimension to have high/low mean pitch, following Tzeng et al. (2018) who found that, on average, participants produced pseudowords with higher/lower pitch in response to brighter/darker colors, respectively. Other predictions follow from the association between brightness/darkness and unvoiced/voiced consonants respectively (Newman, 1933; Hirata et al., 2011; Lacey et al., 2024): the FUF should decrease from bright to dark whereas the pulse number, MAC, and mean HNR should increase (see Lacey et al., 2020, for an example in the shape domain of how these parameters co-vary). Finally, brightness/darkness should correspond to shorter/longer duration (Tzeng et al., 2018). Note that the pitch and amplitude crossmodal associations with brightness were initially established with auditory stimuli that, unlike speech utterances, did not intrinsically vary in frequency or intensity (i.e., steady state pure tones, see Wicker, 1968). Thus, we have no specific prediction as to whether parameters that capture variability in pitch (jitter and PSD) and amplitude (shimmer) will be important.

Arousal

Sonorants and back rounded vowels are associated with pseudowords rated as calming (Sidhu et al., 2022; Lacey et al., 2024) while unvoiced obstruents and front unrounded vowels are associated with pseudowords rated as exciting (Sidhu et al., 2022; Lacey et al., 2024). The relationship between arousal and place of articulation is asymmetric: while calming ratings are strongly associated with bilabial/labiodental articulation, exciting ratings are not associated with a specific place of articulation (Lacey et al., 2024). Accordingly, we expect a relatively steep spectral tilt for calming pseudowords with energy concentrated at the lower frequencies, reflecting sonorants and rounded vowels, but a flatter spectral tilt for exciting pseudowords reflecting the higher-frequency energy associated with unvoiced obstruents. In addition, calming/exciting pseudowords should show smoother/more abrupt changes in energy in the spectrogram, together with continuous/discontinuous speech envelopes, respectively.

For the vocal parameters, since calming/exciting were associated with voiced and unvoiced consonants respectively, the FUF should increase as ratings change from calming to exciting. We would expect that pulse number, MAC, and mean HNR should decrease as ratings transition from calming to exciting, reflecting decreasing periodicity. Although voice stress analysis suggests that jitter and shimmer decrease with arousal (reviewed by Van Puyvelde et al., 2018), this was observed largely in relation to real-life emergency communications, a very different context to the present study. However, this review also notes that general, i.e., non-emergency, arousal produces increases in pitch range, i.e., variability, so we might expect jitter and PSD to increase with arousal. If increasing arousal produced a general increase in vocal variability, then we would also expect shimmer to increase. Intuitively, mean pitch will increase from calming to exciting, and shorter pseudowords will be rated as more exciting than longer pseudowords.

Valence

Our phonetic analysis of the valence domain indicated that the main predictors of pseudowords rated as good were sonorants and front unrounded vowels, while pseudowords rated as bad were associated with obstruents and back rounded vowels (Lacey et al., 2024). Place of articulation was not a strong predictor of good/bad ratings (Lacey et al., 2024). Pseudowords rated as sounding good or bad both involve energy at high and low frequencies: unrounded vowels and sonorants respectively for good pseudowords, and obstruents and rounded vowels respectively for bad pseudowords. Therefore, as for the weight domain, we did not expect major differences in spectral tilt. Good and bad pseudowords should, however, be distinguishable by their spectrograms which should show more gradual and diffuse changes in energy for the good pseudowords and their constituent sonorants, compared to more abrupt changes for the bad pseudowords and their constituent obstruents. The speech envelope should also distinguish good and bad pseudowords in being smoother and more continuous for the former and their associated sonorants, but more discontinuous for the latter and their associated obstruents.

For the vocal parameters, it seems reasonable to assume increasing vocal variability as pseudowords change from good/sonorants to bad/obstruents. We therefore expected that jitter, shimmer, and PSD, would increase as ratings transitioned from good to bad and, accordingly, that HNR, MAC, and pulse number would decrease, and FUF would increase, in tandem with this change (since decreasing periodicity implies a noisier signal). However, the effect of voicing was small (Lacey et al., 2024) in general, so we might expect measures of periodicity to correspondingly be weakly associated with valence ratings. Intuitively, badness would be associated with low, rather than high, pitch (Belyk & Brown, 2014), but there is no obvious prediction for pseudoword duration in this domain.

Summary of domain-specific predictions

We can summarize the predictions as follows. For the spectro-temporal parameters, we expected that pseudowords rated as rounded, soft, dark, and calming would be associated with a steeper spectral tilt than those rated as pointed, hard, bright, and exciting. Pseudowords rated as rounded, light, soft, calming, and good should exhibit a more continuous speech envelope and more gradual variations in power across frequency in the spectrogram, whereas those rated as pointed, heavy, hard, exciting, and bad should result in a discontinuous speech envelope and more abrupt variations in power. For the vocal parameters, we expected that increased periodicity (as measured by the HNR, MAC, pulse number, and FUF) would generally reflect ratings of pseudowords as rounded, small, soft, light, bright, soft, calming, and good. Conversely, decreased periodicity – i.e., more noise in the signal – would reflect pseudowords rated as pointed, big, hard, heavy, dark, exciting, and bad. We expected higher vocal variability (as measured by jitter, shimmer, and PSD) to reflect pseudowords rated as pointed, hard, exciting, and bad, while lower variability would be associated with pseudowords rated as rounded, soft, calming, and good.

Analysis of Acoustic Parameters

Spectro-temporal parameter analysis

Representational similarity analysis (RSA) was originally developed as a method for analyzing functional magnetic resonance imaging (fMRI) data (Kriegeskorte et al., 2008). In a prior study (Lacey et al., 2020), we used RSA to compare perceptual ratings of pseudowords on the rounded/pointed dimension of the sound-symbolic shape domain to the spectro-temporal parameters of these pseudowords. Here we used RSA in a similar manner for the seven domains of the present study. For each domain, a 537 × 537 representational dissimilarity matrix (RDM) was constructed based on the perceptual ratings of all participants for the entire set of pseudowords. Each cell of the RDM gives the dissimilarity for a particular item pair in terms of 1-r, where r is the correlation between the ratings for those two items; the RDM indicates the dissimilarity of each item to every other item. Similarly, we computed RDMs based on the spectro-temporal properties of the pseudowords, vectorizing the multiple measurements obtained for each item, computing pairwise correlations, and using 1-r to index pairwise dissimilarity. The resulting RDMs for each spectro-temporal parameter were compared to the reference perceptual rating RDMs by way of second-order correlations. To the extent that the RDM for a particular spectro-temporal parameter is significantly correlated with the reference RDM for the ratings for a given domain, that parameter can be said to contribute to the sound-symbolic mapping for that domain.

Speech envelope, spectral tilt, FFT and the related RDMs were calculated in MATLAB as described previously (Lacey et al., 2020). We normalized the duration of all pseudowords to the mean of 457 ms using the resampling function in MATLAB 2021a. At the down-sampled rate of 22.05 kHz (see Auditory Pseudoword Stimuli above), this resulted in a common vector length of 10077 data points per pseudoword (22050 × .457 = 10077), at the expense of a small amount of noise proportional in magnitude to the SD of the duration (SD/mean = 62/457 ms, i.e. 13.5%). Thus, we obtained equal numbers of data points per pseudoword for the RDMs (see below). Note that these parameters are not necessarily independent of each other (for example, both spectral tilt and the FFT index the frequency composition of the pseudowords), and that each of these parameters entails multiple datapoints for each pseudoword (see Supplementary Material).

As in our prior study (Lacey et al., 2020), we implemented RSA in MATLAB (version 2021a, The MathWorks, Natick MA). A schematic of the analysis pipeline is shown in Fig. 2 and we describe each step in more detail below.

We first created reference RDMs for pseudowords based on their perceptual ratings for each domain. For each domain, one of the two scales was recoded to the opposite scale, e.g. for the shape domain, the ‘rounded’ scale was recoded to the ‘pointed’ scale, such that 1 (‘not rounded’) on the rounded scale became 7 (‘very pointed’) on the pointed scale, and vice versa, producing a single scale in which the values 1–7 ran from rounded to pointed (Figure 2, Step 1). Since the two scales for each domain were intended to capture the categorical opposites of each dimension, e.g. brightness and darkness, they should be negatively correlated and this was indeed established for each domain (Lacey et al., 2024), thus justifying recoding to a single scale.

Thus, in the RDMs for each domain, the pseudowords were ordered left to right, based on the mean rating for each item on the single scale, from most rounded to most pointed for shape, smallest to biggest for size, hardest to softest for texture, lightest to heaviest for weight, brightest to darkest for brightness, most calming to most exciting for arousal, and best to worst for valence. Once the pseudowords had been ordered in this way, we created the reference RDMs for each domain (Fig. 2, Step 2), by calculating the first-order correlation (Pearson’s r) across participants between the perceptual ratings for each pair of pseudowords using the original, un-recoded data (since the RDMs reflect dissimilarity between items regardless of the rating scale that any individual participant used for any given domain). Pairwise dissimilarity is given by 1-r, the value entered in each cell of the RDM.

The next step was to create RDMs reflecting the pairwise dissimilarity for the spectro-temporal parameters of the pseudowords (Figure 2, Step 3). Although, as described above, there was a common vector length for each pseudoword, measurements were taken from these vectors in different ways for each spectro-temporal parameter, e.g., different window lengths and overlaps (see Supplementary Material for details). Thus, the number of measurements underlying the pairwise correlations for each of these parameters differed. For the speech envelope and the FFT, there were 10077 and 5489 data points respectively and the pairwise dissimilarities were calculated using Pearson correlations. For spectral tilt, however, there were only 8 data points, therefore we calculated pairwise dissimilarities using non-parametric (Spearman) correlations. Pairwise dissimilarity is again given by 1-r, the value entered in each cell of the RDM.

Finally, the parameter RDMs were compared to the reference ratings RDMs for each domain via a second-order, Spearman correlation (Fig. 2, Step 4)4. Significant second-order correlations would indicate which spectro-temporal parameters contributed to perception of the pseudowords as rounded or pointed, big or small, etc.

Voice parameter analysis

As reported previously (Lacey et al., 2020), the voice parameters were measured in PRAAT (Boersma & Weenink, 2012), using the original pseudoword sound files, without the resampling outlined above. As with the spectro-temporal parameters, the vocal parameters are not necessarily independent of each other (for example, both FUF and pulse number reflect how often the vocal folds open and close and are measures of the percentage of a segment which is voiced). Unlike the spectro-temporal parameters, the vocal parameters are expressed as a single value per pseudoword. While most of the voice parameter measurements are straightforward, there are several ways to calculate jitter and shimmer. We calculated local jitter and shimmer, i.e. the mean absolute difference in frequency and amplitude, respectively, between consecutive periods of the speech waveform divided by the mean difference over all periods of the speech waveform and expressed as a percentage (Boersma & Weenink, 2012). The voice parameters, comprising a single measurement per pseudoword, were compared directly to (the recoded) perceptual ratings in conventional correlational analyses.

Power analysis

Sample sizes were largely predetermined by the number of pseudowords, since the analyses were item-based, using ratings averaged across participants. Given the remote testing used in the present study, we aimed for twice the number of participants for each domain compared to our previous study (Lacey et al., 2020), in which 15 participants provided ratings on the roundedness scale, and 16 on the pointedness scale. For the vocal parameters, which have only one value per pseudoword, the power for a sample size of 537 to detect a Pearson correlation of .3 was 1.0. For RSA, the sample size for the second-order Spearman correlation between RDMs is half the off-diagonal data; for a 537 × 537 matrix, this is 143,916 and the power to detect a coefficient of .3 was 1.0. (Note that a power analysis is not required for the first-order pairwise correlations between pseudowords that constitute the 1-r value in each cell of the RDMs because we are not concerned with the strength or significance of the correlation, only with its variability within the pseudoword set.)

RESULTS AND DISCUSSION BY DOMAIN

For the three spectro-temporal parameters, the Bonferroni-corrected alpha for three tests (of the second-order correlations between RDMs of ratings and the spectro-temporal parameters) is .0167; for the nine voice parameters, the corrected alpha for nine tests (of conventional correlations with ratings) is .0056. The results are reported in approximately descending order, from the domain with the strongest parameter-ratings relationships (shape) to that with the weakest relationships (size). The spectro-temporal parameter results are illustrated using representative examples of the most highly rated pseudowords for each domain (Supplementary Figures 1–7). As noted above, the spectro-temporal parameters comprise multiple samples across each pseudoword and are most easily connected to phonetic features. For the vocal parameters, the FUF, mean HNR, MAC, and pulse number all reflect the relative voicing/periodicity involved in each item, and therefore connect to the phonetic features in a broad sense. But it is less easy to pinpoint the relationship between particular phonetic features and vocal variability.

Shape

In line with predictions and consistent with our previous study (Lacey et al., 2020: and see Table 2), the RDMs for all three spectro-temporal parameters were significantly positively correlated with the RDM for ratings of the pseudowords from rounded to pointed (Figure 3). The strongest correlation (r = .41, p < .001) was for spectral tilt, which was steeper for pseudowords rated as more rounded where power is concentrated in low-frequency bands, but flatter for pseudowords rated as pointed in which power migrates to the higher-frequency bands (see representative highly rated rounded and pointed pseudowords in Supplementary Figure 1a). These differences reflected our predictions regarding the contribution of the place of articulation to ratings, relative concentrations of low frequencies being associated with the bilabial sonorants in rounded pseudowords and relatively higher frequencies associated with the alveolar and post-alveolar/velar obstruents in the pointed pseudowords. The next strongest relationship (r = .28, p < .001) was for the FFT: the spectrograms show power variations across frequencies and time that are gradual for a representative rounded pseudoword but abrupt for a representative pointed pseudoword (Supplementary Figure 1b), again reflecting acoustic differences between sonorants and obstruents, respectively. The weakest correlation was for the speech envelope (r = .18, p < .001), which was smoother and more continuous for pseudowords rated as strongly rounded, compared to strongly pointed, where the speech envelope was discontinuous and uneven (illustrated for a representative pseudoword in Supplementary Figure 1c). This reflects the fact that there was sustained higher amplitude for rounded, compared to pointed, pseudowords: sonorants and back rounded vowels were higher in amplitude for longer portions of the pseudoword duration compared to unvoiced obstruents and front unrounded vowels.

For the voice parameters (Figure 4), the mean HNR, pulse number, and MAC were, as predicted, significantly negatively correlated with ratings of the pseudowords. For these parameters, lower values indicate less periodicity (i.e., a noisier signal and the presence of fewer voiced segments) and were associated with higher pointed ratings; higher values for these parameters indicate more periodicity and were associated with lower rounded ratings. In line with predictions, the FUF was significantly positively correlated with ratings, reflecting the greater presence of unvoiced segments (and therefore a decrease in periodicity) in items rated as more pointed. As predicted, shimmer and jitter were significantly positively correlated with ratings, reflecting increasing variability in voice quality, as ratings of the pseudowords transitioned from rounded to pointed. Thus far, the voice parameter results replicated our earlier study (Lacey et al., 2020: see Table 2) but, in the present study, and in contrast to our earlier study, pitch standard deviation was correlated significantly and positively with ratings, increasing as ratings moved from rounded to pointed – this indicates greater variability in pitch for pointed, compared to rounded, pseudowords. Thus, overall, ratings moved from rounded to pointed as noisiness and variability in the speech pattern increased. For the two new voice parameters that were not included in our earlier study, as predicted, there was a tendency for pseudowords to be rated as more pointed as pseudoword duration shortened: this likely reflects the phonetic content of the items (most rounded: /m mo/, 669 ms – most pointed: /tike/, 505 ms) as well as their vocal realization. However, while we expected that the mean pitch would increase from low to high, for pseudowords rated as rounded to those rated as pointed, respectively (see Parise & Pavani, 2011; Knoeferle et al., 2017), it was in fact uncorrelated with ratings; these findings contrast with Villegas et al. (2023) who suggest that roundedness is associated with higher pitch. Notably, the relationships between acoustic parameters and shape ratings were a very close replication of our previous findings, in an independent participant group (Table 2; Lacey et al., 2020). There was an additional replication in that, of the 10 most highly-rated pseudowords at each end of the scale, 9 (5 rounded and 4 pointed) were common to both the current results and those of McCormick et al. (2015: see Supplementary Table 1).

Weight

For the weight domain, the RDMs for all three spectro-temporal parameters were significantly positively correlated with the RDM for ratings of the pseudowords from light to heavy (Figure 5) but, in contrast to the shape domain, the strongest correlation was for the FFT RDM (r = .38, p < .001). As predicted, and consistent with our phonetic analysis (Lacey et al., 2024), the distribution of energy across frequencies over time was more diffuse for pseudowords rated as light, reflecting broadband frequency energy resulting from the combination of sonorants and front unrounded vowels (see the spectrograms for representative pseudowords in Supplementary Figure 2a). By contrast, pseudowords rated as heavy were characterized by more abrupt transitions in energy across frequencies and time, reflecting the associated voiced obstruents and back rounded vowels (Supplementary Figure 2a; see also Lacey et al., 2024). The correlation between the speech envelope and ratings RDMs was the next strongest (r = .28, p <.001): as predicted, the envelope was even and continuous for light pseudowords, reflecting the sustained higher amplitude of the sonorants involved, but uneven and discontinuous for the heavy pseudowords, reflecting the predominance of obstruents, including stops which involve a short period of low energy before a release of higher energy (see representative pseudowords in Supplementary Figure 2b; see also Lacey et al., 2024). Spectral tilt had the weakest relationship to the ratings in terms of between-RDM correlations (r = .19, p < .001) but reflected our prediction that energy would be distributed across low and high frequencies for both light and heavy pseudowords. Each of the representative pseudowords in Supplementary Figure 2c shows energy at both low and high frequencies.

In line with predictions, the mean HNR, pulse number, and MAC were significantly negatively correlated with ratings of the pseudowords (Figure 6) indicating that periodicity decreased as ratings transitioned from light to heavy. The FUF increased as ratings transitioned from light to heavy, also indicating decreasing periodicity as the proportion of unvoiced, aperiodic, segments increased. The relationship between measures of periodicity and weight ratings was consistent with our phonetic analysis, in which lightness ratings were associated with sonorants, which are always voiced; by contrast, obstruents, which can be either voiced or unvoiced, were strongly associated with heaviness ratings and the contribution of voicing was relatively small (Lacey et al., 2024). Mean pitch was significantly lower for heavy, compared to light, pseudowords, in line with prediction. Against our prediction, duration was negatively correlated with ratings, such that shorter items were rated as heavier than longer items. It could perhaps be argued that shorter words are ‘denser’ – since density is mass per unit volume, for two-syllable (mass) words, the shorter (unit volume) they are, the denser and therefore apparently heavier, they are. This would need to be confirmed by reference to ratings for an appropriate density scale. Although we had no specific predictions for shimmer, jitter and PSD, these were significantly positively correlated with ratings, reflecting greater variability in voice quality for pseudowords perceived as heavy, compared to those perceived as light.

Texture

The RDMs for all three spectro-temporal parameters were again significantly positively correlated with the RDM for ratings of the pseudowords from hard to soft (Figure 7). The FFT RDM showed the strongest correlation (r = .36, p < .001). As predicted, hard pseudowords and their associated obstruents exhibited more abrupt changes in power across frequency and time than soft pseudowords, which are associated with sonorants, where such changes were more gradual, and more energy was concentrated at the higher-frequency bands compared to the soft pseudowords (see representative spectrograms in Supplementary Figure 3a). The greater concentration of power at higher frequencies for hard, compared to soft, pseudowords also arises from their different places of articulation: post-alveolar/velar compared to bilabial/labiodental, respectively (Lacey et al., 2024). The RDM of the speech envelope was similarly strongly correlated (r = .33, p < .001); a higher-rated hard pseudoword exhibited a more discontinuous envelope than a higher-rated soft pseudoword, reflecting the sustained higher amplitude for the latter (Supplementary Figure 3b), consistent with the presence of obstruents and sonorants, respectively. The relationship between ratings and spectral tilt was weaker (r = .11, p < .001), with more high-frequency energy for the hard, compared to the soft, pseudowords (Supplementary Figure 3c). This also reflected place of articulation with relatively higher frequencies produced by post-alveolar/velar, compared to bilabial/labiodental, consonants respectively (Lacey et al., 2024).

For the voice parameters, as predicted, soft pseudowords were associated with greater periodicity and less vocal variability, while the reverse was true for the hard pseudowords. The mean HNR, pulse number, and MAC were significantly positively correlated with ratings, while FUF was significantly negatively correlated (Figure 8); thus, there was a higher percentage of voicing and increased periodicity (as found in sonorants) as ratings changed from hard to soft. Shimmer and jitter were significantly negatively correlated with texture ratings, meaning that as ratings moved from hard to soft, there was decreasing variability in amplitude and frequency, respectively. The positive correlation for duration indicated that, in line with expectation, shorter pseudowords were perceived as harder and longer pseudowords as softer. Texture ratings were uncorrelated with mean pitch; there was a weak effect in which pitch variability (PSD) decreased for softer pseudowords but this did not survive correction.

The reason for the association of hardness and softness with shorter and longer duration respectively, is not clear. Perhaps, as with the shape domain (see above) it is simply the preponderance of inherently shorter stop sounds in the harder pseudowords (hardest: /kike/, 533 ms) as compared with the concentration of longer sonorant sounds for the softer pseudowords (softest: /mumo/, 583 ms; Lacey et al., 2024). Alternatively, it may be that hardness and softness are associated with percussive events that are more or less auditorily distinct respectively (for example, a hammer striking a nail compared to hands plumping up a cushion), and that similarly vary in duration. Other dimensions of texture that could be sound symbolically mapped include rough/smooth and sticky/slippery, although the latter dimension has not been robustly established (Hollins et al., 2000). Rough ratings would likely be reflected in more vocal variability and a less periodic speech signal while smooth ratings would be associated with the reverse – less vocal variability and more periodicity.

Arousal

Significant positive correlations with the ratings RDM were found for all three spectro-temporal parameter RDMs (Figure 9) although they were weaker than those for the shape, weight, and texture domains. The highest correlation with the arousal ratings RDM was for the FFT RDM (r = .21, p < .001); as predicted, changes in the distribution of power across frequencies over time were more gradual for calming, compared to exciting, pseudowords where these were more abrupt (Supplementary Figure 4a). This is consistent with the predicted acoustic consequences for the sonorants and obstruents associated with calming and exciting pseudowords, respectively (Lacey et al., 2024). The spectrogram also shows more energy at the higher-frequency bands for the pseudowords rated as highly exciting, reflecting their unvoiced obstruents, particularly fricatives which result in some of the highest frequencies in speech (Ladofoged & Johnson, 2011). In contrast, the pseudowords rated as most calming show more energy at lower frequencies (Supplementary Figure 4a), reflecting their sonorants and the bilabial place of articulation which produces the lowest frequencies among the different places of articulation (Reetz & Jongman, 2020). The RDMs of the speech envelope (r = .18, p < .001) and spectral tilt (r = .12, p < .001) had lower correlations with the ratings RDM. Consistent with predictions, for calming pseudowords, the speech envelope was relatively continuous, reflecting the sustained higher amplitude of sonorants, and spectral tilt was steep, reflecting the lower frequencies associated with sonorants and the bilabial place of articulation (see Lacey et al., 2024). By contrast, for exciting pseudowords the speech envelope was more discontinuous, reflecting the ‘silence-release burst’ profile of stops (Ladofoged & Johnson, 2011), and spectral tilt was flatter as energy was dispersed across all frequency bands, including the higher frequencies produced by fricatives (Supplementary Figure 4b and 4c show the speech envelope and spectral tilt, respectively, for representative pseudowords).

The mean HNR, pulse number, and MAC were negatively correlated with ratings, indicating that periodicity decreased as ratings transitioned from calming to exciting (Figure 10). As predicted, FUF was positively correlated with ratings, meaning that the number of unvoiced segments increased from calming to exciting. These measures of periodicity presumably reflect the change from voiced to unvoiced consonants for pseudowords going from calming to exciting. Increasing variability in frequency and amplitude from cycle to cycle of the speech waveform, as measured by jitter and shimmer respectively, reflected the change in ratings from calming to exciting. However, PSD, mean pitch, and duration were not significantly correlated with arousal ratings.

Although some predictions for the arousal domain could have been made from voice stress analyses (e.g., Van Puyvelde et al., 2018), a cautionary note is that this previous work was carried out with real-life emergency communications where there is a need to speak more clearly in order to avoid miscommunication and error and also presumably more significant levels of arousal in those contexts, regardless of phonetic content. We found that jitter and shimmer increased with increasing arousal, in line with the finding that general arousal produces an increase in these vocal measures (Van Puyvelde et al., 2018). However, in the context of emergency communications, speaking more clearly and loudly might result in a decrease in such vocal variability (see Brockmann-Bauser et al., 2018). An additional point is that our arousal scale was positively valenced as a whole: both ‘calm’ and ‘exciting’ are in the top 10% of items at the positive end of the valence scale in Warriner et al. (2013: see also the Global Observations section below). A restful/stressful scale might have produced a clearer delineation between positive and negative dimensions.

Valence

Weak, but significant, positive correlations were found for all three spectro-temporal parameter RDMs with the ratings RDM (Figure 11). The FFT RDM had the highest correlation of the three (r = .17, p < .001) and, as predicted, the spectrogram showed changes in the distribution of power across frequency over time that were more diffuse for the good, compared to the bad, pseudowords, reflecting the associated sonorants and obstruents, respectively (Supplementary Figure 5a; Lacey et al., 2024). Similarly, the speech envelope, whose RDM had the next highest correlation (r = .16, p < .001), was somewhat more continuous for good, compared to bad, pseudowords (Supplementary Figure 5b), reflecting that sonorants are generally higher in amplitude for longer durations than obstruents. We had no specific prediction for spectral tilt whose RDM, indeed, showed a very weak relationship with the ratings RDM (r = .07, p < .001) reflecting a relatively flat spectral tilt whether pseudowords were rated as bad or bad (Supplementary Figure 5c).

For voice parameters, the mean HNR, pulse number, and MAC were significantly negatively correlated with ratings, and FUF was significantly positively correlated with ratings, consistent with our prediction that periodicity would decrease, and therefore noise would increase, as ratings transitioned from good to bad (Figure 12). Shimmer, jitter, and PSD were significantly positively correlated with ratings, reflecting increasing variability in voice quality for pseudowords perceived as bad, compared to those perceived as good. A significant negative correlation with ratings indicated that mean pitch was lower for bad, compared to good, pseudowords. As expected, ratings were not significantly correlated with pseudoword duration.

Like the arousal domain, the valence domain could have been interpreted in several different ways other than the good/bad dimension (see the Limitations section in the General Discussion). Whether different aspects of the same domain share common phonetic and/or acoustic characteristics would be an interesting area for further research. Emotional state is another context in which both valence and arousal are relevant and both may increase from, for example, sad to happy. In this case, PSD might be sound symbolically related. PSD is a measure of the variation in the fundamental frequency and thus of vocal inflection (Kliper et al., 2016): low PSD manifests as a monotone voice and higher PSD as a livelier voice, each reflecting the speaker’s emotional state (Kliper et al., 2016).

Brightness

As for the valence domain, only weak, albeit significant, correlations were found for all three spectro-temporal parameter RDMs with the ratings RDM (Figure 13). Nonetheless, in line with predictions, spectral tilt (r = .14, p < .001) was flatter for brighter pseudowords, reflecting more high-frequency energy, whereas for dark pseudowords, power was concentrated in the lower-frequency bands (Supplementary Figure 6a). This likely reflected the combination of unvoiced consonants and front unrounded vowels for bright pseudowords, and voiced consonants and back rounded vowels for dark pseudowords (Lacey et al., 2024). Even weaker relationships were found for the FFT (r = .08, p < .001) and speech envelope (r = .04, p < .001) RDMs. For the FFT, although the spectrogram did show greater power in the higher frequencies for bright, compared to dark, pseudowords (consistent with the difference in spectral tilt), there were relatively small differences in the abruptness of power variations across frequencies (Supplementary Figure 6b), reflecting that the main difference between bright and dark pseudowords was voicing rather than manner of articulation (see Supplementary Table 1; Lacey et al., 2024). The speech envelope was discontinuous for both bright and dark pseudowords but amplitude was perhaps greater for the latter (Supplementary Figure 6c), reflecting that bright/dark pseudowords tend to involve unvoiced/voiced consonants, respectively (Supplementary Table 1; Lacey et al., 2024).

Compared to the preceding domains, there were fewer voice parameters that were significantly correlated with brightness ratings (Figure 14). As predicted, FUF decreased significantly as ratings increased towards the dark end of the scale and mean HNR and pulse number were significantly positively correlated with ratings such that voice quality became more periodic and smoother for pseudowords that were rated as darker, i.e. in dark pseudowords, the proportion of voiced segments, and therefore degree of periodicity in the auditory pseudowords increased. In line with predictions, high/low mean pitch were associated with bright/dark pseudowords, respectively (see Tzeng et al., 2018). There was a small effect of duration in which longer pseudowords were rated as darker; this was consistent with Tzeng et al. (2018) but did not survive correction. There were also weak relationships for shimmer and MAC that did not survive correction, and non-significant relationships were found for PSD and jitter.

Size

There was a weak correlation between the RDMs for spectral tilt and size ratings (r = .13, p < .001: Figure 15). Spectral tilt was broadly similar for both small and big pseudowords, the slope of spectral density being initially steep before flattening out (Supplementary Figure 7a), although for big pseudowords, broadband energy remained relatively higher across the frequency range. As was found for brightness, there were weaker relationships with the size ratings RDM for the speech envelope (r = .11, p < .001) and FFT (r = .08, p < .001) RDMs (Figure 15). Both small and big pseudowords showed discontinuous speech envelopes, reflecting the presence of stops in both, the only difference being that big pseudowords showed somewhat greater amplitude for the second syllable, compared to the small pseudowords (Supplementary Figure 7b). Likewise, there were relatively distinct changes in energy for both small and big pseudowords, although there was somewhat more energy in the higher frequencies for the small, compared to the big, pseudowords, perhaps reflecting the contrast between unvoiced and voiced consonants respectively (see Supplementary Figure 7c; Lacey et al., 2024). Although the correlations between ratings and all three spectro-temporal parameters were significant, the relationships were weak. Visual inspection of the RDMs (Figure 15) shows that the effects were diffuse, with higher dissimilarity values appearing throughout the RDM, especially for spectral tilt. Thus, it is hard to say how these parameters relate to sound-symbolic size (see the Limitations section in the General Discussion).

There were no significant correlations between voice parameters and size ratings that survived correction (Figure 16); in particular, we did not replicate the finding that short and long pseudoword duration reflect small and large size respectively (Knoeferle et al., 2017).

Global observations

As stated above, results were Bonferroni-corrected for three tests in the case of the spectro-temporal parameters, and nine tests for the vocal parameters. Since the Bonferroni correction is relatively conservative, aimed at minimizing Type 1 error (erroneously claiming a real effect where none exists), we also carried out the less restrictive stepwise Bonferroni correction suggested by Holm (1979). However, the pattern of results was unchanged, except that the correlation between texture ratings and PSD was significant under the Holm correction. This suggests that, overall, Type 2 error (erroneously rejecting a real effect where it does, in fact, exist) was also unlikely.

The relationships between the RDMs for ratings and the spectro-temporal parameters can be subtle. While greater pairwise dissimilarity for ratings mostly clustered at one end of the scale, as expected (see the top right of the ratings RDMs in Figures 3, 5, 7, 9, 13, and 15), this was not so for the valence ratings RDM in which the dissimilarities are difficult to discern (Figure 11). Concentrations of greater pairwise spectro-temporal dissimilarity, approximately matching those in the ratings RDMs, were clearly seen for spectral tilt in the case of shape (Figure 3) and brightness (Figure 13), as well as for the FFT and, to some extent, the speech envelope, for texture (Figure 7); these generally reflected stronger correlations with the corresponding ratings RDM. In other cases, however, high dissimilarity values occurred throughout the spectro-temporal RDM; for example, spectral tilt for weight (Figure 5) and texture (Figure 7); these generally reflected weaker relationships with the corresponding ratings RDM. Thus, RDMs need to be carefully inspected in order to interpret the relationship between ratings and parameter values.

Because the same pseudowords were rated across seven different domains, we examined the 10 most highly-rated pseudowords for the single, re-coded, scale in each domain (e.g., most rounded and most pointed). It was clear that these highly-rated pseudowords fell into two distinct groups, representing one end of the scale for each domain, that shared phonetic and acoustic features. The top ten pseudowords for the rounded, heavy, soft, calming, good, dark, and big scales were largely composed of sonorants, voiced stops and af/fricatives, and back rounded vowels (Supplementary Table 1A; see also Lacey et al., 2024, for a detailed phonetic analysis). Acoustically, these pseudowords were broadly associated with more distributed variations in power across frequencies, a generally steeper spectral tilt, and a smoother, more continuous speech envelope (see Supplementary Figures 1–7 for examples), together with more periodic voiced segments, and less variability. For the pointed, light, hard, exciting, bad, bright, and small scales, however, the most highly-rated pseudowords were largely composed of unvoiced stops and af/fricatives, and front unrounded vowels (Supplementary Table 1B; see also Lacey et al., 2024). Acoustically, these pseudowords exhibited more abrupt power variations with frequency, a generally flatter spectral tilt, and a discontinuous speech envelope (see Supplementary Figures 1–7 for examples), together with less periodicity and voicing, and more variability. Despite the fact that, across domains, each of the two groups shared many phonetic and acoustic features, very few pseudowords occurred in more than one domain in the top ten list, so that most pseudowords were unique in each domain: 74.3% in the first group (Supplementary Table 1A) and 72.9% in the second group (Supplementary Table 1B), suggesting that these mappings were domain-specific.

Finally, as an additional test of domain-general vs domain-specific accounts of sound symbolism, we considered the words used as the endpoints for each domain – i.e., rounded/pointed, light/heavy, hard/soft, and so on. (These appeared with the rating scale on every trial and therefore participants were most frequently exposed to these compared to the additional, explanatory, anchor words that only appeared in the task instructions.) We extracted the ratings for these endpoint words on scales reflecting arousal (calm/excited), valence (unhappy/happy) and dominance (controlled/in control) from a large-scale rating study of nearly 14,000 words (Warriner et al., 2013). This allowed us to place the endpoints of our domains on continua reflecting two potential domain-general factors and one irrelevant factor (to our knowledge, dominance has not been advanced as an explanation for sound-symbolic associations, and thus serves as a reference condition) according to an independent dataset.

With the endpoints of each domain grouped according to their phonetic and acoustic similarity as described earlier, we reasoned that, if a domain-general account were true, each group of endpoints should have similar ratings of either arousal or valence, but vary in the level of the irrelevant dominance continuum. In fact, this was not the case and, despite similarities in phonetic and acoustic features, each group varied considerably on all three continua (Supplementary Table 2). It follows that, whereas grouping by acoustic similarity resulted in dissimilarities in arousal and valence, grouping by similarity in arousal or valence would have resulted in acoustic dissimilarities across domains.

We examined acoustic dissimilarities in more detail by looking at the specific relationships between those acoustic parameters that were significantly correlated with ratings and the differing levels of arousal, valence, and dominance. For the spectro-temporal parameters, this included all seven domains, but only six domains for the vocal parameters since none of these was significantly correlated with ratings for size. Again, we reasoned that if a domain-general account were true, then the acoustic relationship should be consistent across the differing levels of the underlying factor; i.e., the high arousal endpoint for each domain should always be reflected in a flatter spectral tilt, lower HNR and so on.

For the spectro-temporal parameters, although the high arousal endpoints were consistently associated, across all domains, with flatter spectral tilt, abrupt transitions in the spectrogram of the FFT, and discontinuous speech envelopes, the low arousal endpoints could be associated with either steep or flat spectral tilt, gradual or abrupt transitions, and continuous or discontinuous speech envelopes depending on the domain (Supplementary Table 3A). Similarly, there were no consistent relationships across all domains between positive and negative valence for any of the spectro-temporal parameters (Supplementary Table 4A) and the same was true for the reference condition of dominance (Supplementary Table 5A).

By contrast, vocal parameters that reflected periodicity and vocal variability were consistently associated with levels of arousal. Across all six domains, low/high arousal endpoints were associated with high/low HNR, pulse number, MAC, and low/high FUF, respectively, together with low/high jitter, shimmer, and PSD, respectively (Supplementary Table 3B & C). For valence, there was a high degree of consistency for measures of periodicity, positive/negative valence being associated with high/low HNR, pulse number, MAC, and low/high FUF respectively for all domains except arousal and brightness (Supplementary Table 4B). Additionally, for the vocal variability parameters, low/high jitter and shimmer were associated with positive/negative endpoints respectively, for the shape, weight, texture, and valence domains, while the arousal domain showed the reverse relationship (the variability parameters were uncorrelated with brightness, in addition to size, ratings: Supplementary Table 4C). There was also a relatively high degree of consistency for the reference condition of dominance: low/high dominance was associated with low/high periodicity, respectively, for all domains except shape and brightness, for which the periodicity was reversed (Supplementary Table 5B). Similarly, low/high dominance was reflected in high/low jitter and shimmer, respectively, for the weight, texture, arousal, and valence domains, but was reversed for the shape domain (Supplementary Table 5C). Across all domains, no firm conclusions could be drawn about PSD since this was significantly correlated with ratings for so few domains (Supplementary Tables 3C, 4C & 5C) and the same was true for the other parameters of mean pitch and duration (Supplementary Tables 3D, 4D & 5D).

GENERAL DISCUSSION

Speech perception is underpinned by an acoustic-phonetic speech representation tuned to the spectro-temporal profiles of specific phonetic categories (Mesgarani et al., 2014; see also Oganian et al., 2023; Hamilton et al., 2020). While the phonetic features associated with a few sound-symbolic mappings have been extensively investigated, the associated acoustic features have received less attention (Knoeferle et al., 2017). This study builds on our earlier work showing the relationship of selected spectro-temporal and vocal parameters of spoken pseudowords to sound-symbolic ratings of shape (Lacey et al., 2020). Here we replicated our earlier findings for shape (Table 2) in a different participant group with the inclusion of two additional vocal parameters not featured in our earlier study – the duration and mean pitch of the pseudoword – and extended our approach to six additional domains of meaning.

Domain-general vs domain-specific accounts of sound symbolism

Examining multiple domains allowed us to distinguish between two competing accounts of how sound-symbolic mappings arise. The first, and simplest, is that all sound-symbolic mappings reflect a single factor of association to an abstract domain such as arousal, magnitude, or valence (Aryani et al., 2020; Sidhu & Pexman, 2018; Spence, 2011). On this account, one end of any dimension would evoke more arousal than the other; for example, pointed or bright could be more arousing concepts than rounded or dark. Aryani et al. (2020) found just such an association for the rounded-pointed dimension of shape, but only generalized this to a different set of rounded-pointed pseudowords, rather than to other, non-shape, domains. The present study tested this account, not only by examining multiple domains but also by including the calming-exciting dimension of the arousal domain. If arousal underlies all sound-symbolic mappings, then the acoustic profile (the pattern of the relative importance of the acoustic parameters) of the arousal domain should be preserved across all the other domains; our results argue against this hypothesis. For the arousal domain, RSA for the spectro-temporal parameters showed that the FFT is the most, and spectral tilt the least, strongly correlated with ratings; this pattern was also true for the weight, texture, and valence domains. However, the shape and brightness domains showed a different pattern in which spectral tilt was most strongly correlated with ratings, followed by the FFT and the speech envelope (Table 3). The size domain showed a third pattern, although this must be caveated as to whether the pseudowords were suited to sound symbolism for the size domain (see Limitations section below). The same result applies to a domain-general account based on valence: the acoustic profile of the valence domain was not preserved across the other domains (Table 3).

As Table 3 shows, even domains that shared similar patterns for the spectro-temporal parameters had very different patterns for the vocal properties. Moreover, while some vocal parameters contributed to certain domains (for example, duration was relevant to shape, weight, and texture; and mean pitch to weight, valence, and brightness), they did not contribute at all to others (duration was unrelated to arousal, valence, or brightness, while mean pitch was unrelated to shape, texture, or arousal). Even when sound-symbolic ratings were correlated between domains (for example, pointed ratings were positively correlated with hard ratings but negatively with soft ratings, while the reverse was true for rounded ratings: see Lacey et al., 2024), the pattern of acoustic parameters could still differ (see the shape and texture domains in Table 3). For the same reasons, a domain-general account based on valence or magnitude (at least for size-based magnitude, and subject to the limitations noted below), rather than arousal, is also excluded by these results. Moreover, evaluating the opposing ends of our domains for their level of arousal and valence by reference to an independent dataset (Warriner et al., 2013; see Results – Global Observations, above) showed that grouping the endpoints by their acoustic similarity revealed varying patterns of arousal and valence across domains (Supplementary Table 2). In both cases, the resulting patterns were just as variable as those shown in relation to dominance, a factor that (so far as we know) is irrelevant to sound symbolism.

We also examined the specific relationships between those acoustic parameters that were significantly correlated with ratings and the differing levels of arousal, valence, and dominance (see Results – Global Observations, above). Overall, the picture that emerges from this exercise is not one that supports a domain-general account of sound symbolism. Across all domains, the spectro-temporal parameters, probably the most sensitive since they sample the entire pseudoword, were not systematically associated with low/high arousal or positive/negative valence, nor with the reference condition of dominance (Supplementary Tables 3A, 4A & 5A). The vocal parameters offered more systematic relationships, particularly with levels of arousal (Supplementary Tables 3B & C), but the same was largely true for a valence-based account (Supplementary Tables 4B & C). In any case, it seems implausible that a domain-general factor underlying all sound-symbolic mappings would be reflected in vocal, but not spectro-temporal, parameters.

Thus, none of these abstract dimensions appear to explain global patterns of associations between sound and meaning domains, and although we are only able to explicitly test the domain-general hypothesis in relation to arousal and valence, the fact that the acoustic profile of spectro-temporal and vocal parameters varies widely across domains suggests that other single-factor accounts are also likely unsupported. Thus, these analyses provided converging evidence against the hypothesis that a domain-general factor underlies any sound-symbolic domain. Instead, our results are consistent with the second hypothesis: that sound-symbolic mappings arise from domain-specific relationships with acoustic parameters. This is in keeping with domain-specific differences in the phonetic features contributing to each domain (Lacey et al., 2024) and the fact that even though, across domains, the ends of various scales shared broadly similar spectro-temporal characteristics (see Global Observations above), the individual most highly-rated pseudowords were largely unique and specific to each domain with few repetitions (Lacey et al., 2024, and see Supplementary Table 1).

Finally, it is also worth noting that although we chose domains that reflected both sensory and abstract meanings, this too did not appear to be an organizing principle of sound-symbolic mappings. The abstract domains of arousal and valence shared a pattern of spectro-temporal relationships with the sensory domains of weight and texture, but these patterns differed from those for shape, brightness, and size (Table 3). Considering the sensory domains alone, the spectro-temporal patterns for the modality-specific texture dimension of soft-hard (haptic) and brightness (visual) domains differed from each other, as did the patterns for the (visuo-haptic) shape and size domains and the (haptic) weight domain (Table 3); thus, these did not seem to be organized in any principled way. In any case, the patterns for the voice parameters were specific to each domain (Table 3).

General comments

Sound symbolism is generally investigated using pseudowords rather than real words. This avoids the confound of prior semantic knowledge for real words, i.e., that the sound-symbolic rating of a word like ‘balloon’ would be influenced by knowing that ‘balloon’ typically means something round in shape. Pseudowords also avoid the possibility that, for some real words, sound-symbolic effects could be diluted by the presence of non-sound-symbolic (e.g., etymological) components, for example, the phonaestheme /gl/ in words like ‘glitter’ or ‘gloom’ (Bergen, 2004) that are otherwise sound symbolically bright or dark, respectively (see Lacey et al., 2024). Recent work has shown that the acoustic-phonetic representation of speech in the superior temporal gyrus is based on tuning to specific spectro-temporal profiles for different phonetic features (Mesgarani et al., 2014; see also Oganian et al., 2023; Hamilton et al., 2020) and that cortical responses to speech can be decoded from such relatively low-level acoustic features (Daube et al., 2019). Together with the current findings, this raises the question whether the sound-symbolic nature of real words could be assessed from their acoustic characteristics, thus avoiding the potential confound of these non-sound-symbolic components. One way to do this would be to assign ratings to pseudowords from spectro-temporal and voice parameters and to generalize these to real words independent of semantic knowledge (for an initial proof-of-concept, predicting the sound-symbolic ratings of pseudowords from acoustic parameters, see Kumar et al., 2024).

As noted above, the ten most highly rated pseudowords are largely unique to a single domain (Table 4) and, for the shape domain, exhibit a level of reliability in that nine of them replicate from the most highly rated items found by McCormick et al. (2015). Although these pseudowords share similar spectro-temporal profiles at one end of a scale, their specificity to a particular domain arises from small changes in their phonetic composition; for example, /mulu/ (soft) compared to /munu/ (calming), or /teki/ (pointed) compared to /peki/ (bright). While the evolution of the larynx as facilitating the emergence of language is not a new idea (see Hill, 1972, for example), we can make a speculative connection between the minimal overlap across sound-symbolic domains found here and recent anatomical work on the larynx. Paradoxically, it was simplification of the larynx, specifically the loss of the vocal membranes, that likely facilitated the complex phenomenon of language (Nishimura et al., 2022). The vocal membranes contribute to instability in vocalization and their loss resulted in stable, harmonic-rich, vocalization that is important for producing the variations in formant frequencies that carry most phonetic information (Nishimura et al., 2022; see also Gouzoules, 2022). Equally importantly, in addition to the ability to produce a wider range of fundamental frequencies, harmonics and formant frequency variation, this development would also have produced greater consistency of utterances. Such reliability is essential for language: the same utterance produced the same way each time would result in more stable acoustic-phonetic realizations, which in turn are associated with meaning. With such additional precision in articulating, the utterances /m mo/ (round) and /n mo/ (calming), for example, could be reliably differentiated. We can speculate that it is this development in laryngeal anatomy that enabled the proliferation of sound-symbolic utterances – also suggested as a precursor to language (Swadesh, 1971) – into multiple domains via slight variations in speech sounds of the kind shown in Supplementary Table 1; i.e., evolutionary changes that improved vocal and articulatory precision could have resulted in more nuanced iconic sound-to-meaning associations.

Limitations

One limitation that might be raised for the shape domain is that the speaker who recorded the pseudowords might have unconsciously pronounced them differently according to their expectations about the pseudowords’ roundedness/pointedness (see also Lacey et al., 2020). This is unlikely because the speaker was instructed to employ a neutral intonation and two independent judges rejected items that did not sound both neutral and consistent with the other recordings. Furthermore, it would be difficult to manipulate complex vocal parameters like jitter or shimmer in such a way as to produce the correlations seen in Figure 4, particularly over more than 500 items when these were recorded in random order. For the remaining six domains reported here, the speaker did not know that their recordings would be used for other domains in the future; therefore, they cannot have been affected by prior knowledge or expectations about those domains.

The RSA approach has advantages and disadvantages. On the one hand, it facilitates the comparison of spectro-temporal and other parameters that consist of multiple measurements or samples per pseudoword and that are not amenable to conventional correlation analysis. RSA also compares every item to every other item, thus all possible pseudoword pairs enter the analysis whereas conventional correlation analyses only treat items as a group. However, a disadvantage is that, for any parameter expressed as a single measurement per item, pseudowords would have to be binned into groups resulting in a loss of sensitivity (see, for example, this approach for the voice parameters in Lacey et al., 2020). We should also mention that, because RSA accounts for all possible pseudoword pairs, the sample size quickly becomes extremely large: more than 143,000 pairs for the current set of only 537 items. Thus, very small effects can reach significance without necessarily being meaningful. For example, while the spectral tilt RDM for brightness shows a small correlation with the brightness ratings RDM with clear differences between bright and dark (Figure 13), the weight and size RDMs (Figures 5 and 15) show similarly small correlations with the corresponding ratings RDMs but with diffuse patterns of dissimilarity that are hard to interpret.

The results for the size domain were inconclusive: the spectro-temporal parameters were only weakly related to ratings and none of the vocal parameters showed a significant association. One reason for these results may be that the pseudowords were optimized for the shape, rather than the size, domain. This did not seem to prevent participants making reasonably organized ratings as a gradual increase in dissimilarity can be seen in the ratings RDM (Figure 15) and non-optimization was also true of all the other non-shape domains in which meaningful relationships were found. However, a different set of pseudowords, using the same phonemes as here but with the addition of the vowels / /, / /, and / /, that have sound-symbolic associations to size, does show significant relationships for some vocal parameters (Nayak et al., 2023), thus the size domain requires further examination.

Finally, much may depend on the definition of the domain. As noted above, we might have found different results if arousal had been defined as restful vs stressful rather than calming vs exciting, with potentially a greater role for jitter, shimmer, and variability in pitch (Van Puyvelde et al., 2018; Kliper et al., 2016). Likewise, for texture, rough vs smooth might produce different relationships than hard vs soft since these are independent dimensions of tactile texture (see Hollins et al., 2000). Similarly, good vs bad is a rather ‘all-purpose’ conceptualization of valence and doesn’t necessarily have a more specific moral dimension, e.g. we could have tested virtuous/wicked or honest/dishonest (or even rude/polite given the sound-symbolic associations recently reported for profanity [Lev-Ari et al., 2023]); we could also have tested emotional valence using a sad/happy scale (see Warriner et al., 2013). These different dimensions of valence may have different sound-symbolic profiles. For example, motivationally positive emotions are vocalized with both higher pitch and greater amplitude than motivationally negative emotions, while esthetically positive emotions were vocalized with both lower pitch and amplitude than their negative counterparts (Belyk & Brown, 2014). Furthermore, morally positive emotions are vocalized with higher pitch but lower amplitude than morally negative emotions (Belyk & Brown, 2014). These examples show that any domain may comprise multiple related dimensions with different acoustic profiles. Furthermore, domains may not be orthogonal to each other: as noted above, although the calming/exciting dimension differed in the level of arousal, it was potentially confounded with valence to some extent because the endpoints were both positively valenced. These issues are worth exploring in relation to how different speech sounds differentiate between related concepts.

Conclusions

The findings of the present study build on our prior work in the sound-symbolic shape domain (Lacey et al., 2020) and demonstrate that different combinations of acoustic parameters underlie sound-symbolic mappings depending on the domain of meaning. These findings underscore the robustness of sound symbolism as a core feature of language and, possibly, proto-language, and set the stage for further investigation of its neural basis.

Supplementary Material

Supplement 1

ACKNOWLEDGMENTS

This work was supported by grants to KS and LCN from the National Eye Institute at the NIH (R01EY025978) and the Emory University Research Council. Earlier versions of these data were presented at the 2022 meeting of the Cognitive Neuroscience Society (Hoffmann et al., 2022; Nygaard et al., 2022), and an earlier version of the paper was available as a preprint at bioRxiv, doi: [TBA].

DATA AVAILABILITY

Stimuli are available at https://osf.io/ekpgh/ and rating data and scripts for the spectro-temporal analyses are available at https://osf.io/y9zjc/.

Figure 1: Visual depictions of two pseudowords characterized by stops and high frequency vowels (/kike/, left column) and sonorants and low frequency vowels (/mumo/, right column) as reflected in (a) the speech envelope, (b) spectral tilt: frequency is low to high, left to right, and (c) spectrogram: frequency is low to high, bottom to top; shading = power, low power is lighter and high power is darker.

Figure 2: Schematic illustration of the analysis pipeline. Step 1: auditory pseudowords were rated along different dimensions, e.g., shape (rounded/pointed), brightness (bright/dark). Step 2: these perceptual ratings were used to create reference representational dissimilarity matrices (RDMs). Step 3: RDMs for each of the spectro-temporal parameters were created for each domain. Step 4: in each domain, the spectro-temporal parameter RDMs were compared to the ratings RDM by way of a second-order correlation.

Figure 3: Correlations between RDMs for auditory perceptual ratings for shape and for the spectro-temporal parameters of spectral tilt, FFT, and speech envelope. Pseudowords are ordered, left to right, from most rounded to most pointed. Spectro-temporal parameter correlations are ordered, top to bottom, left to right, from the strongest to the weakest relationship. Color bar shows pairwise dissimilarity (see text for details), where 0 = zero dissimilarity (items are identical) and 2 = maximum dissimilarity (items are completely different). r = Spearman correlation coefficient for second-order correlation between the parameter RDM and ratings RDM; Bonferroni-corrected α = .0167 for three tests; df = 143,914 in all cases.

Figure 4: Correlations between auditory perceptual ratings for shape and the voice quality parameters. Pseudowords are ordered as in Fig. 3. Voice parameter correlations are ordered, top to bottom, left to right, from the strongest to the weakest relationship, whether positive or negative. Note that the rating scale is truncated in all panels because there are no values < 2.5 or > 5.5; the mean autocorrelation scale is truncated because there are no values < 0.7. r = Pearson correlation coefficient; *correlation passes the Bonferroni-corrected α of 0.0056 for nine tests; df = 535 in all cases.

Figure 5: Correlations between RDMs for auditory perceptual ratings for weight and for the spectro-temporal parameters. Pseudowords are ordered, left to right, from lightest to heaviest. All other details as for Fig. 3.

Figure 6: Correlations between auditory perceptual ratings for weight and the voice quality parameters. Pseudowords are ordered as in Fig. 5; all other details as for Fig. 4

Figure 7: Correlations between RDMs for auditory perceptual ratings for texture and for the spectro-temporal parameters. Pseudowords are ordered, left to right, from hardest to softest. All other details as for Fig. 3.

Figure 8: Correlation between auditory perceptual ratings for texture and the voice quality parameters. Pseudowords are ordered as in Fig. 7; all other details as for Fig. 4.

Figure 9: Correlations between RDMs for auditory perceptual ratings for arousal and for the spectro-temporal parameters. Pseudowords are ordered, left to right, from most calming to most exciting. All other details as for Fig. 3.

Figure 10: Correlations between auditory perceptual ratings for arousal and the voice quality parameters. Pseudowords are ordered as in Fig. 9. Note that the rating scale is truncated in all panels because there are no values < 2.5 or > 5.0. All other details as for Fig. 4.

Figure 11: Correlations between RDMs for auditory perceptual ratings for valence and for the spectro-temporal parameters. Pseudowords are ordered, left to right, from good to bad. All other details as for Fig. 3.

Figure 12: Correlations between auditory perceptual ratings for valence and the voice quality parameters. Pseudowords are ordered as in Fig. 11. All other details as for Fig. 10.

Figure 13: Correlations between RDMs for auditory perceptual ratings for brightness and for the spectro-temporal parameters. Pseudowords are ordered, left to right, from brightest to darkest. All other details as for Fig. 3.

Figure 14: Correlations between auditory perceptual ratings for brightness and the voice quality parameters. Pseudowords are ordered as in Fig. 13. All other details as for Fig. 4.

Figure 15: Correlations between RDMs for auditory perceptual ratings for size and for the spectro-temporal parameters. Pseudowords are ordered, left to right, from smallest to biggest. All other details as for Fig. 3.

Figure 16: Correlations between auditory perceptual ratings for size and the voice quality parameters. Pseudowords are ordered as in Fig. 15. All other details as for Fig. 4.

Table 1: Phonetic and articulatory features of consonants and vowels used in creating the CVCV pseudowords. Voiced consonants are in bold type; vowels marked

Consonants			Place of articulation		
	Bilabial/labiodental		Alveolar	Post-alveolar/velar	
		
Manner of articulation		
Sonorants	/m/		/I/, /n/		
Stops	/b/, /p/		/d/, /t/	/g/, /k/	
Af/fricatives	/f/, /v/		/s/, /z/	/ /,/ /	
Vowels		Height			
	Mid		High		
			
Rounding (backness)	
Rounded (back)	/o/		/U/, / / *		
Unrounded (front)	/e/, / /*		/i/, /I/*		
* only occur in the first vowel position as they do not occur in word-final positions in English. Af/fricatives: affricates and fricatives treated as a single category.

Table 2: Correlations between (a) spectro-temporal and (b) voice parameters and ratings of pseudowords for the shape domain. A comparison between results previously reported in Lacey et al. (2020) and the current results shows an almost exact replication, apart from a switch in position between the pulse number and FUF. r-value, Spearman’s r; * parameter was not measured in the previous report; non-significant parameters are greyed-out.

	Previous results (Lacey et al., 2020)	Current results	
	
		r-value		r-value	
(a)	Spectral tilt	0.43	Spectral tilt	0.45	
	FFT	0.25	FFT	0.28	
	Speech envelope	0.14	Speech envelope	0.18	
(b)	HNR	−0.51	HNR	−0.58	
	Pulse number	−0.41	FUF	0.5	
	FUF	0.38	Pulse number	−0.48	
	MAC	−0.36	MAC	−0.42	
	Shimmer	0.34	Shimmer	0.39	
	Jitter	0.24	Jitter	0.3	
	PSD	0.1	PSD	0.17	
			*Duration	−0.14	
			*Mean pitch	0.03	

Table 3: Summary of the different patterns of relationships of (a) spectro-temporal and (b) voice parameters to pseudoword ratings in each domain. Parameters are listed in descending order of correlation strength in each section. Abbreviations as given in the text; non-significant parameters are greyed-out.

	Shape	Weight	Texture	Arousal	Valence	Brightness	Size	
		
(a)	Spectral tilt	FFT	FFT	FFT	FFT	Spectral tilt	Spectral tilt	
	FFT	Speech envelope	Speech envelope	Speech envelope	Speech envelope	FFT	Speech envelope	
	Speech envelope	Spectral tilt	Spectral tilt	Spectral tilt	Spectral tilt	Speech envelope	FFT	
(b)	HNR	HNR	HNR	HNR	HNR	FUF	Mean pitch	
	FUF	Pulse number	Pulse number	FUF	Pulse number	HNR	Duration	
	Pulse number	MAC	MAC	Pulse number	FUF	Mean pitch	HNR	
	MAC	Shimmer	Duration	Shimmer	Shimmer	Pulse number	Shimmer	
	Shimmer	Jitter	FUF	MAC	MAC	Shimmer	FUF	
	Jitter	Duration	Shimmer	Jitter	Jitter	MAC	Jitter	
	PSD	FUF	Jitter	PSD	PSD	PSD	PSD	
	Duration	PSD	PSD	Mean pitch	Mean pitch	Jitter	MAC	
	Mean pitch	Mean pitch	Mean pitch	Duration	Duration	Duration	Pulse number	

1 Broadly speaking, vowels can be identified by the relative frequencies of their formants – the resonance frequencies of the vocal tract when producing the vowel sound – and to a lesser extent by their characteristic fundamental frequency (F0). The first three formants, F1-F3, are the most informative about vowel identity, while higher formants are informative about speaker identity (Knoeferle et al., 2017).

2 Sharpness refers to the high frequency content of a sound and can be thought of as the “bass/treble ratio” (Villegas et al., 2023, p2); sharper sounds have a greater proportion of high frequencies (Fastl & Zwicker, 2007).

3 Stops, fricatives, and affricates are collectively referred to ‘obstruents’ because producing them involves an obstruction in the airflow (Reetz & Jongman, 2020); we will use this collective term going forward to avoid unnecessary repetition.

4 Since the RDM values are symmetrical across the diagonal, the second-order correlations between matrices were calculated using one half of the off-diagonal data, rather than the entire matrix, which avoids artificially inflating the degrees of freedom.
==== Refs
REFERENCES

Aiken S. J. , & Picton T. W. (2008). Human cortical responses to the speech envelope. Ear & Hearing, 29 :139–157.18595182
Akita K. (2021). Phonation types matter in sound symbolism. Cognitive Science, 45 :e12982.34018216
Anwyl-Irvine A. L. , Massonnié J. , Flitton A. , Kirkham N. , & Evershed J. K. (2020). Gorilla in our midst: An online behavioral experiment builder. Behavior Research Methods, 52 :388–407.31016684
Aryani A. , Isbilen E.S. & Christiansen M.H. (2020). Affective arousal links sound to meaning. Psychological Science, 31 :978–986.32662741
Belyk M. & Brown S. (2014). The acoustic correlates of valence depend on emotion family. Journal of Voice, 28 :523.e9–523.e18.
Bergen B.K. (2004). The psychological reality of phonaesthemes. Language, 80 :290–311.
Blasi D. E. , Wichmann S. , Hammarström H. , Stadler P. F. , & Christiansen M. H. (2016). Sound–meaning association biases evidenced across thousands of languages. Proceedings of the National Academy of Sciences, 113 :10818–10823.
Boersma P. & Weenink D. (2012). PRAAT: doing phonetics by computer. Accessed at http://www.praat.org/.
Brockmann M. , Drinnan M.J. , Storck C. & Carding P.N. (2011). Reliable jitter and shimmer measurements in voice clinics: the relevance of vowel, gender, vocal intensity, and fundamental frequency effects in a typical clinical task. Journal of Voice, 25 :44–53.20381308
Brockmann-Bauser M. , Bohlender J.E. & Mehta D.D. (2018). Acoustic perturbation measures improve with increasing vocal intensity in individuals with and without voice disorders. Journal of Voice, 32 :162–168.28528786
Chodroff E. & Wilson C. (2014). Burst spectrum as a cue for the stop voicing contrast in American English. Journal of the Acoustical Society of America, 136 :2762–2772.25373976
Daube C. , Ince R.A.A. & Gross J. (2019). Simple acoustic features can explain phoneme-based predictions of cortical responses to speech. Current Biology, 29 :1924–1937.31130454
Denes P.A. & Pinson E.N. (1993). The Speech Chain: The Physics and Biology of Spoken Language, 2nd edition. WH Freeman & Company: New York, NY.
de Saussure F. (1916/2009). Course in General Linguistics. Open Court Classics: Peru, IL, USA.
Fastl H. & Zwicker E. (2007). Psychoacoustics: Facts & Models. Springer-Verlag: Berlin.
Ferrand C. T. (2002). Harmonics-to-noise ratio: An index of vocal aging. Journal of Voice, 16 :480–487.12512635
Gouzoules H. (2022). When less is more in the evolution of language. Science, 377 :706–707.35951706
Hamilton L.S. , Oganian Y. & Chang E.F. (2020). Topography of speech-related acoustic and phonological feature encoding throughout the human core and parabelt auditory cortex. Preprint, bioRxiv, doi: 10.1101/2020.06.08.121624
Hill J.H. (1972). On the evolutionary foundations of language. American Anthropologist, 74 :308–317.
Hillenbrand J. , Getty L. A. , Clark M. J. , & Wheeler K. (1995). Acoustic characteristics of American English vowels. Journal of the Acoustical Society of America, 97 :3099–3111.7759650
Hirata S. , Ukita J. & Kita S. (2011). Implicit phonetic symbolism in voicing of consonants and visual lightness using Garner’s speeded classification task. Perceptual & Motor Skills, 113 :929–940.22403936
Hockett C. (1960). The origin of speech. Scientific American, 203 (3 ): 89–96.14402211
Hoffmann A.M. , Lacey S. , Matthews K.L. , Kumar V. , Sathian K. & Nygaard L.C. (2022). Acoustic parameters underlying sound-symbolic mapping of auditory pseudowords to different domains of meaning. Abstract, Cognitive Neuroscience Society, San Francisco, April 23–26, 2022.
Hollien H. , Girard G. T. , & Coleman R. F. (1977). Vocal fold vibratory patterns of pulse register phonation. Folia Phoniatrica et Logopaedica, 29 :200–205.
Hollins M. , Bensmaia S. , Karlof K. & Young F. (2000). Individual differences in perceptual space for tactile textures: evidence from multidimensional scaling. Perception & Psychophysics, 62 :1534–1544.11140177
Holm S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6 :65–70.
Hornibrook J. , Ormond T. & Maclagan M. (2018). Creaky voice or extreme vocal fry in young women. New Zealand Medical Journal, 131 :36–40.30496165
Ishi C.T. , Sakakibara K.-I. , Ishiguro H. & Hagita N. (2008). A method for automatic detection of vocal fry. IEEE Transactions on Audio, Speech, & Language Processing, 16 :47–56.
Jongman A. , Wayland R. & Wong S. (2000). Acoustic characteristics of English fricatives. Journal of the Acoustical Society of America, 108 :1252–1263.11008825
Kliper R. , Portuguese S. , & Weinshall D. (2016). Prosodic analysis of speech and the underlying mental state. In Serino S. , (eds.) Pervasive Computing Paradigms for Mental Health: MindCare 2015 Selected Papers, pp52–62. Springer, Switzerland.
Kluender K.R. , Diehl R. L. & Wright B.A. (1988). Vowel-length differences before voiced and voiceless consonants: an auditory explanation. Journal of Phonetics, 16 :153–169.
Knoeferle K. , Li J. , Maggioni E. , & Spence C. (2017). What drives sound symbolism? Different acoustic cues underlie sound-size and sound-shape mappings. Scientific Reports, 7 :5562, doi: 10.1038/s41598-017-05965-y.28717151
Köhler W. (1929). Gestalt Psychology. New York: Liveright Publishing Corporation.
Köhler W. (1947). Gestalt Psychology: An Introduction to New Concepts in Modern Psychology. Liveright: New York, NY.
Kriegeskorte N. , Mur M. , Ruff D. A. , Kiani R. , Bodurka J. , Esteky H. (2008). Matching categorical object representations in inferior temporal cortex of man and monkey. Neuron, 60 :1126–1141.19109916
Kumar G.V. , Lacey S. , Nygaard L.C. & Sathian K. (2024). Acoustic parameter combinations underlying mapping of auditory pseudoword sounds to multiple domains of meaning: a machine learning approach. Preprint bioRxiv, doi: [TBA].
Lacey S. , Jamal Y. , List S.M. , McCormick K. , Sathian K. & Nygaard L.C. (2020). Stimulus parameters underlying sound-symbolic mapping of auditory pseudowords to visual shapes. Cognitive Science, 44 :e12883. (Preprint, BioRxiv, doi: 10.1101/517581)32909637
Lacey S. , Matthews K.L. , Sathian K. & Nygaard L.C. (2024). Phonetic underpinnings of sound symbolism across multiple domains of meaning. Preprint, bioRxiv, BIORXIV/2024/610970
Ladefoged P. & Johnson K. (2011). A Course in Phonetics, 6th edition. Wadsworth: Boston MA.
Lev-Ari S. & McKay R. (2023). The sound of swearing: Are there universal patterns in profanity? Psychonomic Bulletin & Review, 30 :1103–1114.36471228
McCormick K. , Kim J. Y. , List S. , & Nygaard L. C. (2015). Sound to meaning mappings in the bouba-kiki effect. In Noelle D. C. , Dale R. , Warlaumont A. S. , Yoshimi J. , Matlock T. , Jennings C. D. , & Maglio P. P. (Eds.), Proceedings 37th Annual Meeting Cognitive Science Society (pp. 1565–1570). Austin TX, USA: Cognitive Science Society.
Mesgarani N. , Cheung C. , Johnson K. & Chang E.F. (2014). Phonetic feature encoding in human superior temporal gyrus. Science, 343 :1006–1010.24482117
Miller J. L. & Volaitis L. E. (1989). Effect of speaking rate on the perceptual structure of a phonetic category. Perception & Psychophysics, 46 :505–512.2587179
Newman S.S. (1933). Further experiments in phonetic symbolism. American Journal of Psychology, 45 :53–75.
Nishimura T. , Tokuda I.T. , Miyachi S. , Dunn J.C. , Herbst C.T. (2022). Evolutionary loss of complexity in human vocal anatomy as an adaptation for speech. Science, 377 :760–763.35951711
Nuckolls J.B. (1999). The case for sound symbolism. Annual Review of Anthropology, 28 :225–252.
Nygaard L.C. , Lacey S. , Hoffmann A.M. , Matthews K.L. & Sathian K. (2022). Voice parameters underlying sound-symbolic mapping of auditory pseudowords to different domains of meaning. Abstract, Cognitive Neuroscience Society, San Francisco, April 23–26, 2022.
Oganian Y. , Bhaya-Grossman I. , Johnson K. & Chang E.F. (2023). Vowel and formant representation in the human auditory speech cortex. Neuron, 111 :2105–2118.37105171
Parise C.V. , & Pavani F. (2011). Evidence of sound symbolism in simple vocalizations. Experimental Brain Research, 214 :373–80.21901453
Parise C.V. , & Spence C. (2012). Audiovisual crossmodal correspondences and sound symbolism: a study using the implicit association test. Experimental Brain Research, 220 :319–333.22706551
Peer E. , Brandimarte L. , Samat S. & Acquisti A. (2017). Beyond the Turk: Alternative platforms for crowdsourcing behavioral research. Journal of Experimental Social Psychology, 70 :153–163.
Ramachandran V.S. & Hubbard E.M. (2001). Synaesthesia – a window into perception, thought and language. Journal of Consciousness Studies, 8 :3–34.
Reetz H. & Jongman A. (2020). Phonetics: Transcription, Production, Acoustics, and Perception, 2nd edition. Hoboken, NJ: Wiley.
Sapir E. (1929). A study in phonetic symbolism. Journal of Experimental Psychology, 12 :225–239.
Sidhu D.M. & Pexman P.M. (2018). Five mechanisms of sound symbolic association. Psychonomic Bulletin & Review, 25 :1619–1643.28840520
Sidhu D.M. , Vigliocco G. & Pexman P.M. (2022). Higher order factors of sound symbolism. Journal of Memory & Language, 125 :104323.
Spence C. (2011). Crossmodal correspondences: A tutorial review. Attention, Perception, and Psychophysics, 73 :971–995.
Sučević J. , Savić A.M. , Popović M.B. , Styles S.J. & Ković V. (2015). Balloons and bavoons versus spikes and shikes: ERPs reveal shared neural processes for shape-sound-meaning congruence in words, and shape-sound congruence in pseudowords. Brain & Language, 145/146 :11–22.25935826
Svantesson J.-O. (2017). Sound symbolism: the role of word sound in meaning. WIREs Cognitive Sci. 8 :e1441, doi:10.1002/wcs.1441
Swadesh M. (1971). The Origin and Diversification of Language. Routledge & Kegan Paul: London.
Teixeira J. P. , & Fernandes P. O. (2014). Jitter, shimmer and HNR classification within gender, tones and vowels in healthy voices. Procedia Technology, 16 :1228–1237.
Tzeng C. Y. , Duan J. , Namy L. L. , & Nygaard L. C. (2018). Prosody in speech as a source of referential information. Language, Cognition and Neuroscience, 33 :512–526.
Uno R. , Shinohara K. , Hosokawa Y. , Atsumi N. , Kumagai G. (2020). What’s in a villain’s name? Sound symbolic values of voiced obstruents and bilabial consonants. Review of Cognitive Linguistics, 18 :428–457.
Van Puyvelde M. , Neyt X. , McGlone F. & Pattyn N. (2018). Voice stress analysis: A new framework for voice and effort in human performance. Frontiers in Psychology, 9 :1994, doi:10.3389/fpsyg.2018.01994 30515113
Villegas J. , Akita K. & Kawahara S. (2023). Psychoacoustic features explain subjective size and shape ratings of pseudo-words. Proceedings of Forum Acusticum, the 10th Convention of the European Acoustics Association, Turin, Italy, September 2023.
Walker P. & Parameswaran C.R. (2019). Cross-sensory correspondences in language: Vowel sounds can symbolize the felt heaviness of objects. Journal of Experimental Psychology: Learning, Memory, & Cognition, 45 :246–252.29698035
Warriner A.B. , Kuperman V. & Brysbaert M. (2013). Norms of valence, arousal, and dominance for 13,915 English lemmas. Behavior Research Methods, 45 :1191–1207.23404613
Whitehead R.L. , Metz D.E. & Whitehead B.H. (1984). Vibratory patterns of the vocal folds during pulse register phonation. Journal of the Acoustical Society of America, 75 :1293–1297.6725780
Wicker F.W. (1968). Mapping the intersensory regions of perceptual space. American Journal of Psychology, 81 :178–188.5747961
Woods K. J. , Siegel M. H. , Traer J. , & McDermott J. H. (2017). Headphone screening to facilitate web-based auditory experiments. Attention, Perception, & Psychophysics, 79 :2064–2072.
