
==== Front
Commun Biol
Commun Biol
Communications Biology
2399-3642
Nature Publishing Group UK London

6764
10.1038/s42003-024-06764-8
Article
Active sampling as an information seeking strategy in primate vocal interactions
Varella Thiago T. 1
Takahashi Daniel Y. 2
http://orcid.org/0000-0003-1960-7470
Ghazanfar Asif A. asifg@princeton.edu

1
1 https://ror.org/00hx57361 grid.16750.35 0000 0001 2097 5006 Princeton Neuroscience Institute & Department of Psychology, Princeton University, Princeton, NJ 08544 USA
2 https://ror.org/04wn09761 grid.411233.6 0000 0000 9687 399X Brain Institute Federal University of Rio Grande do Norte (UFRN) Av, Nascimento de Castro, 2155—Morro Branco, Natal, RN 59056-450 Brazil
7 9 2024
7 9 2024
2024
7 109815 12 2023
21 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Active sensing is a behavioral strategy for exploring the environment. In this study, we show that contact vocal behaviors can be an active sensing mechanism that uses sampling to gain information about the social environment, in particular, the vocal behavior of others. With a focus on the real-time vocal interactions of marmoset monkeys, we contrast active sampling to a vocal accommodation framework in which vocalizations are adjusted simply to maximize responses. We conduct simulations of a vocal accommodation and an active sampling policy and compare them with actual vocal interaction data. Our findings support active sampling as the best model for real-time marmoset monkey vocal exchanges. In some cases, the active sampling model was even able to partially predict the distribution of vocal durations for individuals to approximate the optimal call duration. These results suggest a non-traditional function for primate vocal interactions in which they are used by animals to seek information about their social environments.

Modeling and data show that marmoset monkeys use active sensing to acquire social information from conspecifics; this challenges traditional views of primate communication.

Subject terms

Social behaviour
Sensorimotor processing
https://doi.org/10.13039/100000065 U.S. Department of Health & Human Services | NIH | National Institute of Neurological Disorders and Stroke (NINDS) R01NS054898 Ghazanfar Asif A. https://doi.org/10.13039/100000001 National Science Foundation (NSF) DGE-2039656 Varella Thiago T. issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Contact calls are produced by a variety of animals, particularly when they are out of sight of one another. The functions of contact calls seem to be context-dependent. Among primate species, the same vocalizations are variously used during territorial encounters, mate attraction and isolation1. In all contexts, contact calls produced by one individual typically elicit a similar vocalization by another. There are two common assumptions as to the immediate purpose of these vocal exchanges. The first is that the goal of each individual is to maximize the probability of a response from the other (e.g., frogs2,3 and monkeys4,5). The second assumption is that, in some cases, vocal plasticity is used to either adjust these vocalizations acoustically to signal social closeness or to optimize signal transmission in noisy environments. This, again, has the purpose of increasing the probability of a vocal response (a phenomenon known as “vocal accommodation”6). Here, we consider another possibility: vocal exchanges are a form of active sensing used to gain information about conspecifics (that is, they are active sampling)7, information that can be used to estimate the probability of response. By active sensing, we mean the purposive use of motor control to emit a self-generated energetic signal to collect sensory information8. Active sampling, on the other hand, is defined as defined as gathering information that is important for a specific task7.

Marmoset monkeys are one species where there has been much focus on the use of their contact calls in vocal exchanges (Fig. 1). When by themselves, adult marmosets will produce contact calls and will do so approximately every 10 s as a function of an autonomic nervous system rhythm9,10. When auditory contact is made, two marmosets will communicate with multiple back-and-forth exchanges of contact calls10. When the distance between them changes or if they become visible to each other, then marmosets adjust the latency, loudness, and/or duration of their contact calls and/or switch to producing different affiliative call types11,12. These and other data demonstrate that marmoset vocal production remains flexible (or plastic) throughout adulthood13–15. Here, we test whether the real-time contact calling behavior of marmoset monkeys is consistent with active sampling. Are they interrogating their social environment using a “question-and-answer” strategy?Fig. 1 Vocal behavior within a single dyadic interaction is diverse and dynamic.

A Experimental set-up of the occluded and dyadic vocal interactions. B–E Vocal properties for individuals A and B, respectively. B Scatter plot for each individual with the time that each call was produced in the x-axis and the call duration of said call in the y-axis. The call duration is calculated as the total duration of a sequence of vocal syllables (continuous bouts of vocalization) with less than 1 s between each other10, as exemplified in the spectrogram. C Value of Bayesian Information Criterion (BIC) versus the number of components for each individual when clustering with a Gaussian mixture model the duration of all vocalizations produced after ¼ of the session. For both individuals, 3 is the optimal number of clusters. D Probability density function of each of the clusters for each individual, as determined by the Gaussian mixture model of the optimal number of clusters. E Dynamics of the time-binned coefficient of variation (CV, standard deviation divided by the mean) of the call durations for each individual, calculated via splitting the data into time windows, and calculating the CV in each window. The solid line represents a sigmoid fit of the obtained CVs. F Scatter plot of vocal durations for the population of six individuals throughout the sessions. G BIC versus the number of clusters shows that 3 is also the optimal number of clusters of call durations considering the population as a whole. H Probability density function illustrating what are the clusters. I Dynamics of the CV of call duration for the population. A solid line is a sigmoid fit.

One proposal for active sampling suggests that animals in complex environments with incomplete information use a belief-based policy16, where “belief” here means the state of an animal’s knowledge at that moment. In a belief-based system, information has value, and this value can motivate exploration under conditions in which they must consider many alternatives and where the potential rewards are not known beforehand. In the context of contact calling by marmoset monkeys who, under natural conditions, live in dense, tropical rainforests where they cannot easily see conspecifics, we hypothesize that vocalizing is a way of seeking information about conspecifics.

Results

To assess the hypothesis that marmosets are active sampling, we measured and analyzed different call durations in the single context of vocal exchanges with an out-of-sight conspecific. During such vocal exchanges, the vast majority (~80%) of vocalizations produced are multi-syllabic contact calls known as “phee calls”; remaining call types were minor variations of the phee call known as “trill-phees” and “trills” (see Methods). All calls are affiliative. We used raw recording data collected from an earlier study12; the vocalizations were from six adult marmoset monkeys, split into three pairs. The pair were put into a sound-attenuated testing room where one marmoset was placed in one corner and another in the other corner, separated by an acoustically transparent visual occluder (Fig. 1A). This context reliably elicits vocal exchanges between marmosets10,12. As in previous studies, vocal responses were defined as a vocalization from a different marmoset that occurred within 12 s of the start of the first vocalization10,12.

When we plot the duration of the contact vocalizations emitted over time for individuals (Fig. 1B) or the population (Fig. 1F), we can see that, qualitatively, most of the vocalizations with a longer duration are emitted at the beginning of the session. We quantified this using Bayesian Information Criteria and a Gaussian Mixture Model (Fig. 1C, D for the individuals and Fig. 1G, H for the population). To account for the possibility of the results being dominated by single individuals instead of general property, we ran the analysis with jackknife resampling and got an average number of clusters of 3.166. As the number of clusters needs to be an integer, it is reasonable to round it to 3 in the final model. Still, these data show that distinct groups of vocal durations are evident. Why are marmosets producing contact vocalizations with such distinct durations in the same context? Vocalizations are shaped by energetic and biomechanical constraints17. Thus, an observer might expect that, if the marmoset knows that the receiver is present after hearing them vocally respond, and they begin responding at an optimal rate, then there will come a time when there is no need to emit different call types. The idea that the marmoset may be acquiring such information throughout the session is supported by the observation that the coefficient of variation of call durations is reduced over time (Fig. 1E, I). The jackknife resampling confirmed the reduction with plausible-looking sigmoid fits and a negative slope in a linear regression with p-value < 10−3 for all resamples. Thus, we now try to provide a quantitative model for what this would look like for marmoset monkeys.

To investigate this question, we began by simulating two scenarios for how a pair of agents would change the duration of their vocalizations during one session of vocal exchanges. Figure 2 illustrates these putative processes using two different behavioral policies. In the first scenario—vocal accommodation (Fig. 2A)—the agents change their vocalizations so that they produce calls with durations that maximize the perceived likelihood of getting a response. For example, if they vocalize and get no response, they might produce a longer duration vocalization, making it more likely that the other agent hears it. If the agent successfully gets a response after a longer call, then they would produce a call with similar duration next time. This would increase the probability of response. In this scenario, the agents can learn from their previous vocalizations produced during the session. The distribution on the left of Fig. 2A, B represents a “belief” acquired prior to (or during) the session, such as over the course of development of the agent18,19. This belief informs the agent how likely it is to get a response depending on certain acoustic features. While we focused on call duration, it is possible that spectral cues are also variable and informative; the value of such cues would be dependent upon the habitat acoustics, more so than call duration. The vocalization chosen for this policy is at the peak of the belief distribution, meaning they will choose the vocalization that they think is the most likely to get a response. Notice, however, that by choosing the vocalization in the peak of the belief distribution, small changes on the choice of the vocalization (x-axis) lead to small changes in the response rate (y-axis). This is because of the flatness of the response rate around the peak. This is consistent with the notion that the agent is more “confident” about what the response rate should be in this area, and, therefore, there is a narrow learning potential.Fig. 2 Different vocalizations chosen by a policy lead to different degrees of information acquisition.

In both models, the agent starts with a belief (represented by a distribution) of what call duration leads to the highest response probability. Using a policy applied to this belief (i.e. the rules determining what action to take), the agent emits a vocalization with a specific duration. Based on the information it gets from hearing or not a response, it updates the belief using the Bayes rule. A In the vocal accommodation policy, the vocalization is chosen to maximize the immediate response probability (i.e., the duration corresponding to the peak of the belief). B Another possibility for the policy is the active sampling policy, in which the vocalization is chosen to maximize the learning (see methods). In the example in the figure, the policy leads to vocalizations on the highest slopes of the belief, rather than the peak, because they have a higher learning potential, achieving an update in the belief with higher magnitude and thus a narrower updated belief than the vocal accommodation model.

If the belief correctly describes the response likelihood, the agent is more likely to get a response when vocalizing in the peak, but if the belief is incorrect, vocalizing in the peak could lead to suboptimal calling behavior and a reduced chance to improve. A more robust strategy for information seeking would be to choose a vocalization with a duration that has more potential for learning something new. In the second scenario—the active sampling policy (Fig. 2B)—the agents change their vocal behavior so that they are more likely to learn something new. For example, even if specific vocalizations are already leading to consistent responses, the agent will try to vocalize longer or shorter as a way of learning something new that could lead to, perhaps, a higher response likelihood later on. Given the distribution of the belief, the agent will not necessarily select the vocalization in the peak (i.e., the one with the highest perceived likelihood to get a response). One example of a vocalization that could be chosen in this scenario is indicated by the dotted line in Fig. 2B. In this area, small variations in the chosen vocalization (perturbations on the x-axis of the distribution) lead to big changes in the expected response rate (green area by the y-axis). The height of the green area can serve as an intuition to the agent’s learning potential since exposure to different response rates might help the agent better represent the correct distribution. For this specific distribution, this is equivalent to the area with the highest slope. A consequence of choosing the vocalization with the highest learning potential is that the update in the belief will be larger, as observed in the updated belief distributions on the right. Here, when the agent gets a response, the updated distribution is a lot narrower, representing a big increase in confidence. Because of the symmetry of this belief distribution, there are two regions with maximal learning potential.

We conducted simulations of affiliative vocal exchanges between agent dyads (Fig. 1A). The simulations used Bayesian inference to investigate the two alternative behavioral policies described above. Specifically, we applied the Bayes rule to the belief of optimal vocalization given a newly sampled vocalization and given a variable that represents the presence or absence of a response. Both agents in the simulation use the same policy (Methods). The policies relate to two different ways of thinking about why marmosets produce the vocalizations they do during vocal exchanges with a conspecific. In both policies, the agent can be thought of as exploring the space of possible call durations and selecting a call duration that ultimately would lead to finding the vocalization that elicits a response with the highest response probability. For the vocal accommodation policy, the vocalization is chosen directly based on the vocalization with the highest response probability, while for the active sensing policy the vocalization is chosen to maximize the information acquired. The first policy considers that they are directly trying to maximize the probability of a response from their communication partner based on their prior knowledge. This policy is consistent with matching their vocalizations to their partner’s vocalizations. While marmosets do exhibit vocal accommodation like this over the course of weeks (and as a function of vocalization type)15, it is unknown to what degree they do so on shorter timescales.

In Fig. 3, each row is the result of implementing one of the two policies: vocal accommodation and active sampling. Figure 3A–C shows that the vocal accommodation policy converges mainly to one big peak along with a second, broader peak, while the active sampling policy (Fig. 3E–G) converges from a broad distribution to two main narrow peaks and a third broader peak. The two main peaks in the active sampling policy are the result of the agents trying to explore an area around the vocalization with the highest learning potential (Fig. 2B). The vocal accommodation policy leads to only one sharp peak since the vocalization converges to the perceived optimal duration. In both cases, the broader peak is related to vocalizations that are important for the initial wider exploration of the different call durations (i.e., information search). The elbow method of cluster analysis was done in Fig. 3B, F to algorithmically determine the number of clusters. Figure 3D, H show that both the vocal accommodation and the active sampling policies exhibit a transition in how variable the vocalization is throughout a session, although the transition in the active sampling policy seems sharper.Fig. 3 The active sampling model predicts diversity (3 clusters) and dynamics (sudden transition) of vocal interaction.

A Scatter plot of the simulated call durations in an interaction between two agents using a vocal accommodation policy. B Value of Bayesian information criterion (BIC) versus the number of components when Gaussian mixed models are used to cluster the call durations after ¼ of the sessions simulated by the vocal accommodation policy. The optimal number of clusters defined via the elbow method is 2. C Probability density function of each of the clusters from the vocal accommodation model, using the optimal number of clusters. D Dynamics of the time-binned coefficient of variation (CV, standard deviation divided by the mean) of the vocal accommodation simulation, calculated via splitting the data into bins in time, and calculating the CV in each bin. The solid line represents a sigmoid fit of the CVs. E Scatterplot of the simulated call durations in an interaction between two agents using the active sampling policy. F BIC vs number of components when clustering vocal durations with a Gaussian mixture model from calls taken after ¼ of the sessions in active sampling simulation. G Probability density function of each of the clusters from the active sampling model, using the optimal number of clusters given by the elbow method. H Dynamics of the time-binned CV of the active sampling simulation, along with a sigmoid fit.

Next, we compared what we observed in these simulations with the real marmoset vocal exchanges in Fig. 1. Figure 1G, H shows the same number of clusters as in Fig. 3F, G (for the active sampling policy), while in Fig. 1I we observe a sharp transition similar to the one observed in Fig. 3H for active sampling. This suggests that active sampling is the best model to concomitantly describe both the dynamics and the diversity of vocal durations when compared to the vocal accommodation policy. Both policies include the idea that the marmoset is updating their belief of the best vocalization in a given context, but the active sampling policy would mean that this update is not just passive, but that information is being actively pursued by the marmoset.

As predicted by the idea that the marmosets are actively sampling throughout a vocal interaction session, Fig. 4A shows a statistically significant increase in the proportion of calls that receive a response from their partners as the session progresses (bootstrapped linear fit had a positive slope with p = 0.042). This curve and the 95% confidence intervals were bootstrapped using different window sizes ranging from 120 to 300 s, using data from the whole population, and smoothed using cubic splines via the Python library csaps. There is a particularly sharp increase at the beginning of the curve, similarly to what we observe from the active sampling simulation. If the active sampling policy is used, there should be two clusters of call durations around the peak of an individual belief. We can plot the experience that a marmoset has with the proportion of responses heard given different call durations emitted, which is similar to what the belief should represent. The plots in Fig. 4C, E show such curves for marmosets A and B (the same individuals from Fig. 1), and they are similar to the beliefs illustrated in Fig. 2. As expected, Fig. 4B, D shows the clusters of vocal durations emitted by the same individuals, and the peaks of the clusters bracket the proportion of vocalizations that elicit a response.Fig. 4 Marmoset behavior matches predictions from the active sampling model.

A Graph showing increase in proportion of calls that get a response as the session progresses for the population of 6 individuals, along with 95% CI. B Probability density function of each of the clusters for individual A, as determined by the Gaussian mixture model of the optimal number of clusters shown in Fig. 1D. C Proportion of calls from individual A that got a response (in the x-axis) given different call durations in the y-axis. The dotted line shows alignment of the vocal clusters similarly to the curve, as predicted by the active sampling model in Fig. 2B. D Probability density function of each of the clusters for individual B, as determined by the Gaussian mixture model of the optimal number of clusters shown in Fig. 1D. E Proportion of calls from individual B that got a response (in the x-axis) given different call durations in the y-axis. Dotted line illustrates alignment of the vocal clusters with the curve. F Comparison between call duration with highest chance of response (brown spiky dot) and vocalization clusters for each marmoset (green dot for the clusters around peak response, black for the extra cluster).

A clear match between the clusters and the proportion of calls that got a response was not found in the other individuals, either because they did not exhibit clear clusters or because there was not enough data to calculate a proper curve for the proportion of responses (e.g., many window sizes for the call durations would include 100% or 0% response rate due to the small number of vocalizations). To account for the lack of data, we forced the number of clusters to be 3 for every individual marmoset, and instead of plotting a cubic spline to represent the proportion of calls that got a response, we fitted a Gaussian curve. Figure 4F shows how the vocal clusters bracket the curve for the proportion of vocalizations that elicit a response in this condition for individuals A, B, D, and F by showing the peak of the clusters and of the curve. Individuals C and E show the peak of the curve is in a higher call duration than all of their clusters. This could be explained by the fact that individual marmosets could have more than 3 clusters, despite the population exhibiting 3 clusters. With 4 clusters, we found indeed that the peak of the response curve was surrounded by the clusters. We do not have enough data to determine whether the need for a fourth cluster is evidence for a different mechanism in some individuals or just an artifact of running the analysis in a small dataset.

Discussion

It is impossible and would be highly inefficient, to build a complete representation of all the sensory information in a particular context. Thus, sampling is a necessity for organisms that can sense much more information than they can fully process. It is a means to acquire in small amounts the very rich information coming in from the different sensory systems7. Active sampling is observed across different sensory modalities, such as smelling, active vision, and active touch20–22. The idea that these senses are used actively by most animals is ubiquitous, but hearing is often left behind in these discussions23. Considerations of actively seeking information using hearing and vocal behavior are typically confined to echolocating animals. In general, our results showed that, in real-time, marmosets out of sight of one another primarily use active sampling to guide their vocal production. This suggests that contact vocalizations can be used as an active sensing mechanism similar (though not identical) to echolocation in other animals. Given our small sample size, however, some caution is warranted. While the population average was upheld following leave-one-out resampling and was in line with most individual marmoset behavioral patterns, two marmosets diverged slightly in their behavior. It was not possible to determine whether the individuals that did not follow the predictions were a consequence of a small dataset or that they were following a different strategy.

Echolocation and electrolocation are paradigmatic cases of active sensing, as stimulus energy (vocalization) is generated by the subject to detect, localize, and discriminate objects in the environment via the energy’s reflection. Sound-based active sensing by bats and other echolocating animals is a highly specialized form of vocal behavior, whereby an animal serves as both the sender and receiver of sensory energy. For example, bats emit one type of vocal signal to search for prey, and when a promising echo signal is returned, they change their vocal output to better sample what it might be24,25. Electric fish sense the surrounding environment by generating weak electric fields with a specialized electric organ and then detect distortions in these fields using electroreceptor organs26. By doing so, electric fish can detect object location and whether something is living or not (via their impedances). We propose here that marmoset contact calling is, in some ways, akin to these processes in bats and electric fish. Like echolocating in the dark or electrolocating in murky waters, contact calls are, in essence, signals generated to find something out of sight (in this case, conspecifics). However, instead of the vocal sound generating an echo reflected off prey (or distortion in an electric field), marmoset contact calls may elicit a vocal signal from a conspecific. Like echolocation and electrolocation signals used for prey capture, contact calls are used to detect, localize and discriminate something of adaptive value (for marmosets: “is anyone out there?”, “who are you?” and “where are you?”), i.e. to learn more about the social environment.

It is important to note that the active sampling policy is not mutually exclusive of vocal accommodation in a broader, long-term sense. Our paper only analyzed short-term vocal changes (within the range of 30 min). The differential with active sensing is that it allows for more flexibility in a rapidly changing environment. As it oscillates around the optimal vocalization, when such optimal vocalization changes (e.g., in the presence of another animal, or when the marmoset vocalizing changes position), the agent is already primed to probe the areas around what it believed to be optimal, instead of insisting on the same vocalization. In other words, it can lead to quicker accommodation, when the environment changes. This explains why active sensing is observed in short-term interactions even though accommodation is still observed in the span of days and months in marmoset monkeys14,15.

Marmoset vocal behavior can be characterized as a balance between energy expenditure and information gain27. Such energy-information balancing is common among other established models of active sampling. For example, the amount of energy a weakly electric fish expends in movement is a function of how much sensory information it seeks to gain via electrolocation28. Weakly electric fishes also make communication decisions that are a necessary part of the active sampling mechanism—a means to continually probe the environment29. Marmoset monkeys also have a diverse vocal repertoire whose function is not completely understood. For example, they use multiple types of affiliative vocalizations that are seemingly produced according to physical distance from conspecifics12,30,31, but sometimes produce (sub-optimally, in terms of energy usage) loud and long duration affiliative vocalizations in close contexts. We believe that, like electric fish, marmosets are using their contact vocalizations to probe the social environment. The information they are acquiring could be related to individual identity, indexical cues such as age, gender body size, and/or the location of the individual: all such information is known to be present in the contact vocalizations of marmoset monkeys32–34.

Our findings represent a departure from previous research on contact vocalizations not only for the claim that we believe they are being used for active sampling but also because we are suggesting that in real time, vocalizations are not emitted solely to maximize the perceived probability of a conspecific’s response. The assumption that a maximal response rate is a goal of the contact communication system is embedded in previous work, which includes the literature on vocal accommodation6, turn-taking (including our previous studies19), and general observations about vocal behavior3. In the simulation of the accommodation policy, vocalizations quickly converged to the optimal vocalization with respect to a higher chance of response, whereas in the active sampling simulation, vocalizations oscillate around that perceived optimal vocalization with a set of quasi-optimal vocalizations. By adopting an active sampling strategy, the individual is maximizing learning by minimizing the uncertainty about the environment, since it is choosing the vocalization that minimizes the range of their beliefs about it. This strategy makes the marmoset vocal communication more robust in complex environments and thus more adaptive. This maximization of learning through vocalizations is also consistent with how human infants communicate with caregivers. Infant vocalizations often result in turn-taking with caregivers, which provides a context for learning from those responses35, and caregivers often change their responses to infant vocalizations to facilitate learning by the infant36. As we are suggesting for marmoset vocal exchanges, these human infant-and-caregiver exchanges represent a feedback loop for learning about each other37.

Methods

Dataset

The dataset reported here was reported previously12. Recordings were obtained from six marmosets separated into three pairs of cagemates, each pair being one female and one male. The marmosets were fed with a commercial marmoset diet, insects, vegetables, and fresh fruits. Their water access was ad libitum. The Princeton University Institutional Animal Care and Use Committee approved all experiments.

A number n = 3117 of marmosets vocalizations were manually extracted and classified into different call types, even though they all serve the same function of contact calls in this context. Almost all of the vocalizations produced in these recordings were phees, trillphees, and trills12, and their proportions for each vocalization remained approximately constant throughout the sessions, with 80% phees, 10% trillphees, and 10% trills. The steadiness of the proportion of trillphees and trills was backed by a linear regression whose slope was not significantly different than 0. In some sessions, there were some twitter vocalizations produced in the first 10 min of recording, but this was not significant enough to change the average call durations. Twitter calls are shorter than phee calls (one-sided t-test p-value < 10−6, t-statistic 7.76 (95% CI is 7.49, 8.03), Cohen’s d 0.99, df = 2526, n = 6 marmosets). Since there’s a slight increase in phees and slight decrease in twitters, and phees are significantly longer than twitters, we would expect an increase in call duration. In the study, we observed a significant decrease in call duration, confirmed by the jackknife analysis (see methods, section Jackknife resampling). Therefore, a change in call types cannot explain the changes reported in the paper.

Overview of the dyadic simulation algorithm

In this paper, we tested three different policies, which are three different ways to determine how to choose the duration of a vocalization given their previous knowledge. For this, we simulated two identical agents interacting with the same policy according to the following pseudocode:

FOR each time-iteration in the simulation:

IF agent 1 will vocalize:

  SET probability distribution of the call duration given by policy and current belief

  SET duration of vocalization based on the distribution

  STORE duration of vocalization

  SET whether agent 2 will respond

  UPDATE belief of agent 1

END IF

IF agent 2 will vocalize:

  SET probability distribution of the call duration given by policy and current belief

  SET duration of vocalization based on the distribution

  STORE duration of vocalization

  SET whether agent 1 will respond

  UPDATE belief of agent 2

END IF

END FOR

The details about the probability distribution are given in section “Active sensing and the Expected Information Density (EID)” of the methods for the active sensing policy, and section “Decision making for vocal accommodation” for the vocal accommodation policy. The determination of whether the other agent will vocalize is explained in section “Sensing and observation model” of the methods. The update of the previous knowledge is explained in section “Dynamics of coefficient of variation of call duration”

Sensing and observation model

We adapted the sensing model from Chen et al.29. In our model, we assume that an agent is sensing the likelihood V to obtain a response from the other animal. After each vocalization, the marmoset receives a measurement (whether there was a response or not), which is used to inform a perceived response likelihood V to that vocalization.

To model V, let’s first define a function ϒ(θ′,x) that, given a sampled call duration x and the optimal call duration θ′, returns the assumed ground truth probability of getting a response to that vocalization. Considering that the measured response from the agent includes an error, we can define V=ϒθ′,x+ϵ, where ϵ is a zero-mean Gaussian measurement noise. The function ϒ(θ′,x) can be assumed as the following Gaussian:ϒθ′,x=exp−x−θ′22σ12Vmax

so that the minimum response probability is 0 when x is very different than θ′, and the response probability is maximum at x=θ′, which is defined as the parameter Vmax, given that no vocalization elicits a response 100% of the time. The parameter σ1 in the function is how broad that function is without changing the peak, which means that a big σ1 implies more values of x can be associated with a high response probability. In our simulations, we used σ1=0.9,Vmax=0.5. For our simulations, without loss of generality, the range of possible vocal durations is from 0 to 5, and the optimal duration is 1.5 (the values were chosen to be similar to real life values in seconds, even though the simulations did not assume a unit of measurement, nor is it intended to fit the exact values observed in real life). Note that the likelihood does not need to sum up to one as it is not a probability distribution.

Active sensing and the expected information density (EID)

Consider that the marmoset has a belief about what that optimal vocal duration is, which is represented here by a probability distribution. If the marmoset only samples the duration where the belief is highest, small variations in the sampled vocalization will not lead to big differences in the measured response probability, because in the peak the belief is locally flat. If the vocalization sampled is in a region with higher slope, any small change on the sampled vocalization might lead to differences in the received information and less obvious results, which can imply more learning. Following Chen et al.29, we model the active sensing as a policy that maximizes the learning.

The belief, represented by pt(θ), is a probability distribution that informs the probability that each possible value θ is equal to the optimal duration θ′, at a time t. It starts as a uniform distribution, assuming that the agent does not have any previous information about the physical and social environment. In real life, the marmoset generally will have some expectation about how to act where they are.We measure the uncertainty associated with a given belief pt using the associated Shannon entropy, which is calculated via:Spt=−∑θptθlogptθ.

The lower the entropy is, the more certainty the agent has in the belief. We can describe the change in entropy as ΔSx,V=Spt−Spt+1θ∣V,x, where pt+1θ∣V,x is the new belief of each possible value of θ after a given sampled x and the measured response likelihood V. This way, the best vocalization to maximize learning is the one that most lowers the expected entropy after the update, i.e., it lowersEΔSx,V=ESpt−Spt+1θ∣V,x=Spt−SEpt+1θ∣V,x

where Ept+1θ∣V,x is the expected updated belief.

For each potential vocalization x, we iterate over possible values of V to calculate the expected updated belief given a possible x sampled and a hypothetical response likelihood V.

For a given x and V, pt+1θ∣V,x is calculated applying a Bayesian filter to update pt(θ). The formula used ispt+1θ∣V,x=pV∣θ,xp(θ∣x)pV∣x=pV∣θ,xptθη.

Here, η is just a normalization constant since the updated belief is still a probability distribution. We use that pθ,∣,x=ptθ, since θ does not depend on x. Finally, the function pV∣θ,x is a likelihood function that informs how likely the marmoset is to actually associate the response probability V with the current belief of θ. Given that V=ϒθ,x+ϵ, we can define p(V∣θ,x) as a Gaussian likelihood centered in ϒ(θ,x) with a deviation σ2, sopV∣θ,x=12πσ22exp−V−ϒθ,x22σ22.

This allows us to calculate the entropy reduction ΔS(x,V) for each x and V. After calculating pt+1θ∣V,x for each possible V, we calculate Ept+1θ∣V,x by marginalizing over V. By integrating over possible values of V,we get a value for how much we expect each sampled vocalization to change the entropy, that is, ΔS(x), which we define as the EID(x). To do this, we first need a probability of V given x,pV∣x, which we marginalize over θ through the formulapV∣x=∫θpθpV∣θ,xdθ

and we then use it as the weight of ΔS(x,V) in the integral that marginalizes over V in the interval from 0 to 1. In summary:EIDx=EΔSx,V=∫VpV,∣,xΔSx,VdV.

Only in the active sensing policy we calculate the EID(x). We do so for every x within the interval of 0–5, and we can define the result of this process as the distribution EID(0, 5). To make this calculation in our simulations, we used σ2 = 0.3.

The call duration is then sampled from the expected information density normalized to a probability distribution, i.e., x~EID(0, 5). That way, the vocalization is chosen from the interval [0, 5] and is the one more likely to maximize the expected entropy reduction in the belief, i.e., the amount of new information that can be acquired from vocalizing within that interval.

Decision-making for vocal accommodation

In all the policies, the call duration chosen for the next vocalization is sampled from a probability distribution. For the vocal accommodation policy, the call duration is sampled from the belief itself, i.e., x~pt(θ). That way, vocalizations are chosen to maximize the expected response rate.

Bayesian update of the belief

After the duration of the call is decided (x), we will observe whether the other agent produces a vocalization or not. For the active sensing and the vocal accommodation policies, the agent will update the belief depending on the response. We modeled the update using a Bayesian filter.

Let’s say that r ∈ Error! Bookmark not defined. is the variable is the variable that represents the response, r = 1 if there is a response. To update the belief given the response, we can use the following Bayesian filter instead:pt+1θ=ptθ∣r=1,x=pr=1∣θ,xp(θ∣x)p(r=1∣x)=pr=1∣θ,xptθη.

Notice, however, that p(r=1∣θ,x) (the chance of receiving a response given θ and x) is directly related to the observational model ϒθ,x, so with η as a normalizing factor, we would havept+1(θ)=ηϒθ,xptθ.

If there was no response, then:ptθ∣r=0,x=pr=0∣θ,xptθη=η1−ϒθ,xpt(θ).

Gaussian Mixture Model used for clustering

We implemented a Gaussian Mixture Model to investigate the number of clusters present in the recording sessions in this study. Given a set of vocal durations extracted by filtering for the second half of the recordings, we used the function “mixture” from the Python library “sklearn” to calculate the Bayesian information criterion by considering the number of clusters any number from 1 to 5. Then, we used an elbow method to decide on the optimal number of clusters by looking at the sharper angle formed in the plot of BIC against the number of clusters. The exact implementation can be seen in the GitHub repository.

Dynamics of coefficient of variation of call duration

We computed the dynamics of the coefficient of variation (CV) over time. To do so, we started by defining a window size, at which the coefficient of variation (standard deviation over mean) would be calculated, and an overlap, to ensure continuity over the different windows/bins. For the population, the window size was 30 s with an overlap of 10 s. All of the windows included more than 30 vocalizations, with most of the bins including more than 50 vocalizations. For the individuals, the window size was 60 s with an overlap of 30 s. The window size was bigger to make sure enough vocalizations were included in every bin. For the first individual shown, there were also more than 30 vocalizations in every bin. For the second individual, some bins included as little as five vocalizations, with an average of more than eight vocalizations per bin.

The data was shown along with a sigmoid fit just for ease of visualization. The r2 of the fits are the following: 0.123 for individual A, 0.596 for individual B, 0.630 for the population, 0.920 for the accommodation model, and 0.930 for the active sensing model. These goodness-of-fit values are not actually used to assess how well the sigmoid fits since this was not informative for comparing the vocal accommodation with the active sensing model.

Jackknife resampling

To address the generalizability of the model beyond individual marmosets, we resampled the data six times, excluding one individual each time and re-running the analysis. The clustering analysis amounted to three clusters for every resample except when individual C was left out. This resulted in an average of 3.166 clusters. In addition to the sigmoid fit, we also ran a linear regression to statistically test for the reduction in the coefficient of variation. The r2 of the sigmoid, the slope of the regression, and the p-value of the two-sided t-test of the coefficient of the regression are shown for the analysis, leaving out individuals A through F, respectively (Table 1). Each resampling used an n = 5 individuals.Table 1 Statistics of the population model run via jackknife resampling

Individual	r2 of sigmoid fit	Slope of regression	p-Value of regression	
A	0.520	−1.3*10−4	<10−3	
B	0.538	−1.2*10−4	<10−3	
C	0.527	−1.1*10−4	<10−3	
D	0.586	−1.3*10−4	<10−3	
E	0.633	−1.5*10−4	<10−3	
F	0.577	−1.3*10−4	<10−3	

Statistics and reproducibility

In this study, the data was collected from six individuals who followed the protocols from the section “Dataset” of the methods. The data is available at the GitHub repository reported in the section “Data Availability”, with the analysis available in the same repository. All confidence intervals are calculated via bootstrapping and reported using the 2.5 and 97.5 percentile. Statistical comparisons were made with t-tests, with a significance threshold of 0.05, with degrees of freedom and p-values reported. All analyses were performed using Python 3.10.0, mostly using the numpy and matplotlib libraries, although the figures were assembled and labeled using Inkscape.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

Reporting Summary

Supplementary information

The online version contains supplementary material available at 10.1038/s42003-024-06764-8.

Acknowledgements

We are very grateful to Gabriel Vercelli for their help with the Bayesian modeling and to Steven Elmlinger for comments and suggestions on an earlier draft of this manuscript. This work was supported by an NSF Graduate Fellowship DGE-2039656 (T.T.V.) and NIH-NINDS R01NS054898 (A.A.G.).

Author contributions

Conceptualization: T.V., D.Y.T., A.A.G.; Methodology: T.V., D.Y.T.; Investigation T.V.; Visualization: T.V.; Funding acquisition: T.V., A.A.G.; Supervision: D.Y.T., A.A.G.; Writing—original draft: T.V., A.A.G.; Writing—review & editing: T.V., D.Y.T., A.A.G.

Peer review

Peer review information

Communications Biology thanks Francisco García-Rosales and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Primary Handling Editors: Julio Hechavarria and Joao Valente.

Data availability

Data used to generate each of the figures is available at this link: https://github.com/ThiagoTVarella/Active-Sampling-in-Primate-Vocal-Interactions/tree/main.

Code availability

Analysis code to generate each of the figures, along with up-to-date information of the algorithm, are available at this link: https://github.com/ThiagoTVarella/Active-Sampling-in-Primate-Vocal-Interactions/tree/main. The latest version at the time of publication is also archived on Zenodo38.

Competing interests

The authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Daniel Y. Takahashi, Asif A. Ghazanfar.
==== Refs
References

1. Miller, C. T., Ghazanfar, A. A. Meaningful Acoustic Units in Nonhuman Primate Vocal Behavior. in The Cognitive Animal: Empirical and Theoretical Perspectives on Animal Cognition (eds. Bekoff, M., Allen, C. & Burghardt, G. M.) 265–274 (The MIT Press, 2002).
2. Capranica, R. Evoked Vocal Response of the Bullfrog. (The MIT Press, 1965).
3. Penna, M. Selectivity of evoked vocal responses in the time domain by frogs of the genus Batrachyla. J. Herpetol. 31, 202–217 (1997).
4. Ghazanfar AA Smith-Rohrberg D Pollen AA Hauser MD Temporal cues in the antiphonal long-calling behaviour of cottontop tamarins Anim. Behav. 2002 64 427 438 10.1006/anbe.2002.3074
Ghazanfar, A. A., Smith-Rohrberg, D., Pollen, A. A. & Hauser, M. D. Temporal cues in the antiphonal long-calling behaviour of cottontop tamarins. Anim. Behav. 64, 427–438 (2002).10.1006/anbe.2002.3074
5. Miller CT Wang X Sensory-motor interactions modulate a primate vocal behavior: antiphonal calling in common marmosets J. Comp. Physiol. A 2006 192 27 38 10.1007/s00359-005-0043-z
Miller, C. T. & Wang, X. Sensory-motor interactions modulate a primate vocal behavior: antiphonal calling in common marmosets. J. Comp. Physiol. A 192, 27–38 (2006).10.1007/s00359-005-0043-z
6. Ruch H Zürcher Y Burkart JM The function and mechanism of vocal accommodation in humans and other primates Biol. Rev. 2018 93 996 1013 10.1111/brv.12382 29111610
Ruch, H., Zürcher, Y. & Burkart, J. M. The function and mechanism of vocal accommodation in humans and other primates. Biol. Rev. 93, 996–1013 (2018).29111610 10.1111/brv.12382
7. Gottlieb J Oudeyer P-Y Towards a neuroscience of active sampling and curiosity Nat. Rev. Neurosci. 2018 19 758 770 10.1038/s41583-018-0078-0 30397322
Gottlieb, J. & Oudeyer, P.-Y. Towards a neuroscience of active sampling and curiosity. Nat. Rev. Neurosci. 19, 758–770 (2018).30397322 10.1038/s41583-018-0078-0
8. Zweifel NO Hartmann MJZ Defining “active sensing” through an analysis of sensing energetics: homeoactive and alloactive sensing J. Neurophysiol. 2020 124 40 48 10.1152/jn.00608.2019 32432502
Zweifel, N. O. & Hartmann, M. J. Z. Defining “active sensing” through an analysis of sensing energetics: homeoactive and alloactive sensing. J. Neurophysiol. 124, 40–48 (2020).32432502 10.1152/jn.00608.2019
9. Borjon JI Takahashi DY Cervantes DC Ghazanfar AA Arousal dynamics drive vocal production in marmoset monkeys J. Neurophysiol. 2016 116 753 764 10.1152/jn.00136.2016 27250909
Borjon, J. I., Takahashi, D. Y., Cervantes, D. C. & Ghazanfar, A. A. Arousal dynamics drive vocal production in marmoset monkeys. J. Neurophysiol. 116, 753–764 (2016).27250909 10.1152/jn.00136.2016
10. Takahashi DY Narayanan DZ Ghazanfar AA Coupled oscillator dynamics of vocal turn-taking in monkeys Curr. Biol. 2013 23 2162 2168 10.1016/j.cub.2013.09.005 24139740
Takahashi, D. Y., Narayanan, D. Z. & Ghazanfar, A. A. Coupled oscillator dynamics of vocal turn-taking in monkeys. Curr. Biol. 23, 2162–2168 (2013).24139740 10.1016/j.cub.2013.09.005
11. Choi JY Takahashi DY Ghazanfar AA Cooperative vocal control in marmoset monkeys via vocal feedback J. Neurophysiol. 2015 114 274 283 10.1152/jn.00228.2015 25925323
Choi, J. Y., Takahashi, D. Y. & Ghazanfar, A. A. Cooperative vocal control in marmoset monkeys via vocal feedback. J. Neurophysiol. 114, 274–283 (2015).25925323 10.1152/jn.00228.2015
12. Liao DA Zhang YS Cai LX Ghazanfar AA Internal states and extrinsic factors both determine monkey vocal production Proc. Natl Acad. Sci. USA 2018 115 3978 3983 10.1073/pnas.1722426115 29581269
Liao, D. A., Zhang, Y. S., Cai, L. X. & Ghazanfar, A. A. Internal states and extrinsic factors both determine monkey vocal production. Proc. Natl Acad. Sci. USA 115, 3978–3983 (2018).29581269 10.1073/pnas.1722426115
13. Pomberger T Risueno-Segovia C Löschner J Hage SR Precise motor control enables rapid flexibility in vocal behavior of marmoset monkeys Curr. Biol. 2018 28 788 794.e783 10.1016/j.cub.2018.01.070 29478857
Pomberger, T., Risueno-Segovia, C., Löschner, J. & Hage, S. R. Precise motor control enables rapid flexibility in vocal behavior of marmoset monkeys. Curr. Biol. 28, 788–794.e783 (2018).29478857 10.1016/j.cub.2018.01.070
14. Zürcher Y Willems EP Burkart JM Are dialects socially learned in marmoset monkeys? Evidence from translocation experiments PloS ONE 2019 14 e0222486 10.1371/journal.pone.0222486 31644527
Zürcher, Y., Willems, E. P. & Burkart, J. M. Are dialects socially learned in marmoset monkeys? Evidence from translocation experiments. PloS ONE 14, e0222486 (2019).31644527 10.1371/journal.pone.0222486
15. Zürcher Y Willems EP Burkart JM Trade-offs between vocal accommodation and individual recognisability in common marmoset vocalizations Sci. Rep. 2021 11 15683 10.1038/s41598-021-95101-8 34344939
Zürcher, Y., Willems, E. P. & Burkart, J. M. Trade-offs between vocal accommodation and individual recognisability in common marmoset vocalizations. Sci. Rep. 11, 15683 (2021).34344939 10.1038/s41598-021-95101-8
16. Loewenstein G Molnar A The renaissance of belief-based utility in economics Nat. Hum. Behav. 2018 2 166 167 10.1038/s41562-018-0301-z
Loewenstein, G. & Molnar, A. The renaissance of belief-based utility in economics. Nat. Hum. Behav. 2, 166–167 (2018).10.1038/s41562-018-0301-z
17. Zhang YS Ghazanfar AA A hierarchy of autonomous systems for vocal production Trends Neurosci. 2020 43 115 126 10.1016/j.tins.2019.12.006 31955902
Zhang, Y. S. & Ghazanfar, A. A. A hierarchy of autonomous systems for vocal production. Trends Neurosci. 43, 115–126 (2020).31955902 10.1016/j.tins.2019.12.006
18. Chow CP Mitchell JF Miller CT Vocal turn-taking in a non-human primate is learned during ontogeny Proc. R. Soc. B 2015 282 20150069 10.1098/rspb.2015.0069 25904663
Chow, C. P., Mitchell, J. F. & Miller, C. T. Vocal turn-taking in a non-human primate is learned during ontogeny. Proc. R. Soc. B 282, 20150069 (2015).25904663 10.1098/rspb.2015.0069
19. Takahashi DY Fenley AR Ghazanfar AA Early development of turn-taking with parents shapes vocal acoustics in infant marmoset monkeys Philos. Trans. R. Soc. B 2016 371 1 12 10.1098/rstb.2015.0370
Takahashi, D. Y., Fenley, A. R. & Ghazanfar, A. A. Early development of turn-taking with parents shapes vocal acoustics in infant marmoset monkeys. Philos. Trans. R. Soc. B 371, 1–12 (2016).10.1098/rstb.2015.0370
20. Bajcsy R Active perception Proc. IEEE 1988 76 966 1005 10.1109/5.5968
Bajcsy, R. Active perception. Proc. IEEE 76, 966–1005 (1988).10.1109/5.5968
21. Gibson JJ Observations on active touch Psychol. Rev. 1962 69 477 491 10.1037/h0046962 13947730
Gibson, J. J. Observations on active touch. Psychol. Rev. 69, 477–491 (1962).13947730 10.1037/h0046962
22. Jones TK Allen KM Moss CF Communication with self, friends and foes in active-sensing animals J. Exp. Biol. 2021 224 jeb242637 10.1242/jeb.242637 34752625
Jones, T. K., Allen, K. M. & Moss, C. F. Communication with self, friends and foes in active-sensing animals. J. Exp. Biol. 224, jeb242637 (2021).34752625 10.1242/jeb.242637
23. Ghazanfar AA Takahashi DY The evolution of speech: vision, rhythm, cooperation Trends Cogn. Sci. 2014 18 543 553 10.1016/j.tics.2014.06.004 25048821
Ghazanfar, A. A. & Takahashi, D. Y. The evolution of speech: vision, rhythm, cooperation. Trends Cogn. Sci. 18, 543–553 (2014).25048821 10.1016/j.tics.2014.06.004
24. Griffin, D. R. Listening in the dark: the acoustic orientation of bats and men. (Yale Univ. Press, 1958).
25. Simmons JA Fenton MB O’Farrell MJ Echolocation and pursuit of prey by bats Science 1979 203 16 21 10.1126/science.758674 758674
Simmons, J. A., Fenton, M. B. & O’Farrell, M. J. Echolocation and pursuit of prey by bats. Science 203, 16–21 (1979).758674 10.1126/science.758674
26. Von der Emde G Active electrolocation of objects in weakly electric fish J. Exp. Biol. 1999 202 1205 1215 10.1242/jeb.202.10.1205 10210662
Von der Emde, G. Active electrolocation of objects in weakly electric fish. J. Exp. Biol. 202, 1205–1215 (1999).10210662 10.1242/jeb.202.10.1205
27. Varella TT Zhang YS Takahashi DY Ghazanfar AA A mechanism for punctuating equilibria during mammalian vocal development PLOS Comput. Biol. 2022 18 e1010173 10.1371/journal.pcbi.1010173 35696441
Varella, T. T., Zhang, Y. S., Takahashi, D. Y. & Ghazanfar, A. A. A mechanism for punctuating equilibria during mammalian vocal development. PLOS Comput. Biol. 18, e1010173 (2022).35696441 10.1371/journal.pcbi.1010173
28. MacIver MA Patankar NA Shirgaonkar AA Energy-information trade-offs between movement and sensing PLOS Comput. Biol. 2010 6 e1000769 10.1371/journal.pcbi.1000769 20463870
MacIver, M. A., Patankar, N. A. & Shirgaonkar, A. A. Energy-information trade-offs between movement and sensing. PLOS Comput. Biol. 6, e1000769 (2010).20463870 10.1371/journal.pcbi.1000769
29. Chen C Murphey TD MacIver MA Tuning movement for sensing in an uncertain world eLife 2020 9 e52371 10.7554/eLife.52371 32959777
Chen, C., Murphey, T. D. & MacIver, M. A. Tuning movement for sensing in an uncertain world. eLife 9, e52371 (2020).32959777 10.7554/eLife.52371
30. Bezerra BM Souto A Structure and usage of the vocal repertoire of Callithrix jacchus Int J. Primatol. 2008 29 671 701 10.1007/s10764-008-9250-0
Bezerra, B. M. & Souto, A. Structure and usage of the vocal repertoire of Callithrix jacchus. Int J. Primatol. 29, 671–701 (2008).10.1007/s10764-008-9250-0
31. Landman R Close-range vocal interaction in the common marmoset (Callithrix jacchus) PLoS ONE 2020 15 e0227392 10.1371/journal.pone.0227392 32298305
Landman, R. et al. Close-range vocal interaction in the common marmoset (Callithrix jacchus). PLoS ONE 15, e0227392 (2020).32298305 10.1371/journal.pone.0227392
32. Caselli CB The role of extragroup encounters in a Neotropical, cooperative breeding primate, the common marmoset: a field playback experiment Anim. Behav. 2018 136 137 146 10.1016/j.anbehav.2017.12.009 37065636
Caselli, C. B. et al. The role of extragroup encounters in a Neotropical, cooperative breeding primate, the common marmoset: a field playback experiment. Anim. Behav. 136, 137–146 (2018).37065636 10.1016/j.anbehav.2017.12.009
33. Miller CT Mandel K Wang X The communicative content of the common marmoset phee call during antiphonal calling Am. J. Primatol. 2010 72 974 980 10.1002/ajp.20854 20549761
Miller, C. T., Mandel, K. & Wang, X. The communicative content of the common marmoset phee call during antiphonal calling. Am. J. Primatol. 72, 974–980 (2010).20549761 10.1002/ajp.20854
34. Miller CT Wren Thomas A Individual recognition during bouts of antiphonal calling in common marmosets J. Comp. Physiol. A 2012 198 337 346 10.1007/s00359-012-0712-7
Miller, C. T. & Wren Thomas, A. Individual recognition during bouts of antiphonal calling in common marmosets. J. Comp. Physiol. A 198, 337–346 (2012).10.1007/s00359-012-0712-7
35. Goldstein, M. H. & Schwade, J. A. in Oxford Handbook of Developmental Behavioral Neuroscience (eds Blumberg, M. S., Freeman, J. H. & Robinson, S. R.) 708–729 (Oxford University Press, 2010).
36. Elmlinger SL Schwade JA Goldstein MH The ecology of prelinguistic vocal learning: Parents simplify the structure of their speech in response to babbling J. Child Lang. 2019 46 998 1011 10.1017/S0305000919000291 31307565
Elmlinger, S. L., Schwade, J. A. & Goldstein, M. H. The ecology of prelinguistic vocal learning: Parents simplify the structure of their speech in response to babbling. J. Child Lang. 46, 998–1011 (2019).31307565 10.1017/S0305000919000291
37. Warlaumont AS Richards JA Gilkerson J Oller DK A social feedback loop for speech development and its reduction in autism Psychol. Sci. 2014 25 1314 1324 10.1177/0956797614531023 24840717
Warlaumont, A. S., Richards, J. A., Gilkerson, J. & Oller, D. K. A social feedback loop for speech development and its reduction in autism. Psychol. Sci. 25, 1314–1324 (2014).24840717 10.1177/0956797614531023
38. Varella, T. T. ThiagoTVarella/Active-Sampling-in-Primate-Vocal-Interactions: Code for paper Active Sampling as an Information Seeking Strategy in Primate Vocal Interactions (v1.0). Zenodo. 10.5281/zenodo.13320214 (2024).
