
==== Front
Nat Commun
Nat Commun
Nature Communications
2041-1723
Nature Publishing Group UK London

39300106
52240
10.1038/s41467-024-52240-6
Article
A geometrical solution underlies general neural principle for serial ordering
http://orcid.org/0009-0008-7572-4157
Di Antonio Gabriele 123
Raglio Sofia 14
http://orcid.org/0000-0002-2356-4509
Mattia Maurizio maurizio.mattia@iss.it

1
1 https://ror.org/02hssy432 grid.416651.1 0000 0000 9120 6856 Natl. Center for Radiation Protection and Computational Physics, Istituto Superiore di Sanità, Rome, Italy
2 https://ror.org/02p77k626 grid.6530.0 0000 0001 2300 0941 PhD Program in Applied Electronics, ‘Roma Tre’ University of Rome, Rome, Italy
3 Research Center ‘Enrico Fermi’, Rome, Italy
4 grid.7841.a PhD Program in Behavioral Neuroscience, ‘Sapienza’ University of Rome, Rome, Italy
19 9 2024
19 9 2024
2024
15 82387 9 2023
29 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
A general mathematical description of how the brain sequentially encodes knowledge remains elusive. We propose a linear solution for serial learning tasks, based on the concept of mixed selectivity in high-dimensional neural state spaces. In our framework, neural representations of items in a sequence are projected along a “geometric” mental line learned through classical conditioning. The model successfully solves serial position tasks and explains behaviors observed in humans and animals during transitive inference tasks amidst noisy sensory input and stochastic neural activity. This approach extends to recurrent neural networks performing motor decision tasks, where the same geometric mental line correlates with motor plans and modulates network activity according to the symbolic distance between items. Serial ordering is thus predicted to emerge as a monotonic mapping between sensory input and behavioral output, highlighting a possible pivotal role for motor-related associative cortices in transitive inference tasks.

How the brain sequentially encodes knowledge is not fully understood. Here authors propose a geometric framework for the elusive neural principles of serial reasoning and sequence encoding. Neural representations are theorized to align along a learned mental line, solving serial position and transitive inference tasks.

Subject terms

Network models
Problem solving
Decision
Classical conditioning
Neural encoding
- EC H2020 Research and Innovation Programme, Grant 945539 (HBP SGA3) - Italian National Recovery and Resilience Plan (PNRR), M4C2, NextGenerationEU (Project IR0000011, CUP B51E22000150006, ‘EBRAINS-Italy’)issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Serial thinking is a cognitive function underpinning almost any of our daily actions. Our brain continuously encodes temporal sequences that we can remember, process, and replay, such as the subsequent steps to make a sandwich or the path from home to work. Serial reasoning underlies logical deduction, categorical and hierarchical thinking, which are crucial to many higher-level cognitive functions (episodic memory, algebraic computation, language, and many others). A still widely open challenge is to find a unique framework in which all these different abilities could be explained, and several hypotheses about how sequences are mentally encoded have been proposed1,2.

One of these hypotheses is that sequence information in serial learning tasks is represented in the brain as an ordinal knowledge of the items in the list, independently from their timing or their spatial location. The ranking of the elements of a sequence based on the assigned order implies an abstraction of their arrangement, together with the simultaneous encoding of the ordinal and item-specific information3,4. Notably, these are the building blocks of a more general framework named “compositional computation”5,6.

Task-relevant information, such as ordinal knowledge in serial thinking, is typically encoded by resorting to specific neuronal representations in the high dimensional state spaces of their activity. These representations in the brain are sparse and involve relatively wide populations of neurons7,8. The high degree of dynamical complexity expressed by such neuronal networks can be exploited to solve relatively difficult cognitive tasks by simply resorting to a linear mapping of their collective neuronal states9–12.

Within this theoretical framework, we conjecture the existence of a geometrical solution for serial learning tasks associated with the reorganization of a specific neuronal subspace. This manifold is a line where the arbitrary elements of a sequence can be ranked in any order as suited projections of their neural representations. We derive its analytical form, hypothesizing it is the so-called “mental line” capable of solving serial ordering tasks13–16. The proposed “geometric” mental line (GML) is capable of explaining all the behavioral effects observed in Humans and other animals performing an implicit serial learning task, the so-called Transitive Inference (TI) task17,18. We found that the noisy representation of the stimuli to be ordered plays an important role in learning and shaping behavioral performances. Our GML is embedded in the low-dimensional space determined by the neural states encoding the items in a sequence, and it turns out to be naturally implemented in recurrent neural networks (RNN). In this framework, the dynamical drift along the GML determining behavioral output results to be an attractive manifold where single units are strongly correlated, further reducing the dimensionality of the latent state space where the neural activity unfolds.

Results

Ranking abstract items: The geometric mental line

Consider for example M = 4 items S1 = ‘A’, S2 = ‘B’, S3 = ‘C’ and S4 = ‘D’. Arbitrarily ranking this set of items means to put them in sequences like ‘ABCD’ or ‘CADB’, eventually assigning to each Sk the chosen rank rk∈Z, with index k ∈ [1, M]. In the former example, we then have {rk}k=1M={4,3,2,1}, while in the latter, the ranks are {3, 1, 4, 2}. A typical example with M = 7 often used in cognitive neuroscience is shown in Fig. 1a. Given that, what kind of computational capabilities must a network of neurons be equipped with to solve this task? In other words, what is the machinery needed to encode the generic rank of a set of items?Fig. 1 Item representations and the geometric mental line (GML).

a An example list of M = 7 abstract items Sk here represented as letters (‘A’, ‘B’, …) with rank rk such that the k-th item is greater than Sj only if rk > rj. b A shallow network (linear Perceptron) with N neurons (orange circles) with activity xi (i ∈ [1, N]) linearly modulated by the presentation of one of the M items in (a). The k-isolated item elicits the neural state x=Sk, which in turn is read out as a projection on the GML ⟨ζ∣ returning the rank rk. Such projection is given by the inner product ⟨ζ∣x⟩=∑i=1Nζixi of the Perceptron state and the read-out synaptic weights ζi={⟨ζ∣}i determining the GML. c, d Two example sequences with different ordering of three abstract items ‘A’, ‘B’ and ‘C’ with the related GML ⟨ζ∣ given by Eq. (3). Each item representation lies on the orthogonal axes Sk. Their projections (colored points on the GML) are differently sorted depending on the GML orientation (see Supplementary Movie 1 for a three-dimensional view). e Same as (c) but considering a sensory (“directional”) noise ξn (see main text) in the item representation that now, at each trial n, (colored dotes) are mildly displaced from the corresponding Sk. As a result, projections on the GML will be distributed around the expected rank rk (colored distributions, see Supplementary Movie 2).

To address this question, we start considering that, under stationary conditions, a network of neurons can represent a stimulus (i.e., one of the items Sk) as a vector in the space of neuronal states. Referring to the rate-coding framework, this representation corresponds to the set ∣Sk⟩=∣x⟩ of firing rates xi that the N neurons (i ∈ [1, N]) have in response to such stimulation (Fig. 1b). For the sake of simplicity here ∣Sk⟩∈RN is a column vector where the i-th element {∣Sk⟩}i=Ski can be both positive or negative. As such, neural activity xi represents the change of firing rate of the i-th neuron in a shallow network like a linear Perceptron19,20.

In associative cortices, neurons respond to different stimuli with mixed selectivity7,21,22. Mixed coding implies input decorrelation enabling the flexible storage of relations between items23, and it is naturally implemented by introducing a layer of randomly connected neurons7,24. According to this, we assume that the inner representations ∣Sk⟩ are independent random vectors such that for any j, k ∈ [1, M]1 SjSk≡∑i=1NSjiSki=δjk.

Here, according to the “bra-ket” notation adopted in quantum mechanics25, ⟨Sj∣ is a row vector given by transposing ∣Sj⟩. Sj∣Sk⟩ is the inner product returning the projection of the vector ∣Sk⟩ onto the axis defined by ⟨Sj∣. This condition naturally holds in the limit of N → ∞, provided that Ski are independent random activities with zero mean and variance 1/N. In this limit, Eq. (1) tells us that the set of M states ∣Sk⟩ are an orthogonal basis (i.e., the axes) of an M-dimensional subspace living in the full N-dimensional neural state-space (Fig. 1c).

In this modeling framework, a linear readout unit (the gray circle in Fig. 1b) returning the arbitrary function f(rk) of the rank rk assigned to the k-th item,2 f(rk)=ζSk,

exists, and it is given by the following linear combination of item representations:3 ζ= ∑j=1Mf(rj)Sj.

Indeed, from this definition ⟨ζ∣Sk⟩=∑j=1Mf(rj)⟨Sj∣Sk⟩=∑j=1Mf(rj)δjk=f(rk) proving Eq. (2) for any set of assigned ranks {rk}k=1M. Note that by adding to the GML any arbitrary vector ⟨γ∣ external to the M-dimensional space (i.e., ⟨γ∣Sk⟩=0 for any k), Eq. (2) still holds leaving unchanged the model output. This means that, in principle, an infinite number of GML exist. We will come back to this point later in the text.

In the particular case f(rk) = rk, the row vector of synaptic weights ⟨ζ∣ allows to readout directly the rank of the k-th item presented to the shallow network in Fig. 1b. It results to be proportional to the projection of the network state ∣x⟩ along the line oriented as ⟨ζ∣, as shown in Fig. 1c for the example sequence ‘ABC’. Here, although ⟨ζ∣ is the vector of synaptic weights between the readout y and the units xi, we can represent it as a line in the network state-space. Indeed, Eq. (3)⟨ζ∣ is a weighted sum of the network activities ∣x⟩=∣Sk⟩ leading to the interpretation of it also as an activity pattern itself. In this geometric framework, the model output (i.e., the inner product y=⟨ζ∣x⟩=∑jζjxj) explicitly corresponds to the projection of the network state onto the line ⟨ζ∣. The geometrically composed ⟨ζ∣ is then the searched mental line where item representations are sorted, solving a serial ordering task. As such, in what follows, we call ⟨ζ∣ the geometric mental line (GML). Changing item order from ‘ABC’ to ‘BCA’ as in Fig. 1d, the GML rotates in the M = 3-dimensional space of the representations ∣Sk⟩ giving a new set of projections sorted according to the new item position in the sequence (see Supplementary Movie 1 for a clearer visualization of the mechanism).

Now we consider the fact that the neural state encoding an item is unavoidably noisy. This is because the unit activities xi in a network fluctuate due to several sources of “thermal” noise, which we will refer to as endogenous noise. Indeed, these units aim at modeling cortical cell assemblies composed of a finite number of neurons, each receiving balanced excitatory and inhibitory currents26,27. Another source of stochasticity is the sensory (“directional”) noise affecting the input received by the network when items are presented. This effect takes into account the fact that item representations are corrupted due to a change of the attentional level, for instance, or a temporary partial view of the visual stimuli. In this case, the actual input received by the network can be modeled as ∣Sk⟩+∣ξn⟩, where the column vector ∣ξn⟩ is a random perturbation occurring at the trial n, affecting, in turn, the direction of the item representation. As shown in Fig. 1e, the result of this sensory noise is that in each trial, the item lies in a different position close to its uncorrupted representation ∣Sk⟩. The resulting projections on the GML will be then distributed around the rank rk with variance σS2=∑n=1T⟨ζ∣ξn⟩2/T estimated across the T trials per item presented (see Supplementary Movie 2). Thus, even under this noisy condition, the GML in Eq. (3) solves, on average, the serial ordering task.

As we will see later, the stochasticity of both network activity and item representations is a key ingredient to explain the wide set of behavioral effects observed in serial learning tasks.

The transitive inference task: learning the GML

Transitive inference (TI) is an implicit serial reasoning task that consists of generalizing the comparison between items based on an arbitrarily assigned rank in a sequence. In particular, knowing how adjacent items relate to each other (i.e., A > B, B > C), the relationship among stimuli never compared before can be inferred (i.e., A > C). Together with Humans, plenty of animal species have been found to show such an ability: birds, rodents, fishes, and non-human primates28,29. Successfully performing a TI task thus means having encoded the order of the presented items, as shown, for instance, in some experimental works in Humans12,30–32. Is this ordering computed relying on the same GML derived in Eq. (3)?

To answer this question, we consider as workbench a standard non-verbal TI task, with a list of M = 7 visual stimuli, conventionally associated with the letters from ‘A’ to ‘G’ and ordered alphabetically as in Fig. 1a. During the learning phase of the task, only pairs of adjacent items (e.g., ‘AB’, ‘BC’, ‘CD’, …) are presented on a screen during different trials (Fig. 2a). Following the appearance of the pair, the subject is required to choose the item with the highest rank (Fig. 2b). If the item, chosen after a reaction time, is correct, a reward is received.Fig. 2 The GML solves the transitive inference task.

a The pairs of adjacent items (with their relationship) presented on a screen during the first learning phase of the task. b Example trial of transitive inference (TI) task. Once a pair (BA) is presented on the screen, after the Go signal (disappearance of the red cue) the subject chooses to move to the right touching the item with the highest rank (i.e., A). If the choice is correct, as in this case, the subject eventually receives a reward. c Shallow network, as in Fig. 1, responding to the presentation of the pair ‘BA’ with a positive readout (y = 1) instructs to touch the right item (‘A’ with the highest rank as r1 = 7 > r2 = 6). The representation Sk of the presented items (k ∈ {1, 2}) is linearly combined as input to the network unit with arbitrary  + 1 and  − 1 for the right and the left positions, respectively. Readout unit y is given by the projection of the network state onto the GML ζ. d In the presence of items corrupted by sensory noise, the state vectors ∣Sj−Sk⟩ (pointed out by a pair of light gray lines) appear as a distribution of perturbed states (dots). Once projected onto the GML (the gray line ζ), they give rise to Gaussian-distributed readouts centered around the expected signed symbolic distances (SD), which for the example trials ‘BA’, ‘DB’, and ‘AD’ shown are 1, 2 and  − 3. e Pairs of items presented during the test phase (top) and the expected distributions of readout activity y of the shallow network in (c) (bottom). The network response y is given by the projection of the GML worked out relying only on the item pairs from the learning set (a).

The shallow network in Fig. 2c implements this task by linearly combining the inner representations ∣Sk⟩ of the pair of items appearing on the screen. Coefficients of the combination are arbitrarily set to  + 1 and  − 1 to account for the right and left location of the items, respectively, leading to the network state ∣x⟩=∣Sk⟩−∣Sj⟩≡∣Sk−Sj⟩. The readout unit y=⟨ζ∣x⟩ will inform to reach the right or the left item by responding  + 1 or  − 1, respectively. We now ask whether a vector, in the M-dimensional subspace of the item representations, ⟨ζ∣=∑n=1Man⟨Sn∣ with suited real an can solve the task. The set of coefficients an should be found taking into account only the learning set of pairs, where the ranks of the presented items always differ by one: ∣rk − rj∣ = 1. Recalling now Eq. (1), the response to a generic pair of items ∣Sk⟩, ∣Sj⟩ results to be y=⟨ζ∣x⟩=∑n=1Man⟨Sn∣Sk−Sj⟩=∑n=1Manδnk−δnj=ak−aj. This implies that for any arbitrary coefficient an holds the relationship ak − aj = rk − rj ∈ { + 1, − 1} for all item pairs in the learning set (which is composed of pairs with adjacent items, i.e., with ∣SD∣ = 1). This is a linear system of M − 1 equations with M unknown variables ak having an infinite number of solutions ak = rk + φ, with φ any arbitrary real number. Hence, the shallow network solves the TI task with4 ζ=∑k=1MrkSk+φ ∑k=1MSk.

This solution is the same as the GML in Eq. (3) with f(rk) = rk + φ. Notably, this family of mental lines gives the right output choice even when the test set of item pairs (unseen during learning) is presented. Indeed, due to the orthonormality of item representations in Eq. (1), the readout results to be5 y=ζSk−Sj= ∑i=1M(ri+φ)(δik−δij)=rk−rj,

which is the so-called “symbolic distance” (SD  = rk − rj) giving for any pair {Sk, Sj} the response to move to the rightmost item if y > 0, i.e., if rk > rj. Of course, if y < 0, the item to choose is the one on the left. The learning set of pairs then allows, in principle, to find the GML solving both the TI task and the serial ordering task introduced in the previous Section “Ranking abstract items: The geometric mental line”. Here, it is important to remark that the particular map f(rk) ∝ rk is tightly related to the specific learning protocol adopted in the TI task. Indeed, if a different set of item pairs would be presented during learning or the amount of reward would depend on the serial position of the involved items, f(rk) ∝̸ rk would result in Eq. (4), as we show later in detail.

Notably, the same GML allows us to readout another relevant quantity in the TI task: the “joint rank”33JR = rk + rj, i.e., the sum of the ranks of the two presented items. Indeed, by summing the item representation, the readout unit gives6 ζSk+Sj=rk+rj+2φ,

which in the general case of φ ≠ 0 leads to projections on the GML containing a biased estimate of the joint rank.

Note that in the subspace of the item representations, pairs of items are no more aligned to the ∣Sk⟩ axis. Indeed, as shown in Fig. 2d, their representation ∣Sk−Sj⟩ bisect the plane determined by ∣Sk⟩ and ∣Sj⟩ of the two presented items. Considering now the sensory noise previously introduced, these representations are randomly distributed as a cloud centered in ∣Sk−Sj⟩. This is because in a given trial n, the actual representation of the pair will be ∣x⟩=∣Sk+ξn(R)⟩−∣Sj+ξn(L)⟩. As in Section “Ranking abstract items: The geometric mental line”, the representational noises of left (∣ξn(L)⟩) and right (∣ξn(R)⟩) items, once projected on the GML, are independent random variables with mean E[⟨ζ∣ξn(R,L)⟩]=0 and variance V[⟨ζ∣ξn(R,L)⟩]=σS2. Hence, the readout of noisy pairs is y=⟨ζ∣Sk−Sj⟩+⟨ζ∣ξn(R)−ξn(L)⟩=rk−rj+2σSΓn, where in each trial n, Γn is an independent random Gaussian variable with zero mean and unit variance. One example pair from the learning set (‘BA’ with SD = + 1) and two pairs from the test set (‘AD’ and ‘DB’ with SD = − 3 and  + 2, respectively) are shown in Fig. 2d, together with the expected Gaussian distribution of the readout symbolic distances.

The distributions of readouts for all the item pairs (i.e., from both learning and test set) are shown in Fig. 2e. In agreement with the so-called “symbolic distance effect” (SDE)34, pairs with large ∣SD∣ lead our shallow network to respond correctly with higher probability. This is because the related Gaussian distribution of the readout y is far from the origin. In contrast, the likelihood of making a mistake increases with decreasing symbolic distance, as the chance to have a readout with the opposite sign of the rank difference (i.e., (rk − rj)y < 0) widens. The arrangement of these distributions also supports another well-documented behavioral effect. In fact, first and last item distributions are less overlapped with the others compared to what happens to the distributions of central items. This was, for instance, observed in the posterior parietal cortex (PPC) and dorsomedial prefrontal cortex (dmPFC) in Humans32, where the terminal items representations, in a reduced 2-dimensional space, appeared to be more distant to the others with respect to the central ones. This leads, in principle, to higher performances for the pairs containing terminal items, namely the “serial position effect” (SPE)18,35, detailed in the following Section.

Learning the serial order of noisy items

Although a geometric mental line can be inferred by observing only the pairs of adjacent items in a sequence, and it works on average also in the presence of sensory noise, can it be learned in a shallow network of linear units? Previous modeling efforts investigated such possibility proving that this is the case in single-layer networks of formal neurons trained with a “delta rule”36,37. Delta rule learning38,39 is equivalent to classical conditioning in behavioral psychology39,40 as it reinforces the network weights minimizing the mean square error between expected and actual responses. The learned weights ⟨ζ∣ are then those resulting from a least mean squares algorithm which can be directly computed from the pseudoinverse of the network activity across trials41. From a geometrical perspective, this learning process is equivalent to a rotation of the GML, eventually leading to projections of pair representations given by Eq. (5) (see Supplementary Movie 3).

Following this approach, we computed the learned ⟨ζL∣ in our network performing the TI task in Fig. 2 with noisy pairs of adjacent items, and varying the presented pairs randomly from trial to trial (see “Methods”). As expected, the response accuracy (i.e., the fraction of correct responses, Fig. 3a) measured after learning depends on both the difficulty of the task, the level of noise, and the rank of involved items (Fig. 3b)36,37. For sufficiently high noise σS, performances increase with the symbolic distance (∣SD∣) between items, giving rise in Fig. 3c to the SDE introduced before. Similarly, U-shaped performances arise only for large enough σS displaying higher accuracies for pairs containing terminal items of the sequence Fig. 3d. This is the mentioned SPE associated with the fact that the first and last items are the ones always winning and losing, respectively, and thus they are easier to discriminate. From these results, a suitable noise level σS≃0.15N allows for the reproduction of experimental psychometric functions measured, for instance, in monkeys28,34,35,42–44.Fig. 3 Learning the mental line from the TI task.

a Accuracy is the fraction of correct responses y, i.e., those pointing to the item with the highest rank (y SD > 0). b Response accuracy after learning with varying levels of sensory noise σS for item pairs grouped by symbolic distance SD. c Symbolic distance effect (SDE) at varying σS. The accuracy in (b) is averaged across pairs grouped by SD in absolute value (i.e., by the easiness of the task ∣SD∣). d Serial position effect (SPE) obtained as in (c) but averaging across items. e Post-learning average readout activity y=⟨ζL∣x⟩ for different SDs and noise levels. Dashed line, y resulting from the theoretical GML (σS = 0). f Overlaps between theoretical ⟨ζ∣ and learned ⟨ζL∣ mental lines, i.e., the cosine of the angle ζζL^. The simulated network (Fig. 2c) has N = 1000 units. The M = 7 items have sensory noise σS/N={0.05,0.10,0.15}. Task sessions have T/(M − 1) = 25 trials per pair of adjacent items, and for each noise level σS, 20 random sessions are simulated. Weights ⟨ζL∣ for each session are computed from the pseudoinverse of the network activity across the T = 150 trials (see “Methods”). Response accuracy of the learned ⟨ζ∣ is estimated for each item pair by extracting 105 random realizations of the sensory noise. Whiskers of box plots in (f) are extreme values of the distribution of results across the 20 simulated sessions, stars are outliers, edges are the first and third quartiles, central marks are the medians and notches represent the 95% confidence interval for median differences.

Inspecting the distribution of responses y=⟨ζL∣x⟩ across trials with varying difficulty (Fig. 3e), we found that the learned mental line is still a geometric combination of the mean representation of the items. Indeed, a linear relationship between readout activity and symbolic distance is apparent, although the slope of the regression now decreases with σS. This is because the learned GML ⟨ζL∣ is no longer parallel to the theoretical one from Eq. (3) (Fig. 3f). In fact, the noisy representations ∣Sk+ξn⟩ averaged across the presented trials are not fully embedded into the M-dimensional space of the exact representations ∣Sk⟩ where the theoretical GML unfolds. Thus, the larger is the noise σS, the smaller is the overlap between these subspaces, meaning that the network state falls in large part into the orthogonal N − M-dimensional space. As a consequence, the representations of items ∣Sk+ξn⟩ only in part are embedded in the subspace occupied by the learned GML, leading to y values linearly distributed on a range that shrinks with the noise level.

Nonlinear mapping of serial order depends on task protocol

In the previous Section, we have shown that the encoding quality of the items to sort, eventually affects the shape of the GML, and hence, the projections y determining the behavior. Indeed, learning in the presence of higher sensory noise reduces the quality of the item representations leading to a GML oriented in directions external to the M-dimensional space of the items (Fig. 3e, f). The quality of such representations can, in general be influenced also by the design of the task and the cognitive resources requested to solve it. An example is the immediate serial recall (ISR) task45–47 consisting of the presentation of a sequence of items, followed by a test phase in which the subject is requested to recall such items in the same order they have been seen. ISR task is performed in this test phase by accessing the short-term memory (STM) of the items with their position in the sequence. The limited amount of cognitive/neuronal resources for STM limits the length of the sequence to reproduce48,49. According to “ordinal theories”, this can be modeled with a decreasing activation of the item representations in memory45,50,51.

Here, to implement this effect in our GML model (i.e., the shallow network in Fig. 1), we consider items as corrupted by an amount of sensory noise which increases with their rank (Fig. 4a-right). This is an alternative way to modulate the strength of the item representations in the sequence according to their position. As a result, the first items have a stronger representation (i.e., higher quality) than the last ones, being less corrupted by sensory noise, similar to what was observed in the prefrontal cortex of monkeys performing a serial movement task52, or in Humans performing a magnitude comparison task31. Sensory noise σS is linearly modulated according to the serial position k of the item: σS(k) = σS(1) + β (k − 1). In this framework, β is the rate of STM resources used per item. It thus determines the corruption degree of the item representations which increases with the length of a sequence. As in the previous Section, we then compute the learned GML ⟨ζ∣ from a random sequence of items ∣Sk+ξn⟩ with representational noise ∣ξn⟩ having standard deviation σS(k) given by the serial position k (see “Methods”). Model performances are tested by recalling multiple times the same sequence of items, as in the single-trial version of the ISR51. By increasing the corruption rate β, the readout y=⟨ζ∣Sk+ξn⟩ averaged across sequences displays a saturation (Fig. 4a-left), leading to a no longer linear dependence on the order k, and thus, on the item rank rk (i.e., y ∝ f(rk) rather than  ∝ rk). Remarkably, in our model, such nonlinearity emerges, although it operates only linear transformations of sensory representations with the GML given again by Eq. (3).Fig. 4 Nonlinear mapping of serial positions and ranks.

a Average readout of M = 6 items stored in short-term memory (STM) to perform an immediate serial-recall (ISR) task. Model network and GML learning as in Fig. 3 but receiving only one item per time. Representations of items to recall have sensory noise σS(k) = σS(1) + β (k − 1), linearly increasing with the position k in the sequence. For each corruption rate β, 100 trials per item are simulated. GML is learned, requiring the readout to give the position of the represented item (y=⟨ζ∣x⟩=k). Right, distribution of readouts per item for β = 0.0075 and 0.12 from 104 trials post-learning. b Fraction of trials wrongly assigning the order of a specific item at different corruption rates. Item order in the model is given by the serial position between 1 and M nearest to the readout y. c Fraction of tested sequences with different lengths M perfectly recalled at varying β. A perfect recall occurs when the serial positions of all the M items are correctly readout by the model with learned GML. The curves are averaged over 100 independent simulations. d Average readout in the same model network learning to perform the TI task as in Fig. 3 but with a probabilistic reward schedule. The correct choice in each trial is only randomly assigned to the item with the highest rank. The probability that the highest item is the one to choose, decreases logarithmically with its position k (see “Methods”). Such decrease is modulated by the scaling parameter σ. σ = 0 implies no bias (standard TI task, right-top), while σ = 0.35 (right-bottom) determines an increased probability of switching the correct choice. e, f Mean accuracy per symbolic distance and per item (i.e., serial position), respectively, for varying σ in the test phase. Results from 104 trials for each pair. All the panels show mean values ± SEM.

Considering now the behavioral response in the modeled ISR task as given by the nearest serial position to the readout y (recall is not multiple choices as in ref. 50, see “Methods”), we measure from the simulations the fraction of sequences reporting for each item a wrongly assigned position/rank (Fig. 4b). For a given corruption rate, the number of errors shows a non-monotonic pattern, highlighting the coexistence of the so-called “primacy” and “recency” effects45,46,53,54. For the former effect, the first item in the sequence has the highest performance. Performances then decrease with serial position k, as the item representations in the model have corruption degrees increasing with k. Differently, the recency effect occurs when the last item is better remembered than the second-last. In the model, this happens because the readouts exceeding the maximum rank always have the length of the sequence as their nearest serial position the length of the sequence. It is intriguing to note that, as β increases, error curves grow and widen their concavity, similar to what is found in ISR experiments involving verbal items with varying phonological similarity46.

The GML framework also allows the description of the “list-length effect” (Fig. 4c). Indeed, the fraction of sequences recalled without errors in ISR experiments decreases with the number of items composing the list, eventually giving rise to an inverted sigmoid47,51. The higher the corruption rate, the lower the performances in our model, shifting to left (i.e., perfect recall only for shorter lists) the psychometric function, similar to what is found in experiments where different sensory modalities (visual or auditory) are taken into account47.

Serial order in the ISR task with strength modulation of the item representation is not the only way to learn a nonlinear mapping of the position in the sequence onto the GML. Indeed, the same nonlinearity can arise in the TI task solved in our framework just by changing the reward schedule. For an unbiased reward not depending on the rank, we have shown in Fig. 3e, f that the linear map (i.e., y = f(rk) ∝ rk) naturally emerges. We now adopt a biased schedule for which, in randomly selected trials, the choice of the lower-rank item in the pair is rewarded. More specifically, if the pair {Sk, Sk±1} appears on the screen, we no longer consider correct the response given by rk − rk±1 = ± 1. Rather, on each trial, we assign a random rank r~k to each of the two items, sampled from Gaussian distributions with mean log(rk) and standard deviation σ. Here, σ determines the scaling factor of this logarithmically biased schedule. In this way the correct response is determined by the sign of the random variable given by the difference of the extracted ranks, having itself a Gaussian distribution with mean logrk−logrk±1=−log(1∓1/rk) and standard deviation 2σ (see “Methods”).

In simulations, learning in a biased TI task leads to a GML given by Eq. (3) with projections displaying a logarithmic trend (Fig. 4d). This trend becomes even more apparent with increasing scaling factor σ. Of course, if the reward value is deterministically assigned (σ = 0), the standard TI task is recovered, and the readout returns to be linear: f(rk) ∝ rk. Intriguingly, the behavioral performances of the model are only mildly affected by such a drastic rotation of the GML, leading to an almost unchanged SDE (Fig. 4e). This is not the case for the SPE (Fig. 4f), as the U-shaped pattern of the mean accuracy per item becomes more asymmetric being skewed toward higher ranks. As the rank determines the serial position in the list of items, the asymmetry in the SPE qualitatively replicates what is shown in Fig. 4b for the ISR task, further strengthening the hypothesis of a common GML-based ranking mechanism.

Serial ordering in recurrent neural networks

So far, we have solved the transitive inference task relying on a simple one-layer pool of uncoupled units whose activity is readout by a linear Perceptron with weights given by Eq. (3). However, cortical networks have recurrent synaptic connections and the firing rate xi(t) of their units is intrinsically stochastic. In this more realistic framework, can the same geometric mental line be used to compute the arbitrary rank assigned to the items of a sequence? To answer this question we consider a recurrent neural network (RNN) with units coupled via a random connectivity matrix and each receiving a stochastic input current intended to model an endogenous source of noise (see “Methods”). As for the shallow network in Fig. 2c, the RNN receives as additional input the representations of the items simultaneously presented in pairs (Fig. 5a). Even in this network model, the activity ∣x⟩ is readout as a projection onto the GML (y=⟨ζ∣x⟩).Fig. 5 TI task in a recurrent neural network (RNN).

a Configuration of an RNN composed of N units with stochastic activity xi (i ∈ [N]) and random connectivity matrix, modeling a cortical network capable of solving the TI task. The coupling strength is assumed to be weak, leading to a linearized dynamics of xi(t). The pair of items presented during a trial provides an input with strength gS given by the difference gSSR−SL between right and left item representations, respectively. The activity randomness is driven by an endogenous white noise received by the network unit as an independent input with the same fluctuation size. The activity x(t) is readout by the vector of weights ⟨ζ∣ defined by the GML in Eq. (3) rescaled by a factor 1/gS. b Examples of the fluctuating dynamics of the readout activity y(t) following Eq. (7). The decision is taken when the readout activity crosses for the first time (i.e., the reaction time) one of the two threshold values  ± θ (dashed lines). The example pairs are ‘BA’ (purple) and ‘AD’ (blue) with symbolic distance SD = + 1 and  − 3, respectively. For ‘BA’ two example trials are shown corresponding to both a correct (choose right) and a wrong (choose left) response. Top, probability densities (p.d.f.) of reaction times for the two example pairs. Shaded areas, deciles of the p.d.f. of y(t). Right, asymptotic Gaussian densities of y(t) are expected for the two example symbolic distances. c Expected accuracy α(SD) from Eq. (8) of the responses at varying symbolic distance and noise level σE. d Accuracy per item from (c) averaged across all item pairs of the test set. e Mean reaction time RT(SD) across symbolic distances and endogenous noise, setting the relaxation time scale τ to have RT∣∣SD∣ = 1 = 450 ms.

Under the hypothesis of relatively weak recurrent connections, the dynamics of the network state ∣x⟩ can be linearized, giving rise to the following stochastic differential equation for the readout unit (see “Methods”):7 τdy=−y+SDdt+σEdW,

which turns out to be a leaky integrator driven by a Gaussian white noise dW(t) with zero mean and covariance dt, an infinitesimal time step55. Here τ is the relaxation time scale of the network units and σE is the size of the endogenous noise affecting each unit of the network. Here, the distribution of the readout activity exponentially adapts to a Gaussian distribution centered around the symbolic distance SD between the presented items (Fig. 5b), thus RNN can solve the TI task relying on the same GML derived above.

In this modeling framework, the RNN readout can play the role of the “decision value” in diffusive models of perceptual decision56. Accordingly, the decision to reach the right (left) item on the screen is taken when the threshold level + θ (− θ) is crossed by y(t) for the first time (Fig. 5b)36. For such a decision process, the accuracy (i.e., the probability of crossing the correct decision threshold and thus making the right choice) and the reaction time (i.e., the time RT when ∣y(RT)∣ > θ for the first time starting from y(0) = 0) can be analytically derived for θ ≪ 1. Indeed, in this limit, Eq. (7) is well approximated by a Wiener process with two absorbing barriers55. For it, the probability of having y(t) crossing the positive threshold θ when SD = rR − rL > 0 is given by8 α(SD)=121+tanhθSDσE2.

This accuracy is an estimate of the fraction of correct trials where the item on the right is chosen as having the highest rank. Similarly to what was found for shallow networks in the presence of sensory noise in Fig. 3c, the symbolic distance effect clearly emerges (Fig. 5c). Indeed, the accuracy decreases with task difficulty (smaller ∣SD∣) and size σE of the endogenous noise. The serial position effect is also recovered (Fig. 5d), such that the terminal items in the sequence are associated with the highest accuracies. Despite such similarities, a more careful inspection allows us to find a significant difference between the case of sensory and endogenous noise. Indeed, performances on terminal items are more affected by σE rather than by σS. An increase in sensory noise σS (Fig. 3d) leads to a more pronounced U-shaped pattern in performance, while the performance of terminal items remains relatively stable. Conversely, Fig. 5d shows a general decline in performance at higher levels of endogenous noise σE, for both central and terminal items. Endogenous noise affects the network activity, while sensory noise directly influences the distributions of the item projections onto the GML. Consequently, central items are more susceptible to this effect than terminal ones. This is consistent with item representations found in Human brain activity, where central items exhibit more overlap with each other than terminal ones30,32. Increasing sensory noise amplifies the overlap of central distributions, whereas the impact on terminal items is mitigated by the expansion of their tail distributions, which aids in distinguishing their rank position.

The reaction time RT is the other behavioral output we can workout analytically from the stochastic dynamics (7) of the readout. In this theoretical framework, RT is the “first-passage time” of the readout activity through the threshold θ. It is a stochastic variable whose mean results to be559 RT(SD)=τθSDtanhθSDσE2.

As shown in Fig. 5e, the mean response time reduces with the ease of the decision to take, i.e., with the symbolic distance between the pairs of presented items. This serial position effect is another hallmark of the TI task28,36,44,57, according to the well-known speed and accuracy trade-off in decision-making56.

In the presence of sensory noise or for specific task protocols, learned GML can lead to readouts y differing from the symbolic distance SD (Figs. 3e, 4a, d). The formalism developed here remains true in this scenario as well, provided that the dependency on SD in Eqs. (8) and (9) are replaced by y. This establishes a direct link between behavior and the GML-based readout of the RNN activity. Indeed, by making use of Eqs. (8) and (9), we can derive the expression RT(y) = [2α(y) − 1]τθ/y, eventually leading to10 y=2α−1RTτθ.

As both the accuracy α and the reaction time RT are experimentally accessible, this expression allows us to infer the neuronal readout y up to an unknown constant factor τθ. Interestingly, this factor is irrelevant if we are interested in following the relative changes of y across time in experiments with continuous learning, as in the TI task.

Optimal balance between different sources of noise

As sensory and endogenous noise differently affect the behavioral performance of the RNN, we further investigate this aspect in a more realistic scenario. We expose the same linear RNN to a continuous stream of trials where the usual item representations incorporate the sensory noise introduced in Section “Learning the serial order of noisy items” (Fig. 6a). In this in silico experiment, the network state at the beginning of each trial is no more set to 0, as is the case in cortical networks. The vector of the readout weights ⟨ζ∣ is learned by resorting to the pseudoinverse of the activity matrix aiming at reproducing the correct choice in the trials with item pairs from the learning set (y = SD = ± 1, see “Methods”).Fig. 6 Behavioral effects in the TI task are reproduced in RNN with noise.

a The RNN of Fig. 5a processes a continuous stream of randomized TI task trials, incorporating both sensory (σS′) and endogenous (σE′) noise. The network is trained on item pairs from the learning set (∣SD∣ = 1), computing the readout weights ⟨ζ∣ from the pseudoinverse of network activity, assuming a  + 1 (− 1) response when the highest-rank item is on the right (left). The RNN makes a decision when the readout activity first crosses a threshold (∣y(t)∣ ≥ θ = 0.5). If y(t) does not cross this threshold within 1.5 s from item pair onset, the trial is excluded from performance computation. b Distribution of the post-learning readout activity when only one item (top panel) or a pair of items (bottom panel) is presented per time. The distributions are estimated by taking the asymptotic (last 10% of the stimulation period) network activity during the stimulus presentation for 100 random trials per item (pair). c Frequency of network realizations (gray shading) capable of reproducing all the behavioral effects expected for the TI task, by varying the size of both the sources of noise (σS′ and σE′). An RNN reproduces such an effect if it simultaneously displays (i) accuracy and a reaction time increasing with ∣SD∣ (SDEAcc and SDERT, respectively) and (ii) if the mean accuracy of terminal items is greater than the central ones (SPE). Contours represent the response rate, i.e., the fraction of trials with RNN taking a decision. d Frequency of the behavioral effects (averaged over 50 independent RNN realizations) for varying balanced levels of noise (σS′=σE′). Errors bars indicate the SEM of the values. e Behavioral effects (averaged over 50 RNNs) co-occurring in different noise regimes pointed out in (c) by colored circles: strongly endogenous (red), balanced (purple), and strongly sensory (cyan) noise. The balanced noise case is the one with the highest frequency of responses. Reaction times are normalized to 1 for ∣SD∣ = 1.

All these additional sources of uncertainty do not affect the capability of the network to encode the serial order of the items. Indeed, by stimulating the trained network with each item presented alone, the readout activity returns a separate distribution of values. According to the expectation, such values correspond to the projection of the network activity onto our GML, where the items are sorted based on their relative rank (Fig. 6b). These low-dimensional rank-ordered representations have been observed in both monkeys and Humans12,31,32,58. Consistently, projecting the representations of item pairs on the same GML yields ordered distributions of SD (Fig. 6b).

However, inspecting different combinations of noise levels, it is apparent that only a tight balance between them allows the simultaneous occurrence of the mentioned behavioral effects (dark gray band in Fig. 6c). In this particular region, when considering an ensemble of 50 randomly selected RNNs, a significant number of networks exhibit (i) an increasing accuracy α(SD) < α(SD + 1) (SDE), (ii) a decreasing reaction time with the symbolic distance RT(SD) > RT(SD + 1) (SDE), and iii) a higher average accuracy for the first and last items with respect to the central ones (SPE). This analysis is performed by changing both sensory and endogenous noise levels, here denoted by (σS′,σE′). These parameters represent the effect of noise directly on the activity of the network units, and they are tightly correlated with the size of noise fluctuations (σS, σE) measured along the GML (see “Methods”). The existence of a sweet spot for the balance between sensory and endogenous noise, is even more strikingly represented by the narrow peak in the rate of occurrence of all effects along the line σS′=σE′ (Fig. 6d). Such balance clearly emerges from the competition of a two-fold role of the noise. A destructive role occurs when it is so large to disrupt the capability of the RNN to produce the correct response, and a constructive one, as noise is a needed ingredient to have both the symbolic distance and serial position effects (SDE and SPE, respectively).

Although along the narrow band evidenced in Fig. 6c, the behavioral effects are similarly reproduced (Fig. 6e), only when both sensory and endogenous noise have a comparable size (e.g., purple circle, σS′=0.2 and σE′=0.4) the RNN has the highest rate of responses (i.e., the fraction of trials in which a decision is taken) compared to the blue and the magenta dots on the same band.

Impact of sequence length and size of the training set

Using the above RNN with balanced noise as an in silico experiment, we now investigate the impact on some relevant parameters of the TI task. This is to gain further insights and make predictions about how learning of serial ordering can be experimentally modulated.

Firstly we focus on the size of the training set (Fig. 7a). For a small number of learning trials, noise is too high to reach significant performances, limiting the capability to reliably estimate the proper vector of readout weights. This leads to indistinguishable projections of the network activity onto the GML with almost equally low performances and long reaction times (lighter curves). Increasing the number of learning trials, all the behavioral effects are recovered. Interestingly, we found the accuracy displays a change of concavity in the performances and a flattening in the SPE for mid-rank items, which can be explained by the sigmoidal dependence highlighted in Eq. (8).Fig. 7 TI task in RNN with varying number of training trials and items to sort.

a Behavioral effects: SDE for the accuracy and the reaction time and SPE from left to right, respectively. Number of learning trials range from 4 to 20. Network as in Fig. 6 with σS′=0.2, σE′=0.4. Accuracy and reaction times are averaged across 50 random realizations of the RNN. b Behavioral effects as in (a) for increasing numbers of items in the sequence to sort. Items number range from 6 to 20. The error bars represent the associated SEM.

By increasing the number of items to be sorted in the TI task, a strong modulation of the psychometric curves can also be observed (Fig. 7b). SDE displays a change of concavity both in the reaction times and in the accuracy, leading to more rapid changes for larger symbolic distances. Another interesting effect is the equalization of the middle-rank representations as soon as the lowest possible accuracy of 0.5 is approached: a barrier limiting the maximum number of items that can be sorted.

Learning the GML in RNN taking motor decisions

So far, we studied how the GML ⟨ζ∣ is learned assuming that the only plastic synapses were those between the network units xi and the readout y. Here we investigate a more realistic framework, where the cerebral network involved in a TI task is represented by an RNN encoding the motor output by confining its activity ∣x⟩ within a low-dimensional latent space59. Such choice follows from the experimental evidence that the premotor cortex (PMC) of monkeys contributes to the motor decision in a TI task with neuronal activity modulated by the symbolic distance between item pairs44.

To this purpose, in addition to the sensory input gS∣SR−SL⟩, the RNN introduced in Fig. 6 also receives the motor-related input gM∣μ⟩ (Fig. 8a) (see “Methods”). As shown later, this additional input determines the axis along which the network activity is forced to unfold as the decision to move is taken. With this, we aim at roughly modeling motor plans known to be encoded in PMC. The relationship between gS and gM modulates the relative strength of the sensory and motor input to the network. As in previous network configurations, the learning phase involves only the set of pairs {Sj, Sk} with symbolic distance ∣SD∣ = ∣rj − rk∣ = 1 (Fig. 2a), and we impose y = + 1 (− 1) when the item to reach is on the right (left). In this phase the mental line ⟨ζ∣ is computed resorting to the pseudo-inverse of the RNN activity in the absence of feedback (i.e., without re-injecting y=⟨ζ∣x⟩ as a factor of ∣μ⟩, see “Methods”).Fig. 8 RNNs with feedback learn the GML reducing the dimensionality of the latent space.

a RNN as in Fig. 5 with the additional input mediated by the weights μ with strength gM eventually modulated by the feedback (solid arrow) provided by the readout y. During the learning phase, only pairs of items with unitary symbolic distance are presented (as in Fig. 3a). The correct action (i.e., move right (left) to reach the winning item) is represented by the input +μ (−μ) (dotted arrow). b Neural trajectories of the RNN during the test phase (as in Fig. 3a). The learned GML ⟨ζ∣ results from the pseudo-inverse performed on the neural activity of the learning. Principal component analysis (PCA) is performed on the neural activity of the whole test phase. Trajectories are shown in the planar subspace determined by the projection of the network activity on ⟨ζ∣ and the first PC (bottom), and in the plane of the first and second PCs (top). Colors code for the symbolic distance between the items of the presented pair {SL, SR}. Here, the RNN has N = 100, x(0)=0, g = 0.1, gS = 0.01, and gM = 0.02 (see “Methods” for other parameters and task details). c Variance explained by the first principal component PC1, varying the feedback strength gM for an RNN with the same parameters as in (b). Inset, variance explained by first 15 PCs for the RNN with gM = 0.01. d Mean correlations 1/N∑i=1N⟨yxi⟩± SEM versus ∣SD∣ are shown for networks with varying feedback strength gM. Average over 100 simulations. Other parameters: g = 0.05 and gS = 0.01.

In simulation we tested the response of the RNN with feedback to all possible item pairs, finding that it successfully solved the task. Indeed, by performing a principal component analysis (PCA) on the network states, the neural trajectories in the latent plane (PC1, y) clearly cluster according to the symbolic distance of the presented pairs (Fig. 8b-bottom). The trajectory slopes are higher for higher SD, reflecting a change in the driving force proportional to task difficulty. Besides, the asymptotic value of the readout limt→∞y(t)=y∞ returns the symbolic distance SD = rk − rj.

For our linear RNN, this can be analytically proven as the asymptotic activity is ∣x∞⟩≡limt≫τ∣x(t)⟩=∣S~R−S~L⟩+y∞∣μ~⟩. Here ∣S~k⟩ and ∣μ~⟩ are the inner representations of the presented items and of the motor plan, respectively, transformed by the activity reverberation in the network (see “Methods”). As a result, the GML has to be11 ζ≡∑k=1M(rk+φ)S~k+γ,

holding for any real φ and row vector ⟨γ∣ orthogonal to the subspace where the item representations ⟨S~k∣ and motor plan ⟨μ~∣ live. This expression is equivalent to the GML derived in Eq. (4), confirming it is a general solution holding also for RNN with feedback. Note that the found ⟨ζ∣ does not depend on the motor plan. Thus, although both the sensory and the motor engrams coexist in the network, they do not interfere.

In Fig. 8b-top, it is interesting to note that a richer structure of neural trajectories emerges in the plane of the first two principal components (PC1, PC2). Trials with the same symbolic distance SD are now split, unfolding the information about the items in the presented pairs. This closely mirrors the curved manifold observed in monkeys lateral intraparietal (LIP) area58, in which the decision axis encodes motion direction (as in Fig. 8b-bottom) and an orthogonal axis encodes stimulus difficulty (represented by SD values in our model). This is a direct consequence of the fact that representations of item pairs occupy an M-dimensional subspace (Fig. 1d), and this information persists in the asymptotic state of the network in Eq. (17). Recurrent activity transforms these internal representations aligning the network state along the axis ∣μ~⟩, which in turn is the average ∣x∞⟩ across trials. For this reason, the stronger the feedback gM, the larger the variance explained by PC1 (Fig. 8c). Not only, as in the above expression for ∣x∞⟩ (i.e., Eq. (17) in “Methods”), the contribution of the motor component is also proportional to the symbolic distance SD = y∞, also the activity level of the network units are expected to grow according to SD. Such correlation is apparent in Fig. 8b-bottom for the example network, and we found it to be weakened by the increase of the feedback strength gM (Fig. 8d). This is due to the sigmoidal activation Φ(h) limiting the activity xi to be in the range [ − 1, + 1].

To summarize, networks with feedback are capable of encoding simultaneously both motor-related activity and the same geometric mental line found in more simplistic conditions. The dimensionality of the latent state-space of the network is compressed according to the strength gM of the learned feedback. This can limit the degree of correlation between the activity of the network units and the symbolic distance of presented pairs. Remarkably, a similar SD-dependent modulation of the motor-selective activity was found in PMC neurons of monkeys performing a TI task44.

Discussion

The brain encodes task relevant information into complex dynamical patterns of neuronal activity, requiring a high-dimensional state space to be embedded9,60. These high-dimensional representations have the great advantage of being linearly separable, as random projections of correlated representations tend to be orthogonal in the target high-dimensional neural space, according to the “efficient coding hypothesis”61,62. Linearly separable representations can be indeed categorized via relatively simple learning rules eventually leading to a dimensional reduction of the output-potent space63. Remarkably, such low-dimensional latent spaces where correlated neural activity is constrained to wander are pervasively found in associative cortices of many species, suggesting it as a successful computational strategy preserved across evolution8,64,65.

Here we widen the realm of cognitive functions by exploiting such “expansion-compression” computational paradigm to the representations of stimuli and behavioral output in neuronal networks. We indeed prove that the task of arbitrarily ranking a set of M abstract items can always be performed by looking at the projections of their inner representations onto a suited low-dimensional subspace. Centroids of these projections can be arbitrarily mapped on any monotonic function of the ranks assigned to the items. The subspace where item ranks are encoded works as a geometric mental line (GML), given by a suited linear combination of the item representations. In a network composed of N neurons, the GML is not unique as all the N − M dimensions of the residual state space are unconstrained, leading to an infinite set of solutions. As a result, other network resources can, in principle, be used simultaneously to encode additional information and/or to perform complementary computations. The existence of a GML thus implies that the dimensionality of the representational space reduces from M to one as a byproduct of the nonlinear dynamics naturally occurring into a recurrent neuronal network, compressing the information about the ranks into a scalar magnitude. The resulting one-dimensional manifold straightforwardly informs about the motor decision (output) to take, which, in turn we have shown to be modulated by task difficulty (i.e., different symbolic distances). Similarly, several monkey studies have observed that task difficulty modulates the representation of stimuli along the decision axis44,58.

Thanks to the fact that our model is based on the linear combination of representations, the GML is highly flexible9,66,67 and different directions can be read out for different arbitrary orders of the same set of items (Fig. 2). This supports the idea that in serial reasoning two different representations must be learned3: the item representations, which are stored in a suited neural subspace, and the ordinal information, which depends on the readouts shaped by learning constrained by the task to perform. This allows the GML to represent not just different ranks for the same item, but also different features ordered in separate directions. Imagine two GMLs – each oriented to read out the correct order for a specific feature. In context-dependent decision tasks (or perceptual decision making), these GMLs would be orthogonal, discriminating rankings for different features (as for two different task58, or for two different orderings in different directions12,65,68,69). Conversely, for positional inference (different items, same ranks), they would be parallel31,32.

Intriguingly, the geometrical mental line successfully faces task sets never seen during the learning phase. This capability to generalize is a hallmark of serial reasoning needed to solve the transitive inference task. The formulation of a geometrical solution of a task is the backbone of many cognitive models of serial reasoning10,67,70,71. The idea is that in cognitive processes, it is important to consider both the content and the geometrical structure of the neural activity, which fixes the relationships between the relevant variables of the task. In the specific case of transitive inference, the geometric framework underlying the task is widely believed to be a linear workspace13,72–75. The GML we derived represents this one-dimensional manifold, allowing us to solve the TI task. Taking into account the intrinsic uncertainty of item representations (sensory noise) and the intrinsic stochasticity of the neuronal activity (endogenous noise), simulated recurrent neural networks (RNNs) reproduced all the behavioral effects observed in experiments: (i) the symbolic distance effect, both in performances and reaction times, and (ii) the serial position effect. These results can be explained in terms of a varying signal-to-noise ratio for different couples of items. Adding noise to the item representations leads RNNs to visit a state subspace with a dimensionality higher than the length M of the item list. Consequently, the learned GML moves away from being entirely contained in the M-dimensional space of the items as prescribed by Eq. (3). A rotation occurs (Fig. 3e) without affecting the effectiveness of the GML in coding the serial order of the items. Such a widening of dimensionality can be countered by encoding simultaneously other task-relevant information like the motor plan of behavioral responses (Fig. 8). From this, we expect that the expansion and compression of the latent state-space visited by cortical networks probed in vivo33,35,44, is differently expressed during the execution of the TI task.

Further investigating the same RNN, some predictions arose about different task settings, e.g., changing the number of training trials or changing the number M of items in the learned sequence. This offers an interesting perspective to understand the predictive power of our theoretical framework looking at direct comparisons with behavioral data from experiments76. Focusing on the specific case of item sequences with increasing lengths, we expect to see a change in the concavity of the accuracy in the SDE (Fig. 7b), together with a flattening of the performances for central items in the SPE. This kind of model prediction can inform about the limitations of the network in the capability to store serial-order information. These limitations are usually worked around in animal experiments resorting to the so-called “linking chain paradigm”42,77. It would be then interesting to investigate this learning paradigm within our theoretical framework. Indeed, low-dimensional rank-ordered representations have been observed in artificial neural networks and in the brain activity of Humans performing a list-linking task32. Given its focus on how learning shapes item representations, this work can, in principle, offer an ideal workbench to investigate the ongoing network rewiring through the lens of our analytical framework.

Furthermore, we proved that the GML can be naturally embedded into an RNN designed to make motor decisions. In this case, a one-dimensional manifold naturally emerges in the latent space visited by network activity, which is aligned to the motor plan encoding the decision to take. Such dynamical organization of the network dynamics accounts for the expected modulation of single-unit activity observed in the premotor cortex (PMC) of monkeys: the larger the SD (higher response value in the RNN, as predicted by the theoretical GML), the larger the number of units in the RNN exceeding a given activity threshold.

Starting from this evidence we speculate that PMC might be an ideal candidate to implement the GML solution when the transitive inference task is learned and successfully performed. We can also speculate that this mental line representation might play a role well before movement onset, as it has been shown that in the delay period following the presentation of the item pairs, the decision is already taken, and the symbolic distance is no more relevant78.

Notably, we also found that in addition to the symbolic distance, the GML allows our network models to provide as a read-out the joint rank (i.e., the rank sum) of item pairs. This happens if the item representations are summed (∣SR+SL⟩) instead of being subtracted (∣SR−SL⟩). Interestingly, in the posterior parietal cortex of non-human primates, neuronal activity was found to correlate with the joint rank despite the fact that it was not relevant information to solve the TI task33. According to this evidence, a reasonable hypothesis to test is that different sensory pathways might coexist implementing the sum and the difference of item representations presented in pair. In this way, cortical networks would be spontaneously ready to respond according to both the symbolic distance and the joint rank, a kind of primitive algebraic competence.

Regarding the linear combination of item representations determining the GML ⟨ζ∣ in Eq. (3), it is important to remark that the coefficients (i.e., the projection values) of such combination are given by the item ranks rk. This is due to the fact that the TI task is learned on trials with item pairs with ∣SD∣ = 1. However, depending on the training procedure adopted in the serial learning task, the response can be shaped as a function f(rk) of the item ranks, as emphasized in Eq. (3). As shown in Fig. 4, this is the case of the biased TI task we introduced, having reward probabilistically assigned as a function of the mean rank of the presented pairs. Note that, in the experimental literature, there are several other possible modifications of the reward schedule that we expect can change the projection values embedded into the GML, without affecting the capability to encode the correct order in the sequence (28,43,79). Alternatively, also the number of trials composing the learning phase can have a similar impact, as it affects the quality of item representations experienced (Fig. 7a). We explicitly tested such a possibility in modeling the ISR task in which all items have equally corrupted representations (β = 0) but are not always present in the list to recall. More specifically, in some learning trials, items are randomly removed from the list, such that each of them is presented a fraction of times decreasing with their serial position (Supplementary Fig. 1). In this way, items seen fewer times will be represented more weakly in the GML. The larger the difference between the number of trials per item, the stronger the nonlinearity of the rank-readout mapping. Starting from this evidence, we can imagine a scaling of the number of presentations leading to the logarithmic function f(rk)=log(rk), as in the case of an over-training of the low-rank items as hypothesized in the ordering of numerical quantities37,80–83. In these conditions, the activity projection onto ⟨ζ∣ (i.e., the network readout y) is no longer equally spaced: the first items have a greater distance along the GML. This might explain some behavioral effects, such as those following Weber’s principle formalized as Fechner’s law84,85, giving rise to the well-known primacy effect possibly observed in various serial learning tasks (e.g., in the simultaneous chain task16,86). Moreover, nonlinear projections on the GML may also arise from decision variables distributed in the neural state space on curved manifolds. In this case, projections may be, in principle, no more equally spaced according to what recently observed in both Humans and monkeys32,58.

Finally, we remark that several models have been proposed to explain how a mental scheme of the item ranks could be learned by humans and animals32,36,37,87,88, and what is the role of the reward in this process76,89. An intriguing and open question is, then, whether such theoretical frameworks, together with the one we presented here, are capable of predicting effects not yet observed, sharpening their range of applicability, and thus shedding further light on the brain mechanisms underlying serial thinking.

Methods

RNN dynamics and its linearization

Recurrent neural networks (RNN) are composed of N units (N = 100, unless otherwise specified in the main text). The j-th unit has activity state xj(t)={∣x(t)⟩}j evolving according to the first-order dynamicsτxj°=Φ(hj)−xj

for any j∈[1,N]⊂Z. The decay time constant τ = 0.1s, and the activation function Φ(hj)≡tanh(hj) is the same for all units. These dynamical systems serve as effective models of cortical networks, as they can be viewed as networks of local cell assemblies, which are the units composing the RNN. Under mean-field approximation, each of these assemblies, composed of excitatory (glutamatergic) and inhibitory (GABAergic) neurons, can be described by similar low-dimensional dynamics of the firing rate x(t)27,90–93. The synaptic input hj={∣h⟩}j is the weighted sumhj(t)= ∑k=1NWjkxk(t)+hj,ext(t)

where Wjk is the element of the synaptic matrix W∈RN×N sampled randomly from a Gaussian distribution with 0 mean and standard deviation g/N. hj,ext(t) is the external input received by the unit j. The network dynamics is integrated relying on the first-order Euler-Maruyama method with time step dt = 0.1τ to take into account the possible sources of noise in the input detailed in the main text and in the following. In matrix form, the above dynamics can be compactly written as12 τ∣x°⟩=∣Φ(h)⟩−∣x⟩

with∣h⟩=W∣x⟩+∣hext⟩,

where the “kets” ∣⋅⟩∈RN×1 are column vectors.

During the serial ranking tasks discussed in the main text, the presentation of items on the screen is modeled as a modulation of the external input. More specifically, the single k-th item elicits a sensory input ∣hext⟩=gS∣Sk⟩ with strength gS set to 1 when not specified otherwise. The vector elements Ski={∣Sk⟩}i are i.i.d. random variables with zero mean and variance 1/N. The “bra” is the row vectors transposed of ∣Sk⟩, such that in the large-network limit, the inner product is13 SjSk≡∑n=1NSjnSkn→N→∞δjk.

When two items are simultaneously presented as in the TI task, the input received by the network is a linear combination of the isolated stimuli: ∣hext⟩=gS(∣Sj⟩−∣Sk⟩)≡gS∣Sj−Sk⟩. For the sake of simplicity, the item k presented on the left side of the screen has negative strength, while the j-th located on the right contributes as it would be alone. Such a peculiar modeling choice is not critical and can be generalized as shown below (Subsection ‘GML with arbitrary item representations’).

When both the spectral radius g of W and the strength gS of the sensory input are small (g, gS ≪ 1), the dynamics of the RNN can be linearized. Indeed, under these conditions assuming a sensory and endogenous noise of the same order of magnitude (see below), the synaptic input hj is expected to be of the same order. Hence, the activation function can be approximated as Φ(hj)≃Φ(0)+hjΦ′(0)=hj, and Eq. (12) reduces to14 τ∣x°⟩=∣h⟩−∣x⟩=−(I−W)∣x⟩+∣hext⟩.

As above, at rest ∣hext⟩=∣0⟩, while when the pair of items j and k appears to the right and to the left of the screen, respectively, ∣hext⟩=gS∣SR−SL⟩ with SR = Sj and SL = Sk. All the RNNs studied in this work operate within the linear regime and as such follow the dynamics (14).

Linear RNN and motor decision

In the linear RNN introduced in Section “Serial ordering in recurrent neural networks”, we impose that the response unfolds along an independent 1-dimensional manifold encoding the motor plan. To this purpose, an additional motor-related input gM∣μ⟩ is added to ∣hext⟩ further modulate by the expected response y′(t) leading to the linear dynamics15 τ∣x°⟩=−(I−W)∣x⟩+gS∣SR−SL⟩+y′gM∣μ⟩.

Here gS and gM modulate the strength of the sensory and motor-related inputs, respectively. As the item representations ∣Sk⟩, the motor axis ∣μ⟩ is extracted randomly making all these engrams orthogonal (i.e., ⟨Sk∣μ⟩=0 for any item k) in the limit of infinitely large networks (N → ∞).

In the learning phase involving only the set of pairs {Sj, Sk} with unitary symbolic distance, i.e., ∣SD∣ = ∣rj − rk∣ = 1 (see Fig. 2a), the expected response is y′=+1 (− 1) when the item to reach is on the right (left). As in the standard reservoir computing framework94, if the GML ⟨ζ∣ solving the task y′=⟨ζ∣x⟩ exists, it results from a linear regression of the activity leading to a new synaptic matrix16 W→W+gMμζ,

where the rank-1 matrix gM∣μ⟩⟨ζ∣ represents the changes induced by learning.

In simulations, we computed the mental line ⟨ζ∣ during the learning phase, resorting to the pseudo-inverse of the neural activity of the RNN without feedback (i.e., without re-injecting ⟨ζ∣x⟩ as the factor of ∣μ⟩). We then updated the synaptic matrix according to Eq. (16), eventually testing the response of the RNN with feedback to all possible item pairs. In this testing phase only the sensory input gS∣SR−SL⟩ is provided, as the motor plan spontaneously emerged from the recurrent activity of the linear RNN.

From Eq. (15), the asymptotic state of the network following the presentation of a pair of items can be carried out by imposing ∣x°⟩=0, which results to be17 ∣x∞⟩≡limt≫τ∣x(t)⟩=(I−W)−1{gS∣SR−SL⟩+y∞gM∣μ⟩}≡∣S~R−S~L⟩+y∞∣μ~⟩.

Here ∣S~k⟩ and ∣μ~⟩ are the inner representations of the presented items and of the motor plan, respectively, transformed by the activity reverberation in the network.

Even in this general case it is easy to prove that the geometric mental line ⟨ζ∣ solving the TI task is given by a linear combination of row vectors ⟨S~k∣≡gS−1⟨Sk∣(I−W). Indeed, for such vectors, the following relationships hold: (i) ⟨S~j∣S~k⟩=δjk and ii) ⟨S~k∣μ~⟩=0. With that, by imposing ⟨ζ∣x∞⟩=rR−rL for all the item pairs with ∣SD∣ = 1, an infinite set of geometric mental lines exist, and they are18 ζ= ∑k=1M(rk+φ)S~k+γ,

i.e., Eq. (11) reported in the main text. Note that by setting W = 0 and gS = 1, the simple linear-Perceptron configuration introduced in the main text is recovered, eventually making the above expression equivalent to Eq. (3).

From Eq. (15) we can work out the evolution in time of the readout y(t) and the projection z(t)≡⟨μ∣x⟩ expected to faithfully represent the loading of the PC1. Indeed applying to both hand sides of such equation, the “bra” ⟨μ∣ and ⟨ζ∣ we obtain the following analytical approximations19 y°=−(1−gM⟨ζ∣μ⟩)y+rR−rLz°=−z+gMy,

where we assumed that both ⟨ζ∣W∣x⟩ and ⟨μ∣W∣x⟩ are approximately equal to 0 as we verified numerically in simulation. From this expression, it can be noted that the network activity starting from ∣x⟩=0 initially unfolds is driven exclusively by the information ∣Sk−Sj⟩. As a result, the speed of changes in the readout y is given by the symbolic distance rR − rL. This explains why the slope of the trajectories in Fig. 8b is steeper for item pairs with larger rank difference. Afterward, the contribution from the activity z(t) along ∣μ⟩ starts to play a role being modulated by a y(t), as it is apparent from Fig. 8b-bottom looking at the slopes of the trajectories around PC1 = 0. We finally remark that the asymptotic value of z∞=limz→∞z(t) returns to be fully correlated with the projection of the network activity onto the GML: z∞ = gMy∞ = gM(rR − rL).

Linear RNN with sensory and endogenous noise

In Section “Serial ordering in recurrent neural networks”, the dynamics (14) of the above linear RNN incorporates units receiving a stochastic input. This is done by adding to the input ∣h⟩ an endogenous noise ∣ξ(E)⟩ uncorrelated in time whose elements ξk(E)(t) are i.i.d. random variables extracted from a Gaussian distribution with zero mean and an arbitrary variance σE′2/dt. Here, the endogenous noise aims at modeling the stochasticity expressed by biological networks of neurons working in a condition of balanced excitation and inhibition26,95. Under this condition, the membrane potential of neurons fluctuates around a subthreshold mean value finely controlled even when the synaptic contacts on the dendritic trees are huge in number. This is because the mean excitatory (positive) input is practically canceled out by the incoming inhibitory (negative) currents. Spike emission is then fluctuation-driven, and the resulting spike trains are irregular. In our theoretical framework, each unit of the RNN represents a local cell assembly composed of a finite number of K of irregularly firing neurons. For this reason, the collective dynamics of a unit (i.e., the population firing rate xk(t)) is a stochastic variable with a fluctuation size proportional to 1/K27,92,96.

Sensory input in this network configuration is given by the superposition of the contributions due to the presented items (∣hext⟩=gS∣SR−SL⟩), eventually leading to the linearized dynamics20 τx°=−(I−W)x+gSSR−SL+Wξ(E).

In this framework, the readout weights ⟨ζ∣ given by Eq. (3) make the network capable of solving, on average, the TI task provided that they are rescaled by 1/gS (i.e., ⟨ζ∣=1/gS∑k=1Mrk⟨Sk∣). Indeed, applying to both hand sides of Eq. (20) the rescaled GML, we obtain the stochastic dynamicsτy°=−y+gSζSR−SL+ζWξ(E),

where we neglected the O(g) terms approximating I − W ≃ I. Here ⟨ζ∣Wξ(E)⟩ results to be a Gaussian random variable with zero mean and variance σE2/dt≡(σE′g/gS)2∑k=1Mrk2/dt which for rk = k gives21 σE2=σE′ggS2M(1+M)(1+2M)6.

This is because the vector and matrix elements Ski, Wij and ξj(E) are i.i.d. random Gaussian variables with zero mean and variance 1/N, g2/N and σE′2/dt, respectively. Considering that for the chosen readout weights gS⟨ζ∣SR−SL⟩=rR−rL, the stochastic dynamics of the readout y(t) eventually reduce to Eq. (7) in Sect. (7) :τdy=−y+rR−rLdt+σEdW,

where y(t) is driven by a Gaussian white noise dW(t) with zero mean and covariance E[dW(t)dW(t′)]=δ(t−t′)dt.

In addition to endogenous noise, the model incorporates sensory noise to account for the variability in encoding visual stimuli across trials. The actual contribution of the presented pair of items in a generic trial n is given by: ∣hext⟩=∣Sk+ξn(R)⟩−∣Sj+ξn(L)⟩, where ∣ξn(L,R)⟩ is a vector of i.i.d. Gaussian values with zero mean and standard deviation σS′. This leads to a random contribution to the item direction ∣Sk⟩ with a variance of V[⟨Sk∣ξn(R,L)⟩]=σS′2, and for the projection along the GML V[⟨ζ∣ξn(R,L)⟩]=σS′2/gS2∑k=1Mrk2=σS2. Given rk = k, the following relationship holdsσS2=σS′gS2M(1+M)(1+2M)6.

In order to preserve the linear approximation of the dynamics, the effective intensity of sensory contribution V[gS∣hext⟩] is fixed rescaling gS by a factor 1/(1+σS′). Finally, the stochastic dynamics of the readout y(t), considering the two noise sources, is updated as followsτdy=−y+rR−rL+2σSΓndt+σEdW

where Γn is an independent random Gaussian variable with zero mean and unit variance.

Linear RNN parameters and task design

Regarding the simulated linear RNN, they have random connectivity with spectral radius g = 0.05 and strength of sensory input gS = 0.2g. Each item is schematized by injecting a square-wave current lasting 1.5 s, smoothed by a Gaussian-weighted moving average filter (0.15 s bandwidth), and followed by a relaxation time of 3 s. The state history associated with the items pairs with adjacent rank (∣SD∣ = 1) randomly presented 20 times each, is used to estimate the vector of readout weights through the pseudo-inverse approach, eventually determining a motor response when the activity of the readout unit y crosses a threshold value of 0.5. The behavior of the network for other threshold values at various feedback strengths is shown in Supplementary Fig. 2. The accuracy of the responses is tested across the whole set of possible item pairs (i.e., the testing set) presented 20 times.

The task visual stimuli are simulated by stimulating groups of nodes with a square wave current, positive when the stimulus is placed on the right side of the screen negative when it is placed on the left. As previously shown, two different sources of noise are considered: sensory (σS′) and endogenous noise (σE′). The interplay between the two sources in the operation of the recurrent neural network is illustrated in Fig. 6.

The network is trained on couples of nearest-neighbor symbols, computing the linear readout such that the state of the system gives  + 1 when the target item is on the right and  − 1 when it is on the left. The resulting output is an accumulation process that triggers a choice once a threshold of 0.5 is overcome. If the signal does not reach the threshold within the 1.5 s of the square wave stimulation, the trial is not taken into account in the performance computation.

GML with arbitrary item representations

For the sake of simplicity in the main text, we assumed that item representations on one side of the screen (left or right) are associated with opposite network states, such that ⟨SkL∣SkR⟩=−1 for any k ∈ [1, M]. Here we prove that even when this assumption is relaxed, the GML is still a linear combination of the decoding weights discriminating both the presence and the position of the k-th item on the screen, as in Eq. (3). To this purpose, let’s consider the general case of correlated representations for the same item presented in different positions (⟨Skα∣Sjβ⟩=0,⟨Skα∣Skβ⟩≠0 for any k ≠ j and α, β ∈ {L, R}). This choice is justified by the efficient coding hypothesis such that sensory information is orthogonalized via random projects onto a high-dimensional representational space61,62. However, it is reasonable to expect that some degree ρk of correlation exists between the representations of the same item appearing in different positions of the screen: ⟨Sk,L∣Sk,R⟩=ρk. In this case, the GML solving the TI task from the learning phase must solve the following set of 2(M − 1) equations⟨ζ∣Sk+1,R+Sk,L⟩=+1⟨ζ∣Sk,R+Sk+1,L⟩=−1

for any k ∈ [1, M − 1].

Making also, in this case, the ansatz ⟨ζ∣=∑α∈{L,R}∑k=1Makα⟨Skα∣, the above set of equations allows to work out the coefficients akα for the decision to choose the highest item on the right+1=ak+1,L⟨Sk+1,L∣Sk+1,R⟩+ak+1,R⟨Sk+1,R∣Sk+1,R⟩+akR⟨SkR∣SkL⟩+akL⟨SkL∣SkL⟩=ak+1,R+ak+1,Lρk+1+akL+akRρk,

or on the left−1=akR+akLρk+ak+1,L+ak+1,Rρk+1.

Summing these two equations, we obtain0=akR+akL1+ρk+ak+1,R+ak+1,L1+ρk+1

which are always solved if akL = − akR for any k ∈ [1, M]. This, together with the difference between the previous two equations, leads to have2=2akL1−ρk−2ak+1,L1−ρk+1,

eventually allowing us to write the solution22 akR=−akL=k1−ρk.

Given these coefficients, the generalized GML result to be23 ⟨ζ∣=∑k=1Mk1−ρk⟨SkR−SkL∣.

Note that in the simplified case discussed in the main text (ρk = − 1) the GML Eq. (3) is recovered as ⟨SkL∣=−⟨SkR∣.

To further prove the generality of our theoretical framework, we finally consider representations of items that are independent but correlated, i.e., ⟨Sk∣Sj⟩≠0 for all k, j ∈ [1, M]. In this scenario, we can isolate the independent component of each item by using a projector matrix P to map the item representations Sk to an orthogonal basis. Notably, previous analyses remain valid with the updated representational vectors ∣Sk*⟩=P∣Sk⟩ and the row vectors ⟨Sk*∣=⟨Sk∣P⊺, ensuring that ⟨Sk*∣Sj*⟩=0 for all k ≠ j.

GML with arbitrary motor plan

In Section “Learning the GML in RNN taking motor decisions”, we consider the motor axis as not belonging to the space of item representations, and we identify the GML family that solves the task. However, we can also find a GML solution in the case of non-orthogonality between items and motor representations.

Let us consider the more general case in which the inner product ⟨S~k∣μ~⟩=ρk, and assume that a GML solution ⟨ζ∣ exists and solves the task such that y∞=⟨ζ∣x∞⟩=SD. The GML solution ⟨ζ0∣, defined for ρk = 0 and presented in Eq. (18), projects the state as24 ⟨ζ0∣x∞⟩=SD+SD⟨ζ0∣μ~⟩.

By isolating the SD, it reads25 ⟨ζ0∣x∞⟩1+∑k(rk+φ)ρk=SD=⟨ζ∣x∞⟩,

For the sake of simplicity, we have neglected the degeneracy vector ⟨γ∣ in the orthogonal space. Therefore, the updated GML is a rescaled version of the solution defined in Eq. (18). In particular, it holds that26 ⟨ζ∣=⟨ζ0∣1+∑k(rk+φ)ρk.

Nonlinear projections on the GML

In Fig. 4 and Supplementary Fig. 1, we delve into the effects of introducing heterogeneity in environmental settings and its impact on shaping the GML. Our exploration is centered around three key variations.

The first variation pertains to the noise intensities Fig. 4a–c. We linearly map items to their ranks considering higher sensory noise associated with higher serial position k according to the linear relation σS(k) = σS(1) + β(k − 1). The network used has N = 100 units. Sensory noise was modulated according to σS(1) = 0.0075 and corruption rate β ranging from 0.0075 to 0.12 (doubling at each increment). Learning and testing trials were 100 per item. Statistics were collected from 100 independent simulations. For violin plots, data are collected from 10000 simulations for both the learning and testing phases.

The second variation involves the number of learning trials in which a specific item of the sequence appears (Supplementary Fig. 1). This presentation schedule eventually leads to the number of learning trials decreasing with the serial position k of the item as NL(k) = NL(1) − βN(k − 1). The network size is the same, and the sensory noise during the testing phase is σS = 0.3, with βN ranging from 2 to 5. Statistics were collected from 100 independent simulations.

The third variation focuses on a probabilistic reward schedule in the transitive inference task Fig. 4d–f. For each learning trial (couples with symbolic distance ∣SD∣ = 1), the reward for a correct answer is given based on a probability that is higher for higher ranks. Specifically, given a couple of symbols {Sk, Sk+1}, the probability of receiving a reward is calculated as:27 Preward(rk,rk±1)=CDF−log(1∓1/rk)2σ,

where CDF is the cumulative distribution function of the standard normal distribution, and σ is the scaling factor of the reward schedule. This approach is equivalent to considering the response target for a trial {Sk, Sk+1} as the sign of a random value η extracted from a Gaussian distribution N(−log(1∓1/rk),2σ). For this pair of items the rewarded scenario corresponds to η > 0 (right motor choice). The probability that a Gaussian random variable N(μ,σ) is positive is given by CDF(μ/σ), leading to Eq. (27). Note that 1 − CDF(− μ/σ) = CDF(μ/σ) so that the same reward probability holds for both pairs {Sk+1, Sk} and {Sk, Sk+1}. Even in this case, the network size is N = 100, and the number of learning and testing trials per pair was 10000. σ ranged from 0 to 0.35.

Dimensionality analysis

Figure 8 depicts the outcomes of the principal component analysis (PCA) conducted on the network activity simulated during the testing phase of the task. To perform PCA, the sampling matrix T × N is initially centered at zero by subtracting its time averages. In this context, T represents the number of time steps of the test phase, and N stands for the number of nodes in the network. Subsequently, the associated covariance matrix and its eigenvectors, which embody the principal components, are computed. Notably, the network state projected onto the eigenvector with the maximum eigenvalue is the projection onto the first principal component.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

Supplementary Information

Peer Review File

Description of additional supplementary files

Supplementary Movie 1

Supplementary Movie 2

Supplementary Movie 3

Reporting Summary

Supplementary information

The online version contains supplementary material available at 10.1038/s41467-024-52240-6.

Acknowledgements

We are deeply indebted to E. Brunamonti and S. Ferraina for having introduced us to the transitive inference task and for the many stimulating discussions. Work partially funded by EU H2020 Research and Innovation Program, Grant 945539 (HBP SGA3) and by the Italian National Recovery and Resilience Plan (PNRR), M4C2, NextGenerationEU (Project IR0000011, CUP B51E22000150006, ‘EBRAINS-Italy’) to M.M. Preliminary results of this work have been presented in97,98.

Author contributions

G.D.A. and M.M. conceived the model and performed the simulations. G.D.A., S.R., and M.M. designed virtual experiments to test the model. G.D.A., S.R., and M.M. contributed to the discussions and wrote the manuscript.

Peer review

Peer review information

Nature Communications thanks the anonymous, reviewers for their contribution to the peer review of this work. A peer review file is available.

Data availability

The data generated in this study can be reproduced by executing the publicly available code, as detailed in the Code Availability statement.

Code availability

All codes are available at https://github.com/gdianto/GeometricMentalLine99.

Competing interests

All the authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Logan GD Serial order in perception, memory, and action Psychol. Rev. 2021 128 1 10.1037/rev0000253 32804525
Logan, G. D. Serial order in perception, memory, and action. Psychol. Rev. 128, 1 (2021).32804525
2. Dehaene S Meyniel F Wacongne C Wang L Pallier C The neural representation of sequences: from transition probabilities to algebraic patterns and linguistic trees Neuron 2015 88 2 19 10.1016/j.neuron.2015.09.019 26447569
Dehaene, S., Meyniel, F., Wacongne, C., Wang, L. & Pallier, C. The neural representation of sequences: from transition probabilities to algebraic patterns and linguistic trees. Neuron 88, 2–19 (2015).26447569
3. Botvinick MM Watanabe T From numerosity to ordinal rank: a gain-field model of serial order representation in cortical working memory J. Neurosci. 2007 27 8636 8642 10.1523/JNEUROSCI.2110-07.2007 17687041
Botvinick, M. M. & Watanabe, T. From numerosity to ordinal rank: a gain-field model of serial order representation in cortical working memory. J. Neurosci. 27, 8636–8642 (2007).17687041
4. Liu Y Dolan RJ Kurth-Nelson Z Behrens TE Human replay spontaneously reorganizes experience Cell 2019 178 640 652 10.1016/j.cell.2019.06.012 31280961
Liu, Y., Dolan, R. J., Kurth-Nelson, Z. & Behrens, T. E. Human replay spontaneously reorganizes experience. Cell 178, 640–652 (2019).31280961
5. Frankland SM Greene JD Concepts and compositionality: in search of the brain’s language of thought Annu. Rev. Psychol. 2020 71 273 303 10.1146/annurev-psych-122216-011829 31550985
Frankland, S. M. & Greene, J. D. Concepts and compositionality: in search of the brain’s language of thought. Annu. Rev. Psychol. 71, 273–303 (2020).31550985
6. Kurth-Nelson Z Replay and compositional computation Neuron 2023 111 454 469 10.1016/j.neuron.2022.12.028 36640765
Kurth-Nelson, Z. et al. Replay and compositional computation. Neuron 111, 454–469 (2023).36640765
7. Rigotti M The importance of mixed selectivity in complex cognitive tasks Nature 2013 497 585 590 10.1038/nature12160 23685452
Rigotti, M. et al. The importance of mixed selectivity in complex cognitive tasks. Nature 497, 585–590 (2013).23685452
8. Mante V Sussillo D Shenoy KV Newsome WT Context-dependent computation by recurrent dynamics in prefrontal cortex Nature 2013 503 78 84 10.1038/nature12742 24201281
Mante, V., Sussillo, D., Shenoy, K. V. & Newsome, W. T. Context-dependent computation by recurrent dynamics in prefrontal cortex. Nature 503, 78–84 (2013).24201281
9. Fusi S Miller EK Rigotti M Why neurons mix: high dimensionality for higher cognition Curr. Opin. Neurobiol. 2016 37 66 74 10.1016/j.conb.2016.01.010 26851755
Fusi, S., Miller, E. K. & Rigotti, M. Why neurons mix: high dimensionality for higher cognition. Curr. Opin. Neurobiol. 37, 66–74 (2016).26851755
10. Xie Y Geometry of sequence working memory in macaque prefrontal cortex Science 2022 375 632 639 10.1126/science.abm0204 35143322
Xie, Y. et al. Geometry of sequence working memory in macaque prefrontal cortex. Science 375, 632–639 (2022).35143322
11. Bengio Y Courville A Vincent P Representation learning: A review and new perspectives IEEE Trans. Pattern Anal. Mach. Intell. 2013 35 1798 1828 10.1109/TPAMI.2013.50 23787338
Bengio, Y., Courville, A. & Vincent, P. Representation learning: A review and new perspectives. IEEE Trans. Pattern Anal. Mach. Intell. 35, 1798–1828 (2013).23787338
12. Flesch T Juechems K Dumbalska T Saxe A Summerfield C Orthogonal representations for robust context-dependent task performance in brains and neural networks Neuron 2022 110 1258 1270 10.1016/j.neuron.2022.01.005 35085492
Flesch, T., Juechems, K., Dumbalska, T., Saxe, A. & Summerfield, C. Orthogonal representations for robust context-dependent task performance in brains and neural networks. Neuron 110, 1258–1270 (2022).35085492
13. Trabasso, T. & Riley, C. A. On the construction and use of representations involving linear order. In Solso, R. L. (ed.) Information Processing and Cognition: The Loyola Symposium, 381–410 (Lawrence Erlbaum, 1975).
14. Fias, W., van Dijck, J.-P. & Gevers, W. Chapter 10 - how is number associated with space? the role of working memory. In Dehaene, S. & Brannon, E. M. (eds.) Space, Time and Number in the Brain, 133–148 (Academic Press, San Diego, 2011).
15. Bonato M Zorzi M Umiltà C When time is space: Evidence for a mental time line Neurosci. Biobehav. Rev. 2012 36 2257 2273 10.1016/j.neubiorev.2012.08.007 22935777
Bonato, M., Zorzi, M. & Umiltà, C. When time is space: Evidence for a mental time line. Neurosci. Biobehav. Rev. 36, 2257–2273 (2012).22935777
16. Gazes, R. P., Templer, V. L. & Lazareva, O. F. Thinking about order: a review of common processing of magnitude and learned orders in animals. Anim. Cogn. 26, 1–9 (2022).
17. Terrace HS Son LK Brannon EM Serial expertise of rhesus macaques Psychol. Sci. 2003 14 66 73 10.1111/1467-9280.01420 12564756
Terrace, H. S., Son, L. K. & Brannon, E. M. Serial expertise of rhesus macaques. Psychol. Sci. 14, 66–73 (2003).12564756
18. Jensen, G. Serial learning. In APA Handbook of Comparative Psychology: Perception, learning, and Cognition, Vol. 2, APA Handbooks in psychology®., 385–409 (American Psychological Association, 2017).
19. Rosenblatt F The perceptron: a probabilistic model for information storage and organization in the brain Psychol. Rev. 1958 65 386 408 10.1037/h0042519 13602029
Rosenblatt, F. The perceptron: a probabilistic model for information storage and organization in the brain. Psychol. Rev. 65, 386–408 (1958).13602029
20. Amit, D. J.Modeling Brain Function (Cambridge University Press, 1989).
21. Asaad WF Rainer G Miller EK Neural activity in the primate prefrontal cortex during associative learning Neuron 1998 21 1399 1407 10.1016/S0896-6273(00)80658-3 9883732
Asaad, W. F., Rainer, G. & Miller, E. K. Neural activity in the primate prefrontal cortex during associative learning. Neuron 21, 1399–1407 (1998).9883732
22. Sigala N Kusunoki M Nimmo-Smith I Gaffan D Duncan J Hierarchical coding for sequential task events in the monkey prefrontal cortex Proc. Natl. Acad. Sci. USA 2008 105 11969 74 10.1073/pnas.0802569105 18689686
Sigala, N., Kusunoki, M., Nimmo-Smith, I., Gaffan, D. & Duncan, J. Hierarchical coding for sequential task events in the monkey prefrontal cortex. Proc. Natl. Acad. Sci. USA 105, 11969–74 (2008).18689686
23. Salzman CD Fusi S Emotion, cognition, and mental state representation in amygdala and prefrontal cortex Annu. Rev. Neurosci. 2010 33 173 202 10.1146/annurev.neuro.051508.135256 20331363
Salzman, C. D. & Fusi, S. Emotion, cognition, and mental state representation in amygdala and prefrontal cortex. Annu. Rev. Neurosci. 33, 173–202 (2010).20331363
24. Barak O Rigotti M Fusi S The sparseness of mixed selectivity neurons controls the generalization-discrimination trade-off J. Neurosci. 2013 33 3844 56 10.1523/JNEUROSCI.2753-12.2013 23447596
Barak, O., Rigotti, M. & Fusi, S. The sparseness of mixed selectivity neurons controls the generalization-discrimination trade-off. J. Neurosci. 33, 3844–56 (2013).23447596
25. Dirac, P. A. M. The Principles of Quantum Mechanics (Oxford University Press, London, 1958).
26. van Vreeswijk C Sompolinsky H Chaos in neuronal networks with balanced excitatory and inhibitory activity Science 1996 274 1724 6 10.1126/science.274.5293.1724 8939866
van Vreeswijk, C. & Sompolinsky, H. Chaos in neuronal networks with balanced excitatory and inhibitory activity. Science 274, 1724–6 (1996).8939866
27. Mattia M Del Giudice P Population dynamics of interacting spiking neurons Phys. Rev. E 2002 66 051917 10.1103/PhysRevE.66.051917
Mattia, M. & Del Giudice, P. Population dynamics of interacting spiking neurons. Phys. Rev. E 66, 051917 (2002).
28. Jensen G Alkan Y Muñoz F Ferrera VP Terrace HS Transitive inference in humans (Homo sapiens) and rhesus macaques (Macaca mulatta) after massed training of the last two list items J. Comp. Psychol. 2017 131 231 245 10.1037/com0000065 28333486
Jensen, G., Alkan, Y., Muñoz, F., Ferrera, V. P. & Terrace, H. S. Transitive inference in humans (Homo sapiens) and rhesus macaques (Macaca mulatta) after massed training of the last two list items. J. Comp. Psychol. 131, 231–245 (2017).28333486
29. Lazareva, O. F. Transitive inference in nonhuman animals. In The Oxford Handbook of Comparative Cognition. 718–735 (Oxford University Press, New York, NY, US, 2012).
30. Luyckx F Nili H Spitzer B Summerfield C Neural structure mapping in human probabilistic reward learning Elife 2019 8 e42816 10.7554/eLife.42816 30843789
Luyckx, F., Nili, H., Spitzer, B. & Summerfield, C. Neural structure mapping in human probabilistic reward learning. Elife 8, e42816 (2019).30843789
31. Sheahan H Luyckx F Nelli S Teupe C Summerfield C Neural state space alignment for magnitude generalization in humans and recurrent networks Neuron 2021 109 1214 1226 10.1016/j.neuron.2021.02.004 33626322
Sheahan, H., Luyckx, F., Nelli, S., Teupe, C. & Summerfield, C. Neural state space alignment for magnitude generalization in humans and recurrent networks. Neuron 109, 1214–1226 (2021).33626322
32. Nelli S Braun L Dumbalska T Saxe A Summerfield C Neural knowledge assembly in humans and neural networks Neuron 2023 111 1504 1516 10.1016/j.neuron.2023.02.014 36898375
Nelli, S., Braun, L., Dumbalska, T., Saxe, A. & Summerfield, C. Neural knowledge assembly in humans and neural networks. Neuron 111, 1504–1516 (2023).36898375
33. Munoz F Learned representation of implied serial order in posterior parietal cortex Sci. Rep. 2020 10 1 14 10.1038/s41598-020-65838-9 31913322
Munoz, F. et al. Learned representation of implied serial order in posterior parietal cortex. Sci. Rep. 10, 1–14 (2020).31913322
34. Vasconcelos M Transitive inference in non-human animals: an empirical and theoretical analysis Behav. Processes 2008 78 313 34 10.1016/j.beproc.2008.02.017 18423898
Vasconcelos, M. Transitive inference in non-human animals: an empirical and theoretical analysis. Behav. Processes 78, 313–34 (2008).18423898
35. Brunamonti E Neuronal modulation in the prefrontal cortex in a transitive inference task: Evidence of neuronal correlates of mental schema management J. Neurosci. 2016 36 1223 1236 10.1523/JNEUROSCI.1473-15.2016 26818510
Brunamonti, E. et al. Neuronal modulation in the prefrontal cortex in a transitive inference task: Evidence of neuronal correlates of mental schema management. J. Neurosci. 36, 1223–1236 (2016).26818510
36. Leth-Steensen C Marley AAJ A model of response time effects in symbolic comparison Psychol. Rev. 2000 107 62 100 10.1037/0033-295X.107.1.62 10687403
Leth-Steensen, C. & Marley, A. A. J. A model of response time effects in symbolic comparison. Psychol. Rev. 107, 62–100 (2000).10687403
37. Verguts T Van Opstal F A delta-rule model of numerical and non-numerical order processing J. Exp. Psychol. Human. Percept. Perform. 2014 40 1092 1102 10.1037/a0035114 24392740
Verguts, T. & Van Opstal, F. A delta-rule model of numerical and non-numerical order processing. J. Exp. Psychol. Human. Percept. Perform. 40, 1092–1102 (2014).24392740
38. Widrow, B. & Hoff, M. E. Adaptive Switching Circuits. Stanford Electronics Laboratories, Stanford University, Stanford (1960).
39. Hertz, J. A., Krogh, A. & Palmer, R. G. Introduction to the Theory of Neural Computation (West view Press, Boca Raton, 1991).
40. Rescorla, R. & Wagner, A. R. A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and non-reinforcement. In Black, A. H. & Prokasy, W. F. (eds.) Classical Conditioning II: Current Research and Theory, 64–99 (Appleton-Century-Crofts, New York, NY, 1972).
41. Duda, R. O., Hart, P. E. & Stork, D. D. Pattern Classification (John Wiley & Sons, New York, NY, 2001).
42. Treichler FR Van Tilburg D Concurrent conditional discrimination tests of transitive inference by macaque monkeys: List linking J. Exp. Psychol. Anim. Behav. Process 1996 22 105 117 10.1037/0097-7403.22.1.105 8568492
Treichler, F. R. & Van Tilburg, D. Concurrent conditional discrimination tests of transitive inference by macaque monkeys: List linking. J. Exp. Psychol. Anim. Behav. Process 22, 105–117 (1996).8568492
43. Jensen G Alkan Y Ferrera VP Terrace HS Reward associations do not explain transitive inference performance in monkeys Sci. Adv. 2019 5 eaaw2089 10.1126/sciadv.aaw2089 32128384
Jensen, G., Alkan, Y., Ferrera, V. P. & Terrace, H. S. Reward associations do not explain transitive inference performance in monkeys. Sci. Adv. 5, eaaw2089 (2019).32128384
44. Mione V Brunamonti E Pani P Genovesio A Ferraina S Dorsal premotor cortex neurons signal the level of choice difficulty during logical decisions Cell Rep. 2020 32 107961 10.1016/j.celrep.2020.107961 32726625
Mione, V., Brunamonti, E., Pani, P., Genovesio, A. & Ferraina, S. Dorsal premotor cortex neurons signal the level of choice difficulty during logical decisions. Cell Rep. 32, 107961 (2020).32726625
45. Page M Norris D The primacy model: a new model of immediate serial recall Psychol. Rev. 1998 105 761 10.1037/0033-295X.105.4.761-781 9830378
Page, M. & Norris, D. The primacy model: a new model of immediate serial recall. Psychol. Rev. 105, 761 (1998).9830378
46. Baddeley AD How does acoustic similarity influence short-term memory? Q. J. Exp. Psychol. 1968 20 249 264 10.1080/14640746808400159 5683764
Baddeley, A. D. How does acoustic similarity influence short-term memory? Q. J. Exp. Psychol. 20, 249–264 (1968).5683764
47. Drewnowski A Murdock BB The role of auditory features in memory span for words J. Exp. Psychol. Hum. Learn. Mem. 1980 6 319 10.1037/0278-7393.6.3.319
Drewnowski, A. & Murdock, B. B. The role of auditory features in memory span for words. J. Exp. Psychol. Hum. Learn. Mem. 6, 319 (1980).
48. Miller GA The magical number seven plus or minus two: some limits on our capacity for processing information Psychol. Rev. 1956 63 81 97 10.1037/h0043158 13310704
Miller, G. A. The magical number seven plus or minus two: some limits on our capacity for processing information. Psychol. Rev. 63, 81–97 (1956).13310704
49. Baddeley A The magical number seven: still magic after all these years? Psychol. Rev. 1994 101 353 6 10.1037/0033-295X.101.2.353 8022967
Baddeley, A. The magical number seven: still magic after all these years? Psychol. Rev. 101, 353–6 (1994).8022967
50. Grossberg S Behavioral contrast in short term memory: serial binary memory models or parallel continuous memory models? J. Math. Psychol. 1978 17 199 219 10.1016/0022-2496(78)90016-0
Grossberg, S. Behavioral contrast in short term memory: serial binary memory models or parallel continuous memory models? J. Math. Psychol. 17, 199–219 (1978).
51. Henson RNA Short-term memory for serial order: the start-end model Cogn. Psychol. 1998 36 73 137 10.1006/cogp.1998.0685 9721198
Henson, R. N. A. Short-term memory for serial order: the start-end model. Cogn. Psychol. 36, 73–137 (1998).9721198
52. Averbeck BB Chafee MV a Crowe D Georgopoulos AP Parallel processing of serial movements in prefrontal cortex Proc. Natl. Acad. Sci. USA 2002 99 13172 7 10.1073/pnas.162485599 12242330
Averbeck, B. B., Chafee, M. V., a Crowe, D. & Georgopoulos, A. P. Parallel processing of serial movements in prefrontal cortex. Proc. Natl. Acad. Sci. USA 99, 13172–7 (2002).12242330
53. Henson RNA Norris DG Page MPA Baddeley AD Unchained memory: Error patterns rule out chaining models of immediate serial recall Q. J. Exp. Psychol. 1996 49 80 115 10.1080/713755612
Henson, R. N. A., Norris, D. G., Page, M. P. A. & Baddeley, A. D. Unchained memory: Error patterns rule out chaining models of immediate serial recall. Q. J. Exp. Psychol. 49, 80–115 (1996).
54. Manoochehri M Up to the magical number seven: An evolutionary perspective on the capacity of short term memory Heliyon 2021 7 e06955 10.1016/j.heliyon.2021.e06955 34013087
Manoochehri, M. Up to the magical number seven: An evolutionary perspective on the capacity of short term memory. Heliyon 7, e06955 (2021).34013087
55. Cox, D. R. & Miller, H. D. The Theory of Stochastic Processes (CRC Press, Boca Raton, 1965).
56. Gold JI Shadlen MN The neural basis of decision making Annu. Rev. Neurosci. 2007 30 535 74 10.1146/annurev.neuro.29.051605.113038 17600525
Gold, J. I. & Shadlen, M. N. The neural basis of decision making. Annu. Rev. Neurosci. 30, 535–74 (2007).17600525
57. Acuna BD Sanes JN Donoghue JP Cognitive mechanisms of transitive inference Exp. Brain Res. 2002 146 1 10 10.1007/s00221-002-1092-y 12192572
Acuna, B. D., Sanes, J. N. & Donoghue, J. P. Cognitive mechanisms of transitive inference. Exp. Brain Res. 146, 1–10 (2002).12192572
58. Okazawa G Hatch CE Mancoo A Machens CK Kiani R Representational geometry of perceptual decisions in the monkey parietal cortex Cell 2021 184 3748 3761 10.1016/j.cell.2021.05.022 34171308
Okazawa, G., Hatch, C. E., Mancoo, A., Machens, C. K. & Kiani, R. Representational geometry of perceptual decisions in the monkey parietal cortex. Cell 184, 3748–3761 (2021).34171308
59. Sussillo D Churchland MM Kaufman MT Shenoy KV A neural network that finds a naturalistic solution for the production of muscle activity Nat. Neurosci. 2015 18 1025 33 10.1038/nn.4042 26075643
Sussillo, D., Churchland, M. M., Kaufman, M. T. & Shenoy, K. V. A neural network that finds a naturalistic solution for the production of muscle activity. Nat. Neurosci. 18, 1025–33 (2015).26075643
60. Yuste R From the neuron doctrine to neural networks Nat. Rev. Neurosci. 2015 16 487 497 10.1038/nrn3962 26152865
Yuste, R. From the neuron doctrine to neural networks. Nat. Rev. Neurosci. 16, 487–497 (2015).26152865
61. Barlow, H. B. Possible principles underlying the transformation of sensory messages. In Rosenblith, W. A. (ed.) Sensory Communication, 217–34 (MIT Press, Cambridge, MA, 1961).
62. Simoncelli EP Olshausen BA Natural image statistics and neural representation Annu. Rev. Neurosci. 2001 24 1193 1216 10.1146/annurev.neuro.24.1.1193 11520932
Simoncelli, E. P. & Olshausen, B. A. Natural image statistics and neural representation. Annu. Rev. Neurosci. 24, 1193–1216 (2001).11520932
63. Perich MG Gallego JA Miller LE A neural population mechanism for rapid learning Neuron 2018 100 964–976 10.1016/j.neuron.2018.09.030 30344047
Perich, M. G., Gallego, J. A. & Miller, L. E. A neural population mechanism for rapid learning. Neuron 100, 964–976 (2018).30344047
64. Vyas S Golub MD Sussillo D Shenoy KV Computation through neural population dynamics Annu. Rev. Neurosci. 2020 43 249 10.1146/annurev-neuro-092619-094115 32640928
Vyas, S., Golub, M. D., Sussillo, D. & Shenoy, K. V. Computation through neural population dynamics. Annu. Rev. Neurosci. 43, 249 (2020).32640928
65. Aoi MC Mante V Pillow JW Prefrontal cortex exhibits multidimensional dynamic encoding during decision-making Nat. Neurosci. 2020 23 1410 20 10.1038/s41593-020-0696-5 33020653
Aoi, M. C., Mante, V. & Pillow, J. W. Prefrontal cortex exhibits multidimensional dynamic encoding during decision-making. Nat. Neurosci. 23, 1410–20 (2020).33020653
66. Sheng J Higher-dimensional neural representations predict better episodic memory Sci. Adv. 2022 8 eabm3829 10.1126/sciadv.abm3829 35442734
Sheng, J. et al. Higher-dimensional neural representations predict better episodic memory. Sci. Adv. 8, eabm3829 (2022).35442734
67. Stiso, J. et al. Neurophysiological evidence for cognitive map formation during sequence learning. Eneuro 9, 10.1523/eneuro.0361-21.2022 (2022).
68. Park SA Miller DS Boorman ED Inferences on a multidimensional social hierarchy use a grid-like code Nat. Neurosci. 2021 24 1292 1301 10.1038/s41593-021-00916-3 34465915
Park, S. A., Miller, D. S. & Boorman, E. D. Inferences on a multidimensional social hierarchy use a grid-like code. Nat. Neurosci. 24, 1292–1301 (2021).34465915
69. Bernardi S The geometry of abstraction in the hippocampus and prefrontal cortex Cell 2020 183 954 967 10.1016/j.cell.2020.09.031 33058757
Bernardi, S. et al. The geometry of abstraction in the hippocampus and prefrontal cortex. Cell 183, 954–967 (2020).33058757
70. Kriegeskorte N Kievit RA Representational geometry: integrating cognition, computation, and the brain Trends Cogn. Sci. 2013 17 401 412 10.1016/j.tics.2013.06.007 23876494
Kriegeskorte, N. & Kievit, R. A. Representational geometry: integrating cognition, computation, and the brain. Trends Cogn. Sci. 17, 401–412 (2013).23876494
71. Chung S Abbott L Neural population geometry: An approach for understanding biological and artificial neural networks Curr. Opin. Neurobiol. 2021 70 137 144 10.1016/j.conb.2021.10.010 34801787
Chung, S. & Abbott, L. Neural population geometry: An approach for understanding biological and artificial neural networks. Curr. Opin. Neurobiol. 70, 137–144 (2021).34801787
72. D’Amato MR Colombo M The symbolic distance effect in monkeys (cebus apella) Anim. Learn. Behav. 1990 18 133 140 10.3758/BF03205250
D’Amato, M. R. & Colombo, M. The symbolic distance effect in monkeys (cebus apella). Anim. Learn. Behav. 18, 133–140 (1990).
73. Davis H Transitive inference in rats (rattus norvegicus) J. Comp. Psychol. 1992 106 342 349 10.1037/0735-7036.106.4.342 1451416
Davis, H. Transitive inference in rats (rattus norvegicus). J. Comp. Psychol. 106, 342–349 (1992).1451416
74. Terrace H A nonverbal organism’s knowledge of ordinal position in a serial learning task J. Exp. Psychol. Anim. Behav. Process. 1986 12 203 10.1037/0097-7403.12.3.203
Terrace, H. A nonverbal organism’s knowledge of ordinal position in a serial learning task. J. Exp. Psychol. Anim. Behav. Process. 12, 203 (1986).
75. Dusek JA Eichenbaum H The hippocampus and memory for orderly stimulus relations Proc. Natl. Acad. Sci. USA 1997 94 7109 7114 10.1073/pnas.94.13.7109 9192700
Dusek, J. A. & Eichenbaum, H. The hippocampus and memory for orderly stimulus relations. Proc. Natl. Acad. Sci. USA 94, 7109–7114 (1997).9192700
76. Jensen G Munoz F Meaney A Terrace HS Ferrera VP Transitive inference after minimal training in rhesus macaques (Macaca mulatta) J. Exp. Psychol. Anim. Learn. Cogn. 2021 47 464 475 10.1037/xan0000298 34855434
Jensen, G., Munoz, F., Meaney, A., Terrace, H. S. & Ferrera, V. P. Transitive inference after minimal training in rhesus macaques (Macaca mulatta). J. Exp. Psychol. Anim. Learn. Cogn. 47, 464–475 (2021).34855434
77. Treichler FR Raghanti MA Van Tilburg DN Serial list linking by macaque monkeys (macaca mulatta): list property limitations J. Comp. Psychol. 2007 121 250 10.1037/0735-7036.121.3.250 17696651
Treichler, F. R., Raghanti, M. A. & Van Tilburg, D. N. Serial list linking by macaque monkeys (macaca mulatta): list property limitations. J. Comp. Psychol. 121, 250 (2007).17696651
78. Ramawat S Different contribution of the monkey prefrontal and premotor dorsal cortex in decision making during a transitive inference task Neuroscience 2022 485 147 162 10.1016/j.neuroscience.2022.01.013 35193770
Ramawat, S. et al. Different contribution of the monkey prefrontal and premotor dorsal cortex in decision making during a transitive inference task. Neuroscience 485, 147–162 (2022).35193770
79. Lazareva OF Wasserman EA Transitive inference in pigeons: Measuring the associative values of stimuli b and d Behav. Process. 2012 89 244 255 10.1016/j.beproc.2011.12.001
Lazareva, O. F. & Wasserman, E. A. Transitive inference in pigeons: Measuring the associative values of stimuli b and d. Behav. Process. 89, 244–255 (2012).
80. Dehaene S Mehler J Cross-linguistic regularities in the frequency of number words Cognition 1992 43 1 29 10.1016/0010-0277(92)90030-L 1591901
Dehaene, S. & Mehler, J. Cross-linguistic regularities in the frequency of number words. Cognition 43, 1–29 (1992).1591901
81. Dehaene S Varieties of numerical abilities Cognition 1992 44 1 42 10.1016/0010-0277(92)90049-N 1511583
Dehaene, S. Varieties of numerical abilities. Cognition 44, 1–42 (1992).1511583
82. Gallistel CR Gelman R Preverbal and verbal counting and computation Cognition 1992 44 43 74 10.1016/0010-0277(92)90050-R 1511586
Gallistel, C. R. & Gelman, R. Preverbal and verbal counting and computation. Cognition 44, 43–74 (1992).1511586
83. Verguts T Fias W Stevens M A model of exact small-number representation Psychon. bull. Rev. 2005 12 66 80 10.3758/BF03196349 15945201
Verguts, T., Fias, W. & Stevens, M. A model of exact small-number representation. Psychon. bull. Rev. 12, 66–80 (2005).15945201
84. Dehaene S The neural basis of the weber-fechner law: a logarithmic mental number line Trends Cogn. Sci. 2003 7 145 147 10.1016/S1364-6613(03)00055-X 12691758
Dehaene, S. The neural basis of the weber-fechner law: a logarithmic mental number line. Trends Cogn. Sci. 7, 145–147 (2003).12691758
85. Algom D The Weber-Fechner law: A misnomer that persists but that should go away Psychol. Rev. 2021 128 757 765 10.1037/rev0000278 34242050
Algom, D. The Weber-Fechner law: A misnomer that persists but that should go away. Psychol. Rev. 128, 757–765 (2021).34242050
86. Terrace HS The simultaneous chain: a new approach to serial learning Trends Cogn. Sci. 2005 9 202 10 10.1016/j.tics.2005.02.003 15808503
Terrace, H. S. The simultaneous chain: a new approach to serial learning. Trends Cogn. Sci. 9, 202–10 (2005).15808503
87. Kumaran D McClelland JL Generalization through the recurrent interaction of episodic memories: a model of the hippocampal system Psychol. Rev. 2012 119 573 10.1037/a0028681 22775499
Kumaran, D. & McClelland, J. L. Generalization through the recurrent interaction of episodic memories: a model of the hippocampal system. Psychol. Rev. 119, 573 (2012).22775499
88. Lippl, S., Kay, K., Jensen, G., Ferrera, V. P. & Abbott, L. A mathematical theory of relational generalization in transitive inference. Proc. Natl Acad. Sci. 121, e2314511121 (2024).
89. Jensen G Terrace HS Ferrera VP Discovering implied serial order through model-free and model-based learning Front. Neurosci. 2019 13 1 24 10.3389/fnins.2019.00878 30740042
Jensen, G., Terrace, H. S. & Ferrera, V. P. Discovering implied serial order through model-free and model-based learning. Front. Neurosci. 13, 1–24 (2019).30740042
90. Abbott LF van Vreeswijk C Asynchronous states in networks of pulse-coupled oscillators Phys. Rev. E 1993 48 1483 1490 10.1103/PhysRevE.48.1483
Abbott, L. F. & van Vreeswijk, C. Asynchronous states in networks of pulse-coupled oscillators. Phys. Rev. E 48, 1483–1490 (1993).
91. Treves A Mean-field analysis of neuronal spike dynamics Network 1993 4 259 84 10.1088/0954-898X_4_3_002
Treves, A. Mean-field analysis of neuronal spike dynamics. Network 4, 259–84 (1993).
92. Brunel N Hakim V Fast global oscillations in networks of integrate-and-fire neurons with low firing rates Neural Comput. 1999 11 1621 71 10.1162/089976699300016179 10490941
Brunel, N. & Hakim, V. Fast global oscillations in networks of integrate-and-fire neurons with low firing rates. Neural Comput. 11, 1621–71 (1999).10490941
93. Knight BW Dynamics of encoding in neuron populations: some general mathematical features Neural Comput. 2000 12 473 518 10.1162/089976600300015673 10769319
Knight, B. W. Dynamics of encoding in neuron populations: some general mathematical features. Neural Comput. 12, 473–518 (2000).10769319
94. Jaeger H Haas H Harnessing nonlinearity: predicting chaotic systems and saving energy in wireless communication Science 2004 304 78 80 10.1126/science.1091277 15064413
Jaeger, H. & Haas, H. Harnessing nonlinearity: predicting chaotic systems and saving energy in wireless communication. Science 304, 78–80 (2004).15064413
95. Renart A The asynchronous state in cortical circuits Science 2010 327 587 90 10.1126/science.1179850 20110507
Renart, A. et al. The asynchronous state in cortical circuits. Science 327, 587–90 (2010).20110507
96. Vinci GV Benzi R Mattia M Self-consistent stochastic dynamics for finite-size networks of spiking neurons Phys. Rev. Lett. 2023 130 097402 10.1103/PhysRevLett.130.097402 36930929
Vinci, G. V., Benzi, R. & Mattia, M. Self-consistent stochastic dynamics for finite-size networks of spiking neurons. Phys. Rev. Lett. 130, 097402 (2023).36930929
97. Di Antonio, G., Raglio, S. & Mattia, M. Ranking and serial thinking: a geometrical solution. In Bernstein Conference 2022 (Bernstein Network, Berlin, 2022).
98. Raglio, S., Di Antonio, G., Brunamonti, E., Ferraina, S. & Mattia, M. Learning to infer transitively: ranking symbols on a mental line in premotor cortex. In 2022 Neuroscience Meeting Planner, vol. Program No. 564.19 (Society for Neuroscience, San Diego, CA, 2022).
99. Di Antonio, G. & Mattia, M. Code for “Ranking and serial thinking: A geometric solution”. Zenodo Deposited on 8 August 2024 (2024).
