
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

71678
10.1038/s41598-024-71678-8
Article
Learning long sequences in spiking neural networks
https://orcid.org/0000-0003-2726-1860
Stan Matei-Ioan matei.stan@manchester.ac.uk

https://orcid.org/0000-0003-1728-2828
Rhodes Oliver
https://ror.org/027m9bs27 grid.5379.8 0000 0001 2166 2407 Department of Computer Science, The University of Manchester, Manchester, UK
20 9 2024
20 9 2024
2024
14 2195730 1 2024
29 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by/4.0/ Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
Spiking neural networks (SNNs) take inspiration from the brain to enable energy-efficient computations. Since the advent of Transformers, SNNs have struggled to compete with artificial networks on modern sequential tasks, as they inherit limitations from recurrent neural networks (RNNs), with the added challenge of training with non-differentiable binary spiking activations. However, a recent renewed interest in efficient alternatives to Transformers has given rise to state-of-the-art recurrent architectures named state space models (SSMs). This work systematically investigates, for the first time, the intersection of state-of-the-art SSMs with SNNs for long-range sequence modelling. Results suggest that SSM-based SNNs can outperform the Transformer on all tasks of a well-established long-range sequence modelling benchmark. It is also shown that SSM-based SNNs can outperform current state-of-the-art SNNs with fewer parameters on sequential image classification. Finally, a novel feature mixing layer is introduced, improving SNN accuracy while challenging assumptions about the role of binary activations in SNNs. This work paves the way for deploying powerful SSM-based architectures, such as large language models, to neuromorphic hardware for energy-efficient long-range sequence modelling.

Keywords

Spiking neural networks
State space models
Sequence modelling
Long range dependencies
Subject terms

Computational models
Long-term memory
Short-term memory
EU’s Horizon Europe Research and Innovation programmeGrant Agreement 101070679 UK Research and Innovation (UKRI) under the UK government’s Horizon Europe funding guaranteeGrant Agreement 10039070 issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Modelling long-range sequences is a fundamental component in solving many real-world challenges, with applications ranging from processing biosignals such as electroencephalograms spanning tens of thousands of time steps1, to comprehending and potentially writing large documents (e.g., novels, scientific papers) using large language models2,3.

Deep learning methods have established themselves as state-of-the-art solutions for numerous challenging tasks, including learning functions defined over variable-length input sequences. Recurrent neural network (RNN) architectures emerged early on as strong contenders for this purpose. They compress sequences by incorporating input elements one at a time, using only O(1) operations with respect to the sequence length to process each input token and sharing parameters between time steps (Fig. 1a). Notably, RNNs are partially inspired by cognitive and neurological computational principles4. Hence, perhaps unsurprisingly, they also underpin another class of biologically grounded architectures - spiking neural networks (SNNs) (Fig. 1b). SNNs process sequences using simplified mathematical models of biological neurons that relay internal computations using sparse patterns of binary spikes5. The aim is to emulate the brain’s efficient neural coding, which not only enables computing with a fraction of the energy required by modern von Neumann machines6, but may also support various cortical functions such as lifelong continual learning7,8.

RNNs are affected by vanishing and exploding gradients9, stemming from unstable recurrent weight initialisation and the use of backpropagation through time (BPTT) (Fig. 1a). These phenomena hinder learning long-range dependencies in RNNs, and while they can be mitigated to some extent by gating mechanisms such as long short-term memory (LSTM)10, they are difficult to eliminate entirely. In addition, traditional RNNs apply nonlinearities at each time step (σ in Fig. 1a), which requires iterative computations. This approach is non-problematic at inference, where input sequence elements are unknown ahead of time. However, RNN forward passes become prohibitively slow at training time for long sequences, since they cannot take advantage of GPU parallelisation, owing to the nonlinear state propagation11–13.

Additional challenges arise in SNN learning, as binary spiking is non-differentiable, which prohibits training SNNs directly with backpropagation. One solution is to train an artificial neural network (ANN) and then convert its continuous activations to spikes14. However, this approach introduces additional latency during inference and is often prone to excessive firing, which can damage the energy efficiency of the network15. ANN-to-SNN conversion has been shown to achieve near-lossless accuracies on static datasets16,17. However, it has been argued that it is not as suitable for tasks requiring temporal information18, which would make it suboptimal in the context of the present study. For learning temporal information, significant research effort has also been directed towards biologically-inspired algorithms which do not require the differentiation of spikes. One of the most well-known examples is Spike-Timing Dependent Plasticity (STDP), which only uses local firing times from pre and post-synaptic neurons to update each synaptic weight18. However, arguably the most successful solution so far for learning temporal dependencies has been to train SNNs directly using BPTT and surrogate gradients in the backward pass19,20, hence this is the method employed in this study. Nevertheless, even with direct training, SNNs are still generally outperformed by ANNs such as LSTMs21.

The RNN limitations mentioned above are overcome by the Transformer22, which directly compresses the context for each token by measuring its relationship to all other elements (Fig. 1c). Besides improving performance, the Transformer’s core component, self-attention, can be easily parallelised through GPU-friendly matrix multiplication, which accelerates training relative to RNNs23. Consequently, Transformer blocks have been crucial in establishing the current golden age of ever-larger pre-trained models24.

The parallel and dense matrix multiplications that have entrenched the Transformer as arguably the de facto standard in sequence modelling also accentuated the structural differences between SNNs and ANNs. SNNs are built for deployment on neuromorphic computing platforms such as Intel Loihi25, which can potentially enable orders of magnitude lower energy consumption compared to traditional computers. These efficiencies are partly supported by representing information as sparse events identified by their address. Spike events then “excite” the targeted synapses asynchronously, with accumulation occurring within the postsynaptic neurons’ internal states. This enables addition-based feature mixing, reducing costly Multiply-and-Accumulate (MAC) operations26. Massive parallel matrix multiplications, as self-attention requires, can be seen as antagonistic to this event-driven and brain-inspired computing philosophy. Therefore, lessons from Transformer-based research have seen relatively limited adoption in SNNs by comparison27–29.

Nevertheless, self-attention suffers a quadratic computational cost with respect to sequence length30, which effectively limits scaling to longer sequences. In addition, training and inference for large-scale Transformer-based models have seen significant increases in energy requirements, leading to considerable carbon emissions31. This highlights the need for energy-efficient models which scale better with input length, a role recurrent SNNs are potentially well-positioned to fill.

The quadratic computational cost has motivated a recent resurgence in RNN research interest. Receptance Weighted Key Value (RWKV)32, exemplifies research focused on reducing the computational complexity of Transformers. It is essentially a recurrent self-attention adaptation allowing O(1) iterative deployment. Another area of research is focused on deriving RNNs with theoretical guarantees regarding long-range modelling properties. For example, the Legendre Memory Unit (LMU), takes inspiration from hippocampal neurons to augment RNN nonlinear state propagation with a linear memory component Voelker et al.33. The memory unit is constructed using linear projections of input signals onto a Legendre orthogonal polynomial basis. The result is a multidimensional cell state that is theoretically guaranteed to encode a sliding window of a given number of past inputs. This enabled the LMU to become the first recurrent model to successfully capture temporal dependencies on the scale of 100,000 time steps33. Chilkuri and Eliasmith34 remove the remaining nonlinear recurrences in the LMU to obtain a linear time-invariant (LTI) structure with position-wise activations (Fig. 1d). LTI systems have the property of having two equivalent formulations: iterative propagation of the system’s state by repeated application of the linear recurrence; or a convolution of the input signal with a global filter implicitly parametrised by the linear recurrence parameters. Crucially, convolutions can be implemented efficiently using subquadratic O(Nlog(N)) fast Fourier transforms (FFTs)35. In sum, linear RNNs have the desirable property of GPU-friendly parallelisability at training time while retaining efficient iterative deployment for inference.Fig. 1 Example Computational Graphs for Sequence Models. (a) shows how basic RNNs perform computations over time. Of note is the inclusion of nonlinearities between time steps, which entail iterative computations. In addition, one can observe how, during the backward pass using BPTT, credit assignment between time steps ∂hp∂hq, where q<<p, involves numerous repeated multiplications which can cause vanishing or exploding gradients. (b) highlights the structural similarities between SNNs and RNNs. One important difference stems from the addition of a linear recurrence based on leaky membrane voltages in neurons such as Leaky Integrate-and-Fire neurons in SNNs. Moreover, the defining feature of SNNs is the neuron outputs consisting of sparse binary spike trains. (c) underlines the parallel nature of Transformers, where input history is no longer compressed within an evolving network state. The attention matrix containing all pair-wise similarities between tokens in the input sequence is multiplied with the V projection of the inputs in dense and large-scale matrix-matrix multiplication, which is unfavourable for neuromorphic hardware implementation. (d) illustrates the dual interpretation of recurrences in linear time-invariant SSMs. In architectures such as S4 Gu et al.36, individual SSM units are single-input single-output (SISO). The scalar input (it) is projected onto high-dimensional space using B∈Rd at each time step. The state of the model (ut) evolves over time using the transition matrix A∈Rd×d. SSM-based neural networks use the initialisation of A and B to implicitly encode projections of input signals onto an orthogonal polynomial basis. To produce a scalar output (yt), the state vector is linearly projected back onto a single dimension using a vector C∈Rd. (a) Unrolled Computations in Basic RNNs, (b) Unrolled Computations in SNNs, (c) Self-Attention, (d) Unrolled Computations in SSMs.

Structured state space models (S4), introduced by Gu et al.36, generalise the parallelisable LMU by exploring the initialisation of recurrent weights based on alternative orthogonal polynomial bases37. This enables input history compression with biases different from sliding windows (e.g., exponentially decaying). Consequently, the stable long-range features learned through S4 orthogonal polynomial initialisations overcome the limitations of RNNs, with strong approximation theory-based guarantees regarding history compression for arbitrarily long sequence lengths and avoiding vanishing/exploding gradients (Sections “State space models” and “State space initialisation”). S4 established the state space model (SSM) class of neural architectures as state-of-the-art methods on several long-range sequence modelling tasks. For example, it outperformed the Transformer in terms of accuracy by an average of 30% on the tasks of the challenging Long Range Arena (LRA) benchmark38. Nevertheless, as highlighted by Fig. 1d), computing the kernel for the global convolutions entails raising the transition matrix (A∈Rd×d) to high powers, which can become slow for large values of d. Gu et al.36 overcome this using efficient multiplication of low-rank approximations of A. Subsequent works have further simplified this process by establishing almost equally effective diagonal initialisation schemes for A∈Cd12,39,40. Diagonal transition matrices are also better suited for deployment to neuromorphic hardware since they allow for iterative state propagation based on Hadamard products rather than dense vector-matrix, thus requiring fewer MAC operations. This also entails state variables evolving independently over time, with an implementation reminiscent of exponentially decaying synapses and membrane voltages in spiking neurons, topics present in neuromorphic hardware research41.

It should be mentioned that the quadratic self-attention complexity of Transformers has also motivated numerous non-recurrent alternatives42. Some strategies reduce the computational cost through sparse or local attention patterns43, such that each query is only compared to a subset of keys (e.g. sliding windows). Others apply hierarchical attention patterns, whereby tokens attend differently to elements in their local neighbourhood and those over longer distances (e.g. word-level attention locally, sentence level globally)44. For long contexts, some architectures reintroduce recurrent computational elements alongside attention, such as averaging over multiple global timescales similar to SSMs45, segmenting input sequences and linking the segments through shared recurrent internal memory caches46, or even providing an explicit external memory bank for long term information47. Another research direction has also been to train models on shorter sequences for memory efficiency and develop extrapolation positional embeddings that enable the Transformers to generalise more robustly to longer inputs at inference48. While relevant as more efficient sequence modelling solutions, the methods mentioned in this paragraph still employ parallel self-attention operations, which, as previously discussed, are antithetical to the event-driven, fully recurrent and temporal computational paradigm of the brain. In terms of performance, most Transformer variants tested in the literature on the LRA have not been able to reach SSM-level accuracy (Table 5, Appendix A), with the notable exception of MEGA45, which employs an SSM-like multi-scale exponential moving average layer along with self-attention.

Related work

Inspired by the success of Transformers, some works have integrated attention into SNNs. A notable example is Temporal-wise Attention SNNs (TA-SNNs)49, which employs attention between input spatiotemporal features to scale them by their salience and relevance to the task before passing them to subsequent spiking neurons. While successful in improving accuracy, this method requires the inherently parallel computation of attention between all sequence frames, effectively using information from the future to process data from the present. Because the focus here is on sequential models with O(1) inference computations at each time step, TA-SNN is beyond the scope of this study.

The renewed interest in RNNs has also inspired works applying these new techniques to SNNs. Some investigate stacking state-of-the-art ANN layers and well-studied neuromorphic Leaky Integrate-and-Fire (LIF) neurons (Eq. 1) – a leading example of this line of research being SpikeGPT28. The authors present the largest SNN language model to date, constructed by feeding outputs from RWKV layers into LIF neurons, which enable sparse spike-based feature mixing. While the RWKV layers could be parallelised as convolutions, the inclusion of LIF neurons imposes iterative computations during training, as highlighted by Fig. 1b. Another example of this research direction is SpikeS450, where LIF neurons are stacked onto S4 layers.

Other works have focused on parallelising the LIF neuron itself. For instance, Fang et al.51 present leaky integration strategies based either on multiplying the entire length-N input sequences by N×N positional encoding matrices in a similar fashion to self-attention matrix multiplication (Fig. 1c) or linearly integrating over an explicit buffer containing sliding windows of the input. Binary spiking is then applied in a position-wise manner. Crucially, neither strategy is formulated for O(1) iterative deployment. Yarga and Wood11 bring SNNs closer to SSMs by exploring both iterative and parallel computations for linear recurrences. However, compared to the linear memory unit of the LMU and other SSMs, both Fang et al.51 and Yarga and Wood11 focus on neurons with scalar internal states. High-dimensional internal states in SSMs constitute linear relationships between input tokens. For single-input single-output (SISO) SSMs with a d-dimensional internal state (u) as used in S4, the same linear relationships can be computed through a convolution of the scalar input signal with a scalar global kernel (Section “State space models”). Crucially, this means the d-dimensional states u do not have to be explicitly stored during training, reducing memory requirements36. If, instead, a nonlinearity is applied position-wise to each of the d dimensions of the state, then they have to be explicitly materialised. The spiking neurons with scalar states in Yarga and Wood11 and Fang et al.51 manifest this structural pitfall, which may prohibit scaling these methods for challenging long-range tasks or using them to build large pre-trained architectures35. Moreover, the initialisation of the decay factor in Yarga and Wood11 and Fang et al.51 is constant between all neurons, which hinders learning dependencies across varying time scales12,52.

Hence, one can notice a significant gap in research at the intersection of state-of-the-art SSMs and SNNs. To the authors’ knowledge, SNNs that borrow powerful initialisation and parameterisation techniques from SSM architectures, such as S4, while retaining their parallelisability, have not been studied so far. Consequently, this paper employs SSM-based SNNs to investigate whether SNNs can eventually become viable energy-efficient alternatives to state-of-the-art ANNs for challenging long-range sequence modelling tasks. Two core questions are addressed in this regard: (a) Do binary spiking activations inherently prevent SNNs from competing with ANNs on long-range sequence modelling? (b) In case they fundamentally hinder performance, should binary spikes necessarily define SNNs?Fig. 2 Binary SSM Layer. At each time step, a Binary SSM layer consists of independent single-input single-output (SISO) SSM “neurons”. Binary activations are applied element-wise per each SSM output before position-wise feature mixing to avoid dense vector-matrix multiplication.

To answer (a), this paper formulates Binary SSMs as SSM-based SNNs (Fig. 2). The models are implemented using S4D initialisation39, and the performance of the resulting Binary S4D (Section “Binary S4D”) is evaluated. Answering (b) requires challenging the role of binary activations in SNNs, which is mainly to avoid MAC operations for feature mixing. Conversely, this approach assumes that mixing continuous features is synonymous with relying on MAC operations. The Gated Spiking Unit (GSU) is formulated here for the first time in order to challenge this assumption (see Section “Gated spiking unit” for further details). The GSU is a position-wise feature mixing layer inspired by the Gated Linear Unit (GLU)53 based on two parallel streams. Continuous SSM features ∈R are mixed using ternary weights ∈{-1,0,1}54, while ternarised SSM outputs are mixed using a continuous-valued linear layer. The final output of the GSU is the Hadamard product of the feature vectors resulting from the two streams. Both streams require only inexpensive additions/subtractions, avoiding MAC operations. Most importantly, as opposed to binarisation in traditional SNNs, the GSU avoids vanishing gradients by allowing backpropagation through non-saturating activations55. GLU-inspired layers in SNNs have been studied before, but only in the context of mixing binary spike features with continuous weights28.

The remainder of the paper is structured as follows. First, Binary SSMs (Section “Binary S4D”) are compared to the GSU (Section “Gated spiking unit”) and baseline state-of-the-art ANN models on the Long Range Arena benchmark38 (LRA). This is the first time SNNs are comprehensively and systematically studied on significantly longer sequences than standard neuromorphic datasets. Second, Binary SSMs and the GSU are compared to, and shown to outperform, current state-of-the-art SNNs on sequential MNIST (sMNIST) classification56, under similar constraints (Section “Sequential MNIST experimental configuration”). Third, the effect of the surrogate gradient function on classification accuracy is highlighted for sequential CIFAR10 (sCIFAR10). Finally, the most difficult long-range modelling task in the LRA, Path-X, is used to compare binary activations and the GSU with continuous-valued saturating activation functions (arctan, fast sigmoid). This highlights that non-differentiable binary activations are upper-bounded in accuracy by continuous-valued saturating activations, which themselves lag far behind non-saturating activations in deep SSM models.

Results

Fig. 3 Input Scales. (a) shows relative sizes of samples from (left to right) MNIST, CIFAR10 and Path-X, with respective resolutions of 28 × 28 (784), 32 × 32 (1024), and 128 × 128 (16384). (b) shows the flattening process used in all image-based sequential tasks (adapted from Bellec et al.57). (c) visualises how the lengths of the flattened image samples compare. One can easily observe from a to c that Path-X contains input sequences more than twenty times longer than sequential MNIST, commonly used for probing SNN long-range dependencies. For further details of the LRA tasks, the reader is referred to the beginning of Section “Results”. (a) Image Samples (To Scale), (b) Image Flattening, (c) Flattened Image Sample Length.

The selection of evaluation tasks is guided by the need to compare the proposed architectures with state-of-the-art in both neuromorphic and broader sequence modelling research. The neuromorphic community has widely embraced variants of the MNIST dataset as standard benchmarks21, therefore its sequential variant (sMNIST) is employed here. sMNIST consists of flattening the 28x28 MNIST samples to 784-long sequences by appending one pixel at a time to a scalar list, row-by-row (Fig. 3b).

The Long Range Arena (LRA)38 is the de facto standard benchmark for measuring long-range dependency modelling capabilities in state-of-the-art architectures39,45,58,59 and it consists of six tasks. Long ListOps, first introduced by Nangia and Bowman60, requires capturing latent hierarchies by parsing nested operations, forming sequences of 2k elements. The Text task is built around the binary classification of byte-level (character-level) IMDB reviews, with sequence lengths fixed at 4k tokens. Retrieval measures how well models can compress byte-level document information to classify two documents’ mutual similarity. Each sample consists of two concatenated documents totalling 8k input sequence elements. The Image task is constructed similarly to sMNIST, flattening out CIFAR10 images61 into 1024-long sequences of gray-scale-valued pixels for classification (Fig. 3c). Finally, Pathfinder and Path-X follow the same aforementioned flattening procedure for the binary classification of images in which two points are either connected or not by a dotted path (rightmost sample in Fig. 3a). While the baseline Pathfinder task consists of images with resolutions of 32 × 32 (sequence length of 1024), Path-X employs samples with resolutions of 128 × 128, resulting in sequences of more than 16k elements (Fig. 3c). To maintain a fair comparison between the developed models and existing architectures, such as the baseline S4D and the Transformer, experimental conditions largely follow the methodology in Gu et al.39.Fig. 4 Accuracy on the LRA Benchmark. Binary S4D performs on average more than 10% worse than the baseline but still over 20% better than the Transformer. The GSU achieves 1.06% lower accuracy than the baseline on average, and at most just 2.83% below the baseline on Image. On Path-X, Binary S4D has 30% lower accuracy than the baseline yet still manages to outperform the Transformer by 11.2%.

LRA accuracy

The S4 model36, and subsequent variants39, have established themselves among state-of-the-art solutions on the LRA benchmark. For example, the original S4 model outperforms the Transformer on the LRA tasks by an average of more than 30% in accuracy. Most importantly, on the longest and most challenging task, Path-X, S4 reaches 88.1%, while the Transformer fails to converge beyond random accuracy (50%). The baseline employed here, an S4D-Lin model (see Sections “State space initialisation” and “LRA experimental setup”), achieves an even higher accuracy of 92.5% on Path-X.

The Binary S4D model (Section “Binary S4D”), proposed in this paper, is trained and evaluated on all tasks within the LRA, to explore how binarisation impacts baseline performance. Figure 4, shows that Binary S4D lags behind the baseline in all tasks of the LRA. On the Image, Listops, Retrieval, and Text tasks, binary spikes impose at most a 6% accuracy penalty. Accuracy is more strongly degraded on Pathfinder and Path-X, where Binary S4D achieves 82.6% and 61.2%, respectively, compared to 91.7% and 92.5% for the baseline. Nonetheless, Binary S4D outperforms the Transformer on all tasks of the LRA, by an average margin of more than 20%. Even where Binary S4D accuracy is significantly lower than the baseline, on Path-X, it still outperforms the Transformer by 11.2%. As such, one can argue that SSM-based SNNs, such as Binary S4D, have stronger long-range modelling capabilities than basic Transformers, as indicated by the results on the LRA. Moreover, for Path-X, Transformer baseline accuracy is obtained with approximately 600k parameters36,38, while Binary S4D uses less than 200k. Since Binary S4D outperforms the Transformer on all tasks of the LRA, one can reasonably argue that binary spiking does not inherently prevent SNNs from exhibiting competitive performance with respect to other ANN architectures. This result contributes to answering question (a) posed in Section “Introduction”.

The GSU (Section “Gated spiking unit”), is also trained and evaluated on all tasks within the LRA, with performance presented in Fig. 4, to highlight the improvements associated with non-saturating activations. The GSU model lags on average 1.06% behind the baseline S4D-Inv model. While the largest discrepancy is on Image, it is still only 2.83%. Interestingly, on Path-X, the GSU is only 0.8% below the baseline (where Binary S4D dropped more than 30%). Generally, it can be argued that the GSU achieves comparable accuracies to the baseline S4D-Inv model on all tasks of the LRA. Similarly, the GSU outperforms the Transformer on average by over 30% on the LRA.Table 1 Accuracy on sMNIST. Binary S4D and GSU outperform current state-of-the-art SNNs, using fewer parameters.

Model	SSM size	Parallelisable	No. trainable parameters (k)	Accuracy (%)	
Binary S4D	2	Yes	68.9	99.1	
64	Yes	118	99.4	
GSU	2	Yes	37.9	99.2	
64	Yes	85.5	99.4	
SRNN62	N/A	No	156 (estimate)	98.7	
LSNN57	N/A	No	66	97.1	

Sequential MNIST accuracy

The proposed Binary S4D and the GSU are compared to state-of-the-art SNNs using the sMNIST task. The landmark findings of Bellec et al.57 brought SNNs closer to matching LSTM accuracy on the well-documented sMNIST classification task. Yin et al.62 further built upon this result establishing current state-of-the-art SNN accuracy. The results in Table 1, show Binary S4D models outperforming both methods, reaching an accuracy of 99.1%, constituting state-of-the-art accuracy for SNNs to the best of the authors’ knowledge. Expanding the state size for Binary S4D (u∈C64) further improves accuracy to 99.4%. The GSU marginally improves accuracy in the u∈C2 configuration to 99.2% and performs identically for the u∈C64 configuration. Both the GSU and Binary S4D models require fewer parameters to achieve higher accuracy than the SRNN62 while also allowing for parallelisable training. It is worth mentioning that while the SSM state size is two (u∈C2), the parameters of the two dimensions are conjugates, meaning the SSM requires training of only one set of parameters for both dimensions39.Fig. 5 Path-X Saturating Activation Functions versus Binary Spiking with Surrogate Gradients Ablation Study. Applying saturating activation functions to SSM outputs leads to reduced accuracy on Path-X, similar to binary spiking activations.

Table 2 Accuracy on sCIFAR10 (Image from LRA) Surrogate Gradient Ablation.

Model	No. layers	Hidden layer size	Surrogate gradient function	Accuracy (%)	
Binary S4D	4	128	Fast sigmoid	69.62	
4	128	Arctan	79.33	
6	512	Fast sigmoid	69.83	
6	512	Arctan	82.00	
GSU	4	128	Fast sigmoid	80.11	
4	128	Arctan	82.49	
6	512	Fast sigmoid	85.01	
6	512	Arctan	85.0	

Surrogate gradient function ablation

Two different functions are ablated to highlight how sensitive Binary S4D and the GSU are to surrogate gradient choice when different depths and widths of the models are considered as well. Previous research suggests that the choice of surrogate gradient function can impact the accuracy of an SNN, with arctan surrogates generally preferred over others18. Table 2 and Fig. 5, reinforce this observation. Binary S4D trained with arctan surrogate gradients achieves 79.33% and 82.00% on sCIFAR10, in the smaller and larger model configurations, respectively. As one could reasonably expect, increasing the model size improves accuracy. In contrast, when using fast sigmoid surrogate gradients, accuracy falls to 69.62% for the smaller configuration and to 69.83% for the larger one. Therefore, fast sigmoid surrogate gradients cause a more than 10% drop in accuracy, below 70%, regardless of the size of the Binary S4D model. The trends identified on sCIFAR10 are also replicated when analysing the results on Path-X (Fig. 5). Binary S4D with arctan surrogate gradients reaches 61.2% in accuracy, while fast sigmoid gradients cause the accuracy to collapse to near-random (51.7%). Hence, Binary S4D is generally sensitive to the surrogate gradient function used.

When analysing GSU results, the effect of the surrogate gradient choice is greatly diminished. For the smaller configuration on sCIFAR10, adopting arctan boosts accuracy by more than 2%, from 80.11% to 82.49%, in the smaller network size configuration. For the larger configuration, accuracy is essentially unchanged between the two surrogate functions, both approximately equalling 85% (Table 2). For the GSU, training on Path-X is also nearly indistinguishable between the two surrogates, although convergence is slightly faster for arctan surrogates than fast sigmoid (orange and blue curves in Fig. 5).

Baseline saturating activations ablation

Results in Section “Surrogate gradient function ablation” suggest that the choice of surrogate gradient function can impact accuracy, especially for Binary S4D. The evaluation of continuous saturating activations is used in this section to explore the potential performance of Binary S4D on long sequences if, hypothetically, an optimal surrogate gradient method was developed.

When replacing spiking activations in Binary S4D with baseline continuous saturating activations and keeping all other hyperparameters unchanged, accuracy improves to some extent on Path-X. The baseline Arctan + ReLU and Fast Sigmoid + ReLU networks achieve 76.4% compared to 75.44%, respectively (Fig. 5). This contrasts the higher sensibility of Binary S4D to surrogate gradient function selection, where fast sigmoid surrogates failed to converge beyond random selection accuracy. In addition, the GSU outperforms all baseline networks with continuous saturating activations, reaching 91.6% and 91.4% with arctan and fast sigmoid surrogate gradients. This suggests that the inherently saturating behaviour of binary spiking activations significantly influences the degraded accuracy.

Both arctan and fast sigmoid functions have negative outputs for negative inputs and intersect the origin (Section “Surrogate gradients”). Nesting the saturating activations within ReLU, emulates the subthreshold regime of binary spiking activations, where the output would be zero for negative inputs. Consequently, this runs the risk of “dead neurons” (those which always output zero and thus cannot learn)63,64, which may affect model performance. Nonetheless, the baseline Arctan trained model manages to reach 74.12% on Path-X, slightly lower than Arctan + ReLU (pink and purple curves in Fig. 5). This could mean that including ReLU does not produce “dead neurons” to a degree that would damage accuracy, and by extension, binary spiking SSMs may not be heavily affected by this phenomenon either.

In answering question (a) from Section “Introduction”, the results here suggest that binarisation does not render SSMs completely uncompetitive since they can still outperform the Transformer (Section “LRA accuracy”) and state-of-the-art SNNs (Section “Sequential MNIST accuracy”). However, this section provides evidence that it does inherently lower their performance. Regardless of the surrogate gradient function, binarisation is still a saturating activation. Hence, the upper bound of binary spiking accuracy is taken to be that of continuous saturating activations65, which experiments in this section show to be lower than non-saturating counterparts such as the GSU, given all other factors are constant.

Energy cost

One of the most important advantages of SNNs is the energy efficiency brought by replacing dense vector-matrix multiplications relying on MAC operations with more efficient Accumulate (ACC) operations. Using the same methodology as62 and66, one can estimate the energy used by MAC and ACC operations as 3.1pJ and 0.1pJ using 45nm CMOS hardware67. The total number of MAC and ACC operations at a single time step for one layer of LIF neurons (Eqs. 20 and 21), the baseline S4D SSM (Eqs. 22, 23 and 24), Binary S4D (Eqs. 25 to 28), and the GSU are provided in Section “Deriving the energy cost”. A notable result of estimating energy consumption with these methods is that given a hidden layer with nl=96 independent SSMs, nl+1=96, and state sizes of dl=64, as used for Path-X (Section “LRA experimental setup”), it can be estimated that the GSU consumes at most ≈51.86% of the baseline SSM with GLU mixing energy per time step, while incurring a degradation in accuracy of less than 1% on Path-X (Fig. 4). This implies a favourable trade-off between energy cost and performance for the proposed GSU.

Discussion

This work has explored the effect of output binarisation in state-of-the-art SSMs in order to assess the viability of SSM-based SNNs as alternatives to ANN sequence models (question (a) in Section “Introduction”). Results show that binarisation lowers accuracy to some extent compared to baselines, and the degradation is inherent to the saturating nature of binary spiking. Sections “LRA accuracy” and “Baseline saturating activations ablation” highlight how the GSU can overcome the vanishing gradient challenges of binary spikes while retaining efficient addition/subtraction-based feature mixing. This suggests that exclusively binary activations may not be necessary for SNNs, providing an answer to question (b) from Section “Introduction”.

Section “Sequential MNIST Accuracy” helps compare Binary SSMs with traditional SNNs. The reset mechanism in LIF neurons may help keep membrane voltages (preactivations) more closely centred around the threshold when firing. This means that when spikes occur, the gradient with respect to the membrane voltage is more likely to be from the non-saturate regime of the surrogate derivative68, helping mitigate vanishing gradients. In contrast, Binary SSMs lack resetting mechanisms, meaning preactivations may stray further from the firing threshold. Arguably, this may cause binary spiking in SSMs to be more strongly affected by vanishing gradients than in LIF neurons. SSM normalisation strategies could potentially be employed in future work to help avoid this pitfall12. Nevertheless, Section “Sequential MNIST accuracy” shows how the SSM backbone can still help Binary S4D outperform current state-of-the-art SNNs on sMNIST, with fewer parameters.

Given the results of the ablation study in Section “Surrogate gradient function ablation”, one can infer that the choice of surrogate gradient determines, to a certain extent, the accuracy of the Binary SSM. This is best highlighted by the complete failure of Binary S4D to converge on Path-X when using fast sigmoid surrogate gradients, compared to the 61.1% accuracy when using arctan (Fig. 5). Furthermore, the discrepancy between training with surrogate gradients and equivalent continuous activations underlines that there is potential for improving surrogate gradient training. However, the disparity between baseline continuous saturating activations and the GSU (Section “Baseline saturating activations ablation”) highlights the intrinsic limitation that binary spiking activations inherit from saturating counterparts - vanishing gradients55.

Non-saturating activations such as ReLU are known to allow the construction of much deeper models than saturating nonlinearities69, effectively avoiding vanishing gradients. The results in Section “LRA accuracy” reflect this fact. The GSU, which allows the propagation of non-saturating values via ternary weights, outperforms Binary S4D on all tasks of the LRA. Section “Baseline saturating activations ablation”, shows how continuous saturating activation functions are also outperformed by the GSU on Path-X. In addition, the GSU manages to incorporate spiking nonlinearities while retaining comparable accuracy to the baseline S4D model on LRA (Section “LRA accuracy”). These results suggest that the saturating behaviour of binary spiking, rather than its discontinuity, limits SNN performance. The overarching observation is that while there is room for improving surrogate-gradient training for SSM-based SNNs, even an ideal unbiased surrogate would struggle to compete with non-saturating activations. Hence, one could argue that implementing state-of-the-art large-scale SSM architectures on neuromorphic hardware should also include efficient forward propagation of non-saturated values. The proposed GSU shows that this is possible while still only using efficient addition/subtraction-based feature mixing.

Certain neuromorphic platforms, such as Intel’s Loihi, have begun to support integer graded spikes25,70. Therefore, the feasibility of techniques such as the proposed GSU could hinge on developing quantisation and sparsification strategies for the weights and activations of this new class of SNNs. One could also argue that current SNN methodologies already make use of propagating integer values. For example, effective residual connections, employed by state-of-the-art SNNs such as SpikeGPT, rely on spike-addition71. This results in layerwise integer outputs that scale with network depth, which differ from binary spikes72.

The contributions of this paper can be summarised as follows. First, this study formulates SSM-based SNNs and tests SNNs for the first time on the LRA, which contains sequence learning tasks with lengths much larger than traditional benchmarks used in neuromorphic research18. Moreover, for the first time, it is shown that SNNs can outperform Transformers on these established long-range sequence benchmarks. Second, this work demonstrates that SSNs built using SSMs can outperform current state-of-the-art SNNs on Sequential MNIST, while using fewer parameters. Finally, this work provides evidence to suggest that the saturating behaviour of spiking activations, not necessarily their discontinuity, can be considered the main challenge to scaling SNNs for long sequences and larger models. By introducing the GSU, it is further highlighted how this problem can be avoided without using dense vector-matrix multiplications relying on MAC operations.

The significance of this paper’s contributions stems from working towards bringing powerful SSMs to energy-efficient neuromorphic hardware. Recently proposed large language models based on SSMs have shown great potential in rivalling and even outperforming Transformer-based architectures73–75, all while avoiding quadratic computational costs. Binary S4D and the GSU retain to a great extent the desirable properties of SSMs for sequence modelling, as highlighted by outperforming the Transformer on the LRA. This paves the way for deploying SSM-based SNNs to neuromorphic hardware, which could drastically reduce the energy requirements of sequential models. Taking into consideration the efficient scaling of computations with respect to sequence length, SSM-based SNNs could have the potential to replace current solutions such as GPU-deployed GPT476.

SSMs have already been shown to achieve remarkable performance on real-world tasks beyond language modelling. For example, S4 and subsequent SSMs have achieved state-of-art accuracy on the Google Speech Commands classification tasks for speech recognition, as required for on-device voice assistants36,77. S4 has also been applied successfully to speech generation78, and also large-scale video activity recognition on the HMDB-51 dataset79,80. S4D has also been used as a backbone for diffusion models81. Given how close in accuracy the GSU is to the S4 baselines, as shown by the performance on the challenging LRA benchmark, one could expect that the GSU would generalise to real-world datasets and tasks as well and is recognised as an interesting avenue for future work.

The task of training and deploying the proposed models to neuromorphic hardware would be an important avenue for future research. One would have to adapt the models and their training to the inherently noisy computational environment of neuromorphic platforms, adopting strategies for improving noise robustness such as the information bottleneck framework82–84. Local learning rules such as using nonlinear dendritic predictions85 or modern STDP variations compatible with supervised learning86 would be most suitable for training SSM-based SNNs on neuromorphic hardware. Deploying the solutions to edge-computing contexts would also require minimising the memory footprint of the models through rewiring and pruning techniques87

It should be mentioned that preliminary evidence has suggested that SSMs face some limitations. While a significant number of studies has supported the superiority of SSMs in terms of accuracy compared to Transformers on very long tasks such as Path-X39,58,59, research has also suggested that SSMs may perform worse than Transformers on shorter language modelling tasks88. This could be caused by a gap in expressivity compared to self-attention89 or may even reflect the effect of different inductive biases90. The gap in language modelling performance has been shrinking, however, with the introduction of selective state space models75,89. Nevertheless, selective SSMs have not yet been shown to improve accuracy on long-range tasks such as the LRA89, and, therefore, remain outside the scope of the present study, which focuses on learning long-range dependencies. Future work should focus on combining selectivity with SNNs for language modelling as well. This could then lead to studying selective SSM-based SNNs as potential drop-in replacements for Transformer blocks in large language models.

Methods

Leaky integrate-and-fire neurons

Spiking networks are most commonly built using Leaky Integrate-and-Fire (LIF) neurons18. They consist of discretising a simplified RC circuit dynamical system (Eq. 1)91. Input currents (it∈R) are linearly accumulated within the membrane voltage (ut∈R) of the neuron (Eq. 2). Current leakage refers to the exponential decay of inputs over time, controlled by the time constant τ∈R and its discrete-time equivalent (β∈R) (Eq. 3). Thus, the membrane voltage contains at any moment useful information about input spike patterns and has been used for downstream tasks such as reconstructing images from even-based data by models even beyond LIF neurons or SNNs92. Once the membrane potential crosses the firing threshold (θ∈R), a spike s is emitted (Eq. 4). As spikes are discrete events highly localised in time, they can be represented by either presence or absence, i.e. binary values s∈{0,1}. Firing is followed by a refractory period when spiking is more difficult. This is implemented using feedback connections, whereby spiking causes the membrane voltage to be either set to a reset value or the threshold value θ is subtracted from the membrane potential (Eq. 3). This reset mechanism imposes iterative computations at training time, much like nonlinearities in RNNs (Fig. 1a). Removing feedback connections converts LIF neurons into LTI filters, which can be implemented as parallelisable convolutions. Equation 5, shows this convolutional view in continuous time, with κ being the global kernel implicitly parametrised by β.1 τdu(t)dt=-u(t)+iR

2 u[t]=βu[t-1]+(1-β)i[t]

3 u[t]=u[t],s[t-1]=0u[t]-θ,s[t-1]=1β=e-Δtτ

4 s[t]=1,u[t]>θ0,u[t]≤θ

5 u(t)=∫0∞κ(s)i(t-s)dsκ(t)=βt∗(1-β)

State space models

State space models (SSMs) are widely used tools in fields such as engineering and neuroscience36. Borrowing LIF concepts, SSMs can be understood as first projecting one-dimensional input currents (i(t)∈R) onto higher dimensions, using a vector B∈Rd, and the result is added to the state of the model (u(t)∈Cd) (Eq. 6). The state is then propagated forward in time using a transition matrix A. For the purposes of this work, A∈Cd is taken to be diagonal. Complex values are required by A and also passed on to the state u in order to retain the same level of expressivity as most full-rank dxd real matrices39. The state ut is then projected back to scalar output values (yt∈R) using C∈Rd and taking the real part of the product (Eq. 6). SSM parameters A and B in architectures such as S436 are typically parameterised in continuous time, and a discretisation scheme is required. All experiments reported in this work are conducted using bilinear discretisation, following Gu et al.36,39 (Eq. 7). The time step size parameter Δ in Eq. 7 also plays the role of determining how quickly the kernel decays over time, setting its time scale37,39. I in Eq. 7, represents the d×d identity matrix. At training time, the discretised parameters A¯ and B¯, along with C, can be used to precompute the global convolution kernel (K¯) (Eq. 9) for each batch. The convolutional theorem (Eq. 10), states that element-wise multiplication (⊙) in the Fourier domain is equivalent to convolution in the time domain. This means the computational cost is dominated by the Fourier transformation (F(.)), and its inverse (F(.)-1), which can be computed efficiently in discrete settings using Fast Fourier Transforms (FFTs) in O(Llog(L)) time, for sequence length L.6 u′(t)=Au(t)+Bi(t)y(t)=Cu(t)+Di(t)

7 A¯=(I-Δ/2A)-1(I+Δ/2A)B¯=(I-Δ/2A)-1·ΔB

8 u[t]=A¯u[t-1]+B¯i[t]y[t]=Cu[t]+Di[t]

9 y[t]=Σp=0tK¯[p]·i[t-p]K¯[p]=CA¯pB¯

10 K¯∗i=F-1(F(K¯)⊙F(i))

State space initialisation

SSM memory properties are deeply influenced by the choice of initialisation for A and B. The eigenvalues of A determine the asymptotic behaviour of Ap required in computing K¯ (Eq. 9), as p→∞. Since A is taken to be diagonal, the eigenvalues (λn) are just its entries (An), which have been established above as being complex. SSMs parametrise the real and imaginary parts of these eigenvalues to encode an orthogonal basis. To ensure long-term stability and avoid exponential growth for large p, the real parts need to be negative (Re(An)<0). Gu et al.39 present -1-122 to be an optimal choice for initialising the real part. During training, to ensure that the real part remains negative, it is typically enclosed within an exponential function (-eln(Re(An)))12,39. The imaginary parts of the eigenvalues Im(An) determine the spectral distribution of the basis and thus the expressivity of the global kernels K¯ they span. To avoid confusion with the input (i), the imaginary unit in Eqs. 11 and 12 is denoted by j. Gu et al.39 propose several initialisation strategies for Im(An), for example, S4D-Lin employs linearly-spaced Im(An) (Eq. 11). S4D-Lin implements a damped Fourier basis, which, notably, has been examined in neuromorphic research before, e.g. within Resonate-and-Fire neurons70 and resonator reservoirs52. S4D-Inv is another proposed initialisation scheme, where Im(An) are distributed by an inverse law (Eq. 12). S4D-Inv has been shown to outperform S4D-Lin on the LRA, especially Path-X39, therefore all models in this work are based on the S4D-Inv initialisation scheme.11 An=-12+jπn

12 An=-12+jdπd2n+1-1

Binary S4D

As highlighted in Fig. 2, the Binary SSM models examined in this work are built by applying the spiking function from LIF neurons, without the reset, to the scalar outputs (y[t]) of each independent SSM (Eq. 13). More precisely, using the schematic structure from Fig. 2, spikes are introduced between the outputs of each SISO SSM “neuron” and the position-wise feature mixing layer, as also highlighted by Eqs. 22, 27, and 25. In all experiments reported here, the firing threshold (θ) is set to zero. The baseline SSMs being binarised are parametrised using the S4D-Inv scheme. Hence, Binary SSM models are referred to as Binary S4D throughout. The binary spikes ensure that feature mixing between different SSM channels does not require dense vector-matrix multiplication. However, it should be mentioned that some additional MAC operations are present in Binary S4D compared to LIF neurons. Namely, one can notice that integrating inputs in high-dimensional states u is more expensive than scalar membrane voltages. Moreover, the dimensionality reduction step Cu[t-1] in Eq. 6 also requires additional MAC operations compared to LIF neurons. These added operations may increase energy costs over traditional LIF neurons. However, the focus here is on the effect on the accuracy of binary spiking activations. Energy-efficient neuromorphic implementations of Binary S4D can be reserved for future work. For example, Voelker et al.33 propose using population spike probability to implement SSM state operations at inference.13 s(y[t])=1,y[t]>θ0,y[t]≤θ

Surrogate gradients

Fig. 6 Fast Sigmoid and Arctan Gradients Decaying on a Log-Scale. Arctan gradients decay slower than fast sigmoid as x→∞.

To account for the non-differentiability of the binary spike function, surrogate gradients are used in the backward pass of the training process19. The gradients of two functions are adopted here (Fig. 6) - fast sigmoid (Eq. 14) and arctangent (Eq. 15). The hyperparameter α∈R in Eq. 14, is set to 25, following the defaults of the snnTorch library18. For the baseline saturating activations with continuous values tested on Path-X, each function is nested in a ReLU activation. As argued in Section “Baseline saturating activations ablation”, this is done to emulate the subthreshold behaviour of binary spiking activations. The arctan activation is also tested without ReLU nesting to check whether saturating at zero may result in “dead neurons” that negatively impact training.14 σ(x)=x1+abs(x)∗αdσdx=1(α∗abs(x)+1)2

15 σ(x)=1π∗arctan(π∗x)dσdx=11+(πx)2

16 σ=ReLU(FastSigmoid(x))orReLU(arctan(x))

Gated spiking unit

Vanishing gradients arise in Binary S4D when constructing deep models since gradients may be greatly attenuated for early layers after passing through several saturating nonlinearities (binary spiking activations)55. This can be mitigated to some extent by using residual connections between layers71,72. However, even with residual connections, issues remain with gradient backpropagation to internal SSM parameters (A, B, C, D) since they can only flow through the saturating bottlenecks (binary spiking activations). The Gated Spiking Unit (GSU) is intended to serve as an example solution to avoid this bottleneck without introducing additional MAC operations.

The GSU is inspired by the Gated Linear Unit (GLU)53 and is intended as a drop-in replacement for the GLU as a position-wise feature mixing layer (Fig. 2). GLU (Eq. 17) passes inputs (x∈Rd) through two linear projections in parallel, resulting in two feature vectors. A sigmoid nonlinearity (σ) is then applied to one of the vectors. The output of the GLU layer is the Hadamard product (⊙) of the two feature vectors, the sigmoid output acting as a scaling factor for the linear projection. Similarly, the GSU mixes input features (x∈Rd) via two parallel routes (Eq. 18). First, each feature of x is ternarised, i.e. continuous values xi are converted to values in {-1,0,1} following a thresholding method adapted from Zhu et al.54 (Eq. 19). The parameter α∈R is typically set to 0.15 by default and controls how sparse the nonzero features are. The ternary features Ter(x) are then linearly projected using weights W∈Rd×k and biases b∈Rk. Simultaneously, x∈Rd is also multiplied by the ternarised weights Ter(W)∈{-1,0,1}d×k and biases c∈Rk are added (the weights W are shared between the two streams). Ternarising W works by computing the maximum function in ΔW and iterating Wij over both dimensions of the weight matrix in Eq. 19. The output of the GSU is the Hadamard product of the two streams. One can observe that both matrix operations, Ter(x)∗W and x∗Ter(W), can be implemented using additions/subtractions, which are efficient mask operations29. It can also be noted that gradients can flow to both x and W via non-saturating routes, avoiding vanishing problems. In all experiments in this paper where it is present, the GSU layer is also followed by layer normalisation and Gaussian Error Linear Unit (GELU) activations93.17 GLU(x)=(x∗W+b)⊙σ(x∗V+c)

18 GSU(x)=(Ter(x)∗W+b)⊙(x∗Ter(W)+c)

19 Ter(xi)=xi=1,xi>=Δxxi=-1,xi<=-Δxxi=0,otherwiseΔx=α∗max(abs(x))

Table 3 LRA Experimental Configuration. WD refers to weight decay and LR to learning rate.

Task	No. layers	No. features	Dropout	LR	Batch size	Epochs	WD	Norm	Pre-norm	(Δtmin, Δtmax)	
ListOps	8	128	0	0.01	50	40	0.05	BN	False	(0.001, 0.1)	
Text	6	256	0	0.01	16	32	0.05	BN	True	(0.001, 0.1)	
Retrieval	6	256	0	0.01	32	11	0.05	BN	True	(0.001, 0.1)	
Image	6	512	0.1	0.01	50	200	0.05	LN	False	(0.001, 0.1)	
Pathfinder	4	92	0	0.004	64	200	0.03	BN	True	(0.001, 0.1)	
Path-X	4	92	0	0.0005	32	50	0.05	BN	True	(0.0001, 0.1)	
BN signifies batch normalisation and LN layer normalisation.

Table 4 Baseline S4D Accuracy on Pathfinder and Path-X. Because of the memory constraints of training on a single Nvidia A100 GPU, the sizes of the Binary S4D and GSU models used for Pathfinder and Path-X had to be reduced from the ones used by Gu et al.39. To provide a more accurate comparison in Sections “LRA accuracy” and “Baseline saturating activations ablation”, smaller baseline models are evaluated here. This table highlights how the baseline S4D with GELU activations used in this paper compare to the larger models employed by39. Besides the number of layers and features per layer, all other hyperparameters for the large models used by39 are identical to the ones used here (Table 3).

Configuration	Activation function	No. layers	No. features	Pathfinder Acc. (%)	Path-X Acc. (%)	
Small (this work)	GELU	4	92	91.7	92.5	
Large39	GELU	6	256	93.78	92.80	

Deriving the energy cost

One of the driving forces behind neuromorphic research is reducing computational power consumption. As previously mentioned (Section “Introduction”), reducing the number of MAC operations is greatly beneficial in this regard, as multiplications are an order of magnitude more energy-intensive compared to additions94. To obtain the energy consumption estimates in Section “Energy cost”, Eqs. 20 to 30 are used. The main goal of this section is to show how binary spiking and efficient feature mixing strategies, such as the GSU, help offset the additional MAC operations introduced by handling the high-dimensional internal states within individual SSMs (Fig. 2).

A widely adopted neuromorphic solution is Recurrent Spiking Neural Networks (RSNNs)57,62,95, which endow the LIF neuron (Section “Leaky integrate-and-fire neurons”) with recurrent weights akin to RNNs, besides the feedback connection of the reset mechanism (Eq. 3). Equation 20 illustrates this mechanism through the formulation of an RSNN layer computation at a given time step t. At a layer l with nl LIF neurons, the membrane voltage vector Ul[t]∈Rln decays and integrates the incoming input vector Il[t]∈Rln, according to Eq. 2. The input comprises the sum of the output from the previous layer Ol-1[t]∈Rln and the spikes of the current layer from the previous time step mixed using a recurrent weight matrix Wrecl∈Rnl×nl. The resulting spikes Sl[t]∈{0,1}]ln (Eq. 4) are then mixed with the output weights Woutl∈Rnl×nl+1 and sent to the next layer. Equation 21 shows the total energy cost of the operations of an RSNN layer at a single time step, following the analyses of62 and96. MAC operations are only required for decaying the membrane voltages and integrating the incoming inputs, while all feature mixing with Wrecl and Woutl can be implemented with efficient Accumulate (ACC) operations.

Equation 22 shows the operations of an SSM layer of nl independent neurons (Eq. 8) at a given time step. It is important to highlight that, as shown in Fig. 2, each SSM neuron p has an internal state vector Upl∈Rdl, as opposed to the scalar membrane voltage of LIF neurons. Consequently, SSM neurons require more MAC operations than LIF neurons (Eq. 24) to integrate incoming inputs, which need to be expanded from nl scalar values to dl-sized vectors for each neuron. In addition, decaying the internal states of each neuron requires dl MACs for the Hadamard product with the diagonal transition matrix A¯ (Sections “State space models” and “State space initialisation”). Reducing the internal states to scalars requires an additional dl per each of the nl neurons due to the dot product with C. Taking into consideration the GLU feature mixing to produce the layer output vector (Eqs. 17 and 23), the total energy cost of the SSM operations increases meaningfully compared to LIF neurons (Eq. 24). The feature mixing in the Binary S4D models employed throughout this study reduces the number of MAC operations by 2nlnl+1 (Eq. 26) by introducing spiking activations before the GLU mixing (Section “LRA experimental setup”, Eq. 25). Importantly, the GSU retains the same energy cost as Binary S4D with GLU (Eqs. 29, 30) while significantly improving performance (Section “LRA accuracy”). Energy costs could be further improved marginally by appending a linear mixing layer instead of GLU to mix the spikes of Binary S4D (Eqs. 27, 28)20 Ul[t]=βUl[t]+(1-β)Il[t]Ol[t]=WoutlS[t]Il[t]=WreclS[t-1]+Ol-1[t]

21 ETotalLIF=2nlEMAC+(nl+1nl+nlnl)EACC

22 Upl[t]=A¯l⊙Upl[t-1]+B¯Ip[t]Ypl[t]=C·Upl[t]+DplIpl[t]Ipl[t]=Opl-1[t]

23 Ol[t]=σ(WlYl[t]+bl)⊙(VlYl[t]+cl)

24 ETotalSSM=(2nlnl+1+3nldl+nl+nl+1)EMAC

25 Ol[t]=σ(WlSl[t]+bl)⊙(VlSl[t]+cl)

26 ETotalBinarySSM=(3nldl+nl+nl+1)EMAC+2nlnl+1EACC

27 Ol[t]=WlSl[t]+bl

28 ETotalBinarySSM=(3nldl+nl)EMAC+nlnl+1EACC

29 Ol[t]=(Ter(Yl[t])∗Wl+bl)⊙(Yl[t]∗Ter(Wl)+cl)

30 ETotalGSU=(3nldl+nl+nl+1)EMAC+2nlnl+1EACC

LRA experimental setup

Overall, evaluation of the proposed methods on the LRA benchmark closely followed the experiments conducted in Gu et al.39 to ensure that the object of the investigation is only the binary spiking activation (Table 3). More precisely, Binary S4D and the GSU are evaluated on ListOps, Text and Image using identical hyperparameters to Gu et al.39. Retrieval is evaluated using batch sizes reduced from 64 to 32 and fewer epochs (eleven compared to the original twenty). This is done to reduce training time and accommodate memory constraints on a single Nvidia A100 GPU. In addition, Pathfinder and Path-X are evaluated using smaller models than employed by Gu et al.39 for the same reason. The baseline S4D-Inv results for Pathfinder and Path-X (Fig. 4) replicated here are obtained by replacing the binary spiking activation in Binary S4D with GELU activations and GLU feature mixing. In both Pathfinder and Path-X, Gu et al.39 use six layers with 256 features each, compared to four layers with 92 features used here, reporting accuracies of 93.78% and 92.80%, respectively (Table 4). Higher accuracies for the baseline S4D-Inv in the smaller configuration might have been possible if further hyperparameter tuning were employed. However, the goal here is only to isolate the effect of including binary spiking activation, all other hyperparameters being equal, such as the number of layers, SSM state size (u∈Cd), number of epochs, etc. Models trained on all tasks of the LRA have SSM state u with dim(u)=64.

To maintain comparability with the S4D-Inv baselines, Binary S4D models use GLU layers for position-wise feature mixing after applying binary spiking activations. In addition, both Binary S4D and GSU networks are bidirectional. Finally, all models are implemented using addition-based residual connections. For Binary S4D, to ensure that only binary values are passed to the feature mixing layers, the residual addition takes place after feature mixing. The results in Fig. 4, are obtained using arctan surrogates for both Binary S4D and the GSU in all tasks except Retrieval, where fast sigmoid surrogates are used.

The reported Transformer baseline accuracies on the LRA are extracted from the original study proposing the benchmark38. Consequently, they have been used as baseline Transformer accuracies in numerous major studies incorporating the LRA36,39,45,58,59,89. The hyperparameters used by Tay et al.38 are included in Table 6 in Appendix B. Accuracy information on the LRA for additional models besides the vanilla Transformer is included in Table 5 in Appendix A.

Model validation follows the same methodology as Gu et al.39. For the Text task, the train-test split is static, following the original selection from Maas et al.97, with 25,000 samples in each set and no validation set. Retrieval is also based on a static sample partition, with 147,086 samples in the train set, 18,090 in the validation set and 17,437 in the test set. Image consists of a static partition between train and test sets, with the train set including 50,000 samples and the test set 10,000 samples. A validation set is constructed at random for each experiment, with 2% of the training samples. Pathfinder and Path-X consist of 200,000 samples in total each. The validation and test sets each consist of 10% of the total samples and are selected at random for each experiment. All experimental results reported in this study are based on the test set classification accuracy.

All code developed for the experiments in this study is built on the S4 publicly available repository https://github.com/state-spaces/s436.

Sequential MNIST experimental configuration

The choice of hyperparameters for sMNIST classification is largely motivated by the intent to closely emulate traditional SNN computational principles. Hence, residual connections are removed in favour of exclusively spike-based communication between recurrent layers. Bidirectionality is also disabled. Models with both state dimensions (dim(u)∈{2,64}) for GSU and Binary S4D are implemented in networks with two layers with 128 features (independent SSMs).

Decoding

Temporal features from the last spiking SSM layer need to be compressed before being processed by the label prediction output layer for all of the classification tasks presented. The output of the final SSM layer can be viewed as a tensor x of shape [B×L×H], where B is the batch size, L input sequence length, and H is the number of hidden features. The output layer requires condensed inputs of shape [B×H]. To reduce the L dimension, average pooling is applied over it (e.g., torch.mean (x, dim = 1) in PyTorch). This effectively entails that the final spiking layer is employing rate-coding98.

Appendix

A: Extended long range arena accuracies

See Table 5.Table 5 Extended Table of LRA Accuracies.

Model	ListOps	Text	Image	Retrieval	Pathfinder	Path-X	Average	
Transformer22	36.37	64.27	57.46	42.44	71.40	FAIL	53.66	
Reformer99	37.27	56.10	53.40	38.07	68.50	FAIL	50.56	
BigBird43	36.05	64.02	59.29	40.83	74.87	FAIL	54.17	
Linear Trans.100	16.13	65.90	53.09	42.34	75.30	FAIL	50.46	
Performer101	18.01	65.40	53.82	42.77	77.05	FAIL	51.18	
FNet102	35.33	65.11	59.61	38.67	77.80	FAIL	54.42	
Nyströmformer103	37.15	65.52	79.56	41.58	70.94	FAIL	57.46	
Luna-25645	37.25	64.57	79.29	47.38	77.72	FAIL	59.37	
H-Transformer-1D104	49.53	78.69	63.99	46.05	68.78	FAIL	61.41	
CCNN105	43.60	84.08	FAIL	88.90	91.51	FAIL	68.02	
Mega (O(L2)))45	63.14	90.43	91.25	90.44	96.01	97.98	88.21	
Mega-chunk (O(L))45	58.76	90.19	90.97	85.80	94.41	93.81	85.66	
S4D-Inv39	60.18	87.34	91.09	87.83	93.78	92.80	85.50	
Liquid-S458	62.75	89.02	91.20	89.50	94.80	96.66	87.32	
S559	62.15	89.31	91.40	88.00	95.33	98.58	87.46	
H389	57.50	88.20	91.00	87.30	93.00	91.80	84.8	
Binary S4D (this study)	54.8	82.50	82.00	85.03	82.60	61.20*	74.69	
GSU (this study)	59.6	86.50	85.00	90.22	91.30	91.70*	84.05	
The table includes the accuracies of notable efficient Transformer architectures and other SSMs on the LRA (expanding on the list from59).

The spiking methods proposed here are competitive with other SSM architectures. Asterisks (*) signify that a smaller configuration was used compared to the rest of the models in the table, due to memory constraints (Section “LRA experimental setup”).

B Transformer baseline hyper-parameters

See Table 6.Table 6 LRA Experimental Configuration.

Task	Emb. Dim.	No. heads	No. layers	QKV Dim.	Positional MLP Dim.	WD	LR	Batch size	
ListOps	512	8	4	512	1024	0.1	0.05	32	
Text	256	4	4	256	1024	0.1	0.05	32	
Retrieval	128	4	4	128	512	0.1	0.05	32	
Image	128	8	1	64	128	0.0	0.0005	256	
Pathfinder	32	4	4	16	32	0.0	0.01	512	
Path-X	32	2	1	16	32	0.0	0.001	64	
WD refers to weight decay, LR to learning rate and Emb to embedding.

The hyper-parameters have been obtained from the LRA code repository: https://github.com/google-research/long-range-arena/.

Acknowledgements

This research is supported in part through the NimbleAI project, which has received funding from the EU’s Horizon Europe Research and Innovation programme (Grant Agreement 101070679), and by the UK Research and Innovation (UKRI) under the UK government’s Horizon Europe funding guarantee (Grant Agreement 10039070). See: https://www.nimbleai.eu.

Author contributions

M.I.S conceived the models under investigation, conducted the experiments and contributed to their design, prepared figures and wrote the main manuscript text. O.R. supervised the research, contributed to devising the experiments and provided major revisions to the final manuscript text and figures. All authors reviewed the final manuscript.

Data availibility

The datasets used in this study are publicly available: Long-Range Arena GitHub repository: https://github.com/google-research/long-range-arena MNIST: http://yann.lecun.com/exdb/mnist/ CIFAR10 and CIFAR 100: https://www.cs.toronto.edu/~kriz/cifar.html.

Competing interests

The authors declare no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Tang, S., Dunnmon, J. A., Liangqiong, Q., Saab, K. K., Baykaner, T., Lee-Messer, C. & Rubin, D. L. Modeling multivariate biosignals with graph neural networks and structured state space models. In Conference on health, inference, and learning, 50–71 (PMLR, 2023).
2. Zhou, W., Jiang, Y. E., Cui, P., Wang, T., Xiao, Z., Hou, Y., Cotterell, R. & Sachan, M. Recurrentgpt: Interactive generation of (arbitrarily) long text. arXiv preprint arXiv:2305.13304 (2023).
3. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F. & Liang, P. Lost in the middle: How language models use long contexts. arXiv preprint arXiv:2307.03172 (2023).
4. Lipton, Z. C., Berkowitz, J. & Elkan, C. A critical review of recurrent neural networks for sequence learning. arXiv preprint arXiv:1506.00019 (2015).
5. Maass W Networks of spiking neurons: the third generation of neural network models Neural Netw. 1997 10 9 1659 1671 10.1016/S0893-6080(97)00011-7
Maass, W. Networks of spiking neurons: the third generation of neural network models. Neural Netw. 10(9), 1659–1671 (1997).
6. Hasler J Special report: Can we copy the brain?-a road map for the artificial brain IEEE Spectr. 2017 54 6 46 50 10.1109/MSPEC.2017.7934231
Hasler, J. Special report: Can we copy the brain?-a road map for the artificial brain. IEEE Spectr. 54(6), 46–50 (2017).
7. McClelland JL McNaughton BL O’Reilly RC Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of connectionist models of learning and memory Psychol. Rev. 1995 102 3 419 10.1037/0033-295X.102.3.419 7624455
McClelland, J. L., McNaughton, B. L. & O’Reilly, R. C. Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of connectionist models of learning and memory. Psychol. Rev. 102(3), 419 (1995).7624455
8. Shen, J., Ni, W., Qi, X. & Tang, H. Efficient spiking neural networks with sparse selective activation for continual learning. In Proceedings of the AAAI Conference on Artificial Intelligence 38, 611–619 (2024).
9. Pascanu, R., Mikolov, T. & Bengio, Y. On the difficulty of training recurrent neural networks. In International conference on machine learning, 1310–1318 (Pmlr, 2013).
10. Hochreiter S Schmidhuber J Long short-term memory Neural Comput. 1997 9 8 1735 1780 10.1162/neco.1997.9.8.1735 9377276
Hochreiter, S. & Schmidhuber, J. Long short-term memory. Neural Comput. 9(8), 1735–1780 (1997).9377276
11. Yarga, S. Y. A. & Sean U. N. W. Accelerating snn training with stochastic parallelizable spiking neurons. In 2023 international joint conference on neural networks (IJCNN), 1–8, (2023). 10.1109/IJCNN54540.2023.10191884.
12. Orvieto, A., Smith, S. L., Gu, A., Fernando, A., Gulcehre, C., Pascanu, R. & De, S. Resurrecting recurrent neural networks for long sequences. arXiv preprint arXiv:2303.06349 (2023).
13. Kalchbrenner, N., Espeholt, L., Simonyan, K., van den Oord, A., Graves, A. & Kavukcuoglu, K. Neural machine translation in linear time. arXiv preprint arXiv:1610.10099 (2016).
14. Diehl, P. U., Neil, D., Binas, J., Cook, M., Liu, S. -C. Pfeiffer, M. Fast-classifying, high-accuracy spiking deep networks through weight and threshold balancing. In 2015 International joint conference on neural networks (IJCNN), 1–8. (IEEE, 2015).
15. Davidson S Furber SB Comparison of artificial and spiking neural networks on digital hardware Front. Neurosci. 2021 15 345 10.3389/fnins.2021.651141
Davidson, S. & Furber, S. B. Comparison of artificial and spiking neural networks on digital hardware. Front. Neurosci. 15, 345 (2021).
16. Garg, I., Chowdhury, S. S. & Roy, K. Dct-snn: Using dct to distribute spatial information over time for learning low-latency spiking neural networks. arXiv preprint arXiv:2010.01795 (2020).
17. Liu, F., Zhao, W., Chen, Y., Wang, Z. & Jiang, L. Spikeconverter: An efficient conversion framework zipping the gap between artificial neural networks and spiking neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence 36, 1692–1701 (2022).
18. Eshraghian, J. K., Ward, M., Neftci, E. O., Wang, X., Lenz, G., Dwivedi, G., Bennamoun, M., Jeong, D. S. & Lu, W. D. Training spiking neural networks using lessons from deep learning. In Proceedings of the IEEE, (2023).
19. Neftci EO Mostafa H Zenke F Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based optimization to spiking neural networks IEEE Signal Process. Mag. 2019 36 6 51 63 10.1109/MSP.2019.2931595
Neftci, E. O., Mostafa, H. & Zenke, F. Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based optimization to spiking neural networks. IEEE Signal Process. Mag. 36(6), 51–63 (2019).
20. Yujie W Deng L Li G Shi L Spatio-temporal backpropagation for training high-performance spiking neural networks Front. Neurosci. 2018 12 323875
Yujie, W., Deng, L., Li, G. & Shi, L. Spatio-temporal backpropagation for training high-performance spiking neural networks. Front. Neurosci. 12, 323875 (2018).
21. Malcom, K. & Casco-Rodriguez, J. A comprehensive review of spiking neural networks: Interpretation, optimization, efficiency, and best practices. arXiv preprint arXiv:2303.10780 (2023).
22. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł. & Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, (2017).
23. Zeyer, A., Bahar, P., Irie, K., Schlüter, R. & Ney, H. A comparison of transformer and lstm encoder decoder models for asr. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 8–15 (IEEE, 2019).
24. Min B Ross H Sulem E Veyseh APB Nguyen TH Sainz O Agirre E Heintz I Roth D Recent advances in natural language processing via large pre-trained language models: A survey ACM Comput. Surv. 2023 56 2 1 40 10.1145/3605943
Min, B. et al. Recent advances in natural language processing via large pre-trained language models: A survey. ACM Comput. Surv. 56(2), 1–40 (2023).
25. Davies M Wild A Orchard G Sandamirskaya Y Guerra GAF Joshi P Plank P Risbud SR Advancing neuromorphic computing with loihi: A survey of results and outlook Proc. IEEE 2021 109 5 911 934 10.1109/JPROC.2021.3067593
Davies, M. et al. Advancing neuromorphic computing with loihi: A survey of results and outlook. Proc. IEEE 109(5), 911–934 (2021).
26. Li, G. et al. Wolfgang Maass (Brain-Inspired Computing: A Systematic Survey and Future Trends, 2023)
27. Zhou, Z., Zhu, Y., He, C., Wang, Y., Yan, S., Tian, Y. & Yuan, L. Spikformer: When spiking neural network meets transformer. arXiv preprint arXiv:2209.15425 (2022).
28. Zhu, R.-J., Zhao, Q. & Eshraghian, J. K. Spikegpt: Generative pre-trained language model with spiking neural networks. arXiv preprint arXiv:2302.13939 (2023).
29. Yao, M., Hu, J., Zhou, Z., Yuan, L., Tian, Y., Xu, B. & Li, G. Spike-driven transformer. arXiv preprint arXiv:2307.01694 (2023).
30. Tay, Y., Dehghani, M., Bahri, D. & Metzler, D. Efficient transformers: A survey. arXiv preprint arXiv:cs.LG/2009.06732, (2020b).
31. Strubell, E., Ganesh, A. & McCallum, A. Energy and policy considerations for deep learning in nlp. arXiv preprint arXiv:1906.02243 (2019).
32. Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., Kranthi Kiran, G.V. et al. Rwkv: Reinventing rnns for the transformer era. arXiv preprint arXiv:2305.13048 (2023).
33. Voelker, A., Kajić, I. & Eliasmith, C. Legendre memory units: Continuous-time representation in recurrent neural networks. Advances in neural information processing systems, 32, (2019).
34. Chilkuri, N. R. & Eliasmith, C. Parallelizing legendre memory unit training. In International conference on machine learning, 1898–1907. (PMLR, 2021).
35. Albert G Johnson I Goel K Saab K Dao T Rudra A Ré C Combining recurrent, convolutional, and continuous-time models with linear state space layers Adv. Neural. Inf. Process. Syst. 2021 34 572 585
Albert, G. et al. Combining recurrent, convolutional, and continuous-time models with linear state space layers. Adv. Neural. Inf. Process. Syst. 34, 572–585 (2021).
36. Gu, A., Goel, K. & Ré, C. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396 (2021a).
37. Albert G Dao T Ermon S Rudra A Ré C Hippo: Recurrent memory with optimal polynomial projections Adv. Neural. Inf. Process. Syst. 2020 33 1474 1487
Albert, G., Dao, T., Ermon, S., Rudra, A. & Ré, C. Hippo: Recurrent memory with optimal polynomial projections. Adv. Neural. Inf. Process. Syst. 33, 1474–1487 (2020).
38. Tay, Y, Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S. & Metzler, D. Long range arena: A benchmark for efficient transformers. arXiv preprint arXiv:2011.04006 (2020a).
39. Albert G Goel K Gupta A Ré C On the parameterization and initialization of diagonal state space models Adv. Neural. Inf. Process. Syst. 2022 35 35971 35983
Albert, G., Goel, K., Gupta, A. & Ré, C. On the parameterization and initialization of diagonal state space models. Adv. Neural. Inf. Process. Syst. 35, 35971–35983 (2022).
40. Gupta A Albert G Berant J Diagonal state spaces are as effective as structured state spaces Adv. Neural. Inf. Process. Syst. 2022 35 22982 22994
Gupta, A., Albert, G. & Berant, J. Diagonal state spaces are as effective as structured state spaces. Adv. Neural. Inf. Process. Syst. 35, 22982–22994 (2022).
41. Eissa, S., Stuijk, S. & Corporaal, H. Hardware approximation of exponential decay for spiking neural networks. In 2021 IEEE 3rd international conference on artificial intelligence circuits and systems (AICAS), 1–4. (IEEE, 2021).
42. Huang, Y., Xu, J., Jiang, Z., Lai, J., Li, Z., Yao, Y., Chen, T., Yang, L., Xin, Z. & Ma, X. Advancing transformer architecture in long-context large language models: A comprehensive survey. arXiv preprint arXiv:2311.12351 (2023).
43. Zaheer M Guruganesh G Avinava Dubey K Ainslie J Alberti C Ontanon S Pham P Ravula A Wang Q Yang L Big bird: Transformers for longer sequences Adv. Neural. Inf. Process. Syst. 2020 33 17283 17297
Zaheer, M. et al. Big bird: Transformers for longer sequences. Adv. Neural. Inf. Process. Syst. 33, 17283–17297 (2020).
44. Yang, Z., Yang, D., Dyer, C., He, X., Smola, A. & Hovy, E. Hierarchical attention networks for document classification. In Proceedings of the 2016 conference of the North American chapter of the association for computational linguistics: human language technologies, 1480–1489 (2016).
45. Ma, X., Zhou, C., Kong, X., He, J., Gui, L., Neubig, G., May, J. & Zettlemoyer, L. Mega: Moving average equipped gated attention. arXiv preprint arXiv:2209.10655 (2022).
46. Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V. & Salakhutdinov, R. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860 (2019).
47. Lewis P Perez E Piktus A Petroni F Karpukhin V Goyal N Küttler H Lewis M Yih W Rocktäschel T Retrieval-augmented generation for knowledge-intensive nlp tasks Adv. Neural. Inf. Process. Syst. 2020 33 9459 9474
Lewis, P. et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Adv. Neural. Inf. Process. Syst. 33, 9459–9474 (2020).
48. Anil C Yuhuai W Andreassen A Lewkowycz A Misra V Ramasesh V Slone A Gur-Ari G Dyer E Neyshabur B Exploring length generalization in large language models Adv. Neural. Inf. Process. Syst. 2022 35 38546 38556
Anil, C. et al. Exploring length generalization in large language models. Adv. Neural. Inf. Process. Syst. 35, 38546–38556 (2022).
49. Yao, M., Gao, H., Zhao, G., Wang, D., Lin, Y., Yang, Z. & Li, G. Temporal-wise attention spiking neural networks for event streams classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 10221–10230 (2021).
50. Du, Y., Liu, X. & Chua, Y. Spiking structured state space model for monaural speech enhancement. arXiv preprint arXiv:2309.03641 (2023).
51. Fang, W., Yu, Z., Zhou, Z., Chen, Y., Ma, Z., Masquelier, T. & Tian, Y. Parallel spiking neurons with high efficiency and long-term dependencies learning ability. arXiv preprint arXiv:2304.12760 (2023).
52. Hermans M Schrauwen B Memory in linear recurrent neural networks in continuous time Neural Netw. 2010 23 3 341 355 10.1016/j.neunet.2009.08.008 19748225
Hermans, M. & Schrauwen, B. Memory in linear recurrent neural networks in continuous time. Neural Netw. 23(3), 341–355 (2010).19748225
53. Dauphin, Y. N., Fan, A., Auli, M. & Grangier, D. Language modeling with gated convolutional networks. Computation and Language, (2016).
54. Zhu, C., Han, S., Mao, H. & Dally, W. J. Trained ternary quantization. arXiv preprint arXiv:1612.01064 (2016).
55. Gulcehre, C., Moczulski, M., Denil, M. & Bengio, Y. Noisy activation functions. In International conference on machine learning, 3059–3068 (PMLR, 2016).
56. Le, Q. V., Jaitly, N. & Hinton, G. E. A simple way to initialize recurrent networks of rectified linear units. arXiv preprint arXiv:1504.00941 (2015).
57. Bellec, G., Salaj, D., Subramoney, A., Legenstein, R. & Maass, W. Long short-term memory and learning-to-learn in networks of spiking neurons. Adv. Neural Inf. Process. Syst. 31, (2018).
58. Hasani, R., Lechner, M., Wang, T.-H., Chahine, M., Amini, A. & Rus, D. Liquid structural state-space models. arXiv preprint arXiv:2209.12951 (2022).
59. Smith, J. T. H., Warrington, A. & Linderman, S. W. Simplified state space layers for sequence modeling. arXiv preprint arXiv:2208.04933 (2022).
60. Nangia, N. & Bowman, S. R. Listops: A diagnostic dataset for latent tree learning. arXiv preprint arXiv:1804.06028 (2018).
61. Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images (2009).
62. Yin B Corradi F Bohté SM Accurate and efficient time-domain classification with adaptive spiking recurrent neural networks Nat. Mach. Intell. 2021 3 10 905 913 10.1038/s42256-021-00397-w
Yin, B., Corradi, F. & Bohté, S. M. Accurate and efficient time-domain classification with adaptive spiking recurrent neural networks. Nat. Mach. Intell. 3(10), 905–913 (2021).
63. Eshraghian, Jason K, & Lu, Wei D. The fine line between dead neurons and sparsity in binarized spiking neural networks. arXiv preprint arXiv:2201.11915 (2022).
64. Douglas, S. C, & Yu, J. Why Relu units sometimes die: analysis of single-unit error backpropagation in neural networks. In 2018 52nd Asilomar conference on signals, systems, and computers, 864–868. (IEEE, 2018).
65. Roberts, D. A., Yaida, S. & Hanin, B. The principles of deep learning theory. Cambridge University Press Cambridge, MA, USA, (2022).
66. Lemaire, E., Cordone, L., Castagnetti, A., Novac, P.-E., Courtois, J. & Miramond, B. An analytical estimation of spiking neural networks energy efficiency. In International Conference on Neural Information Processing, 574–587 (Springer, 2022).
67. Horowitz, M. 1.1 computing’s energy problem (and what we can do about it). In 2014 IEEE international solid-state circuits conference digest of technical papers (ISSCC), 10–14 (IEEE, 2014).
68. Herranz-Celotti, L. & Rouat, J. Surrogate gradients design. arXiv preprint arXiv:2202.00282 (2022).
69. Glorot, X., Bordes, A. & Bengio, Y. Deep sparse rectifier neural networks. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, 315–323 (JMLR Workshop and Conference Proceedings, 2011).
70. Orchard, G., Frady, E. P., Rubin, B. D., Daniel, S., S., Shrestha, S. B., Sommer, F. T. & Davies, M. Efficient neuromorphic signal processing with loihi 2. In 2021 IEEE Workshop on Signal Processing Systems (SiPS), 254–259 (IEEE, 2021).
71. Fang W Zhaofei Yu Chen Y Huang T Masquelier T Tian Y Deep residual learning in spiking neural networks Adv. Neural. Inf. Process. Syst. 2021 34 21056 21069
Fang, W. et al. Deep residual learning in spiking neural networks. Adv. Neural. Inf. Process. Syst. 34, 21056–21069 (2021).
72. Chen, G., Peng, P., Li, G. & Tian, Y. Training full spike neural networks via auxiliary accumulation pathway. arXiv preprint arXiv:2301.11929 (2023).
73. Dao, T., Fu, D. Y, Saab, K. K, Thomas, A. W, Rudra, A. & Ré, C. Hungry hungry hippos: Towards language modeling with state space models. arXiv preprint arXiv:2212.14052 (2022).
74. Poli, M., Massaroli, S., Nguyen, E., Fu, D. Y., Dao, T., Baccus, S., Bengio, Y., Ermon, S. & Ré, C. Hyena hierarchy: Towards larger convolutional language models. arXiv preprint arXiv:2302.10866 (2023).
75. Gu, A., & Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752 (2023).
76. OpenAI, R. Gpt-4 technical report. 2303–08774, (2023).
77. Warden, P. Speech commands: A dataset for limited-vocabulary speech recognition. arXiv preprint arXiv:1804.03209 (2018).
78. Goel, Karan, Gu, Albert, Donahue, Chris, & Ré, Christopher. It’s raw! audio generation with state-space models. In International conference on machine learning, 7616–7633 (PMLR, 2022).
79. Kuehne, H., Jhuang, H., Garrote, E., Poggio, T. & Serre, T. Hmdb: A large video database for human motion recognition. In 2011 International conference on computer vision, 2556–2563 (IEEE, 2011).
80. Nguyen, E., Goel, K., Gu, A., Downs, G. W., Shah, P., Dao, T., Baccus, S. A. & Ré, C. S4nd: Modeling images and videos as multidimensional signals using state spaces. arXiv preprint arXiv:2210.06583 (2022).
81. Yan, J. N., Gu, J. & Rush, A. M. Diffusion models without attention. arXiv preprint arXiv:2311.18257 (2023).
82. Yang, S., Wang, H. & Chen, B. Sibols: robust and energy-efficient learning for spike-based machine intelligence in information bottleneck framework. IEEE Transactions on Cognitive and Developmental Systems, (2023b).
83. Yang, Shuangming, & Chen, Badong. Snib: improving spike-based machine learning using nonlinear information bottleneck. IEEE Transactions on Systems, Man, and Cybernetics: Systems, (2023b).
84. Yang, S. & Chen, B. Effective surrogate gradient learning with high-order information bottleneck for spike-based machine intelligence. IEEE transactions on neural networks and learning systems, (2023a).
85. Yang S Pang Y Wang H Lei T Pan J Wang J Jin Y Spike-driven multi-scale learning with hybrid mechanisms of spiking dendrites Neurocomputing 2023 542 126240 10.1016/j.neucom.2023.126240
Yang, S. et al. Spike-driven multi-scale learning with hybrid mechanisms of spiking dendrites. Neurocomputing 542, 126240 (2023).
86. Liu F Zhao W Chen Y Wang Z Yang T Jiang L Sstdp: Supervised spike timing dependent plasticity for efficient spiking neural network training Front. Neurosci. 2021 15 756876 10.3389/fnins.2021.756876 34803591
Liu, F. et al. Sstdp: Supervised spike timing dependent plasticity for efficient spiking neural network training. Front. Neurosci. 15, 756876 (2021).34803591
87. Shen, J., Xu, Q., Liu, J. K., Wang, Y., Pan, G. & Tang, H. Esl-snns: An evolutionary structure learning strategy for spiking neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, 86–93 (2023).
88. Vardasbi, A., Pires, T. P., Schmidt, R. M. & Peitz, S. State spaces aren’t enough: Machine translation needs attention. In EAMT, (2023). arXiv:2304.12776.
89. Fu, D. Y., Dao, T., Saab, K. K., Thomas, A. W., Rudra, A. & Ré, C. Hungry hungry hippos: Towards language modeling with state space models. arXiv preprint arXiv:2212.14052 (2022).
90. Atz K Grisoni F Schneider G Geometric deep learning on molecular representations Nat. Mach. Intell. 2021 3 12 1023 1032 10.1038/s42256-021-00418-8
Atz, K., Grisoni, F. & Schneider, G. Geometric deep learning on molecular representations. Nat. Mach. Intell. 3(12), 1023–1032 (2021).
91. Gerstner W Kistler WM Naud R Paninski L Neuronal dynamics: From single neurons to networks and models of cognition 2014 Cambridge University Press
Gerstner, W., Kistler, W. M., Naud, R. & Paninski, L. Neuronal dynamics: From single neurons to networks and models of cognition (Cambridge University Press, 2014).
92. Zhu, L., Li, J., Wang, X., Huang, T. & Tian, Y. Neuspike-net: High speed video reconstruction via bio-inspired neuromorphic cameras. In Proceedings of the IEEE/CVF international conference on computer vision, 2400–2409 (2021).
93. Hendrycks, D. & Gimpel, K. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415 (2016).
94. Han, S., Pool, J., Tran, J. & Dally, W. Learning both weights and connections for efficient neural network. Advances in neural information processing systems, 28, (2015).
95. Shen J Liu JK Wang Y Dynamic spatiotemporal pattern recognition with recurrent spiking neural network Neural Comput. 2021 33 11 2971 2995 34474470
Shen, J., Liu, J. K. & Wang, Y. Dynamic spatiotemporal pattern recognition with recurrent spiking neural network. Neural Comput. 33(11), 2971–2995 (2021).34474470
96. Hunger, R. Floating point operations in matrix-vector calculus, Vol. 2019 (Munich University of Technology, Inst. for Circuit Theory and Signal, 2005).
97. Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y. & Potts, C. Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, 142–150, Portland, Oregon, USA, June (2011). Association for Computational Linguistics. URL http://www.aclweb.org/anthology/P11-1015.
98. Zhu, Y., Fang, W., Xie, X., Huang, T. & Yu, Z. Exploring loss functions for time-based training strategy in spiking neural networks. Adv. Neural Inf. Process. Syst., 36, (2024).
99. Kitaev, N., Kaiser, Ł. & Levskaya, A. Reformer: The efficient transformer. arXiv preprint arXiv:2001.04451 (2020).
100. Katharopoulos, A., Vyas, A., Pappas, N., Fleuret, F. Transformers are RNNS: Fast autoregressive transformers with linear attention. In International conference on machine learning, 5156–5165 (PMLR, 2020).
101. Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L. et al. Rethinking attention with performers. arXiv preprint arXiv:2009.14794 (2020).
102. Lee-Thorp, J., Ainslie, J., Eckstein, I. & Ontanon, S. Fnet: Mixing tokens with fourier transforms. arXiv preprint arXiv:2105.03824 (2021).
103. Xiong, Y. et al. Nyströmformer: A nyström-based algorithm for approximating self-attention. In Proceedings of the AAAI Conference on Artificial Intelligence 35, 14138–14148 (2021).
104. Zhu, Z. & Soricut, R. H-transformer-1d: Fast one-dimensional hierarchical attention for sequences. arXiv preprint arXiv:2107.11906 (2021).
105. Romero, D. W., Knigge, D. M., Gu, A., Bekkers, E. J., Gavves, E., Tomczak, J. M. & Hoogendoorn, M. Towards a general purpose cnn for long range dependencies in d. arXiv preprint arXiv:2206.03398 (2022).
