
==== Front
bioRxiv
BIORXIV
bioRxiv
2692-8205
Cold Spring Harbor Laboratory

39131281
10.1101/2024.07.17.603927
preprint
4
Article
Reward history guides focal attention in whisker somatosensory cortex
http://orcid.org/0000-0003-4064-5148
Ramamurthy Deepa L. 1
Rodriguez Lucia 12
Cen Celine 1
Li Siqian 1
Chen Andrew 1
http://orcid.org/0000-0003-4646-8170
Feldman Daniel E. 13*
1. Department of Neuroscience and Helen Wills Neuroscience Institute, UC Berkeley
2. Neuroscience PhD Program, UC Berkeley
3. Lead Contact
AUTHOR CONTRIBUTIONS

D.L.R. and D.E.F. designed the study. D.L.R., L.R., C.C., S.L., and A.C. performed the experiments. D.L.R., L.R., C.C., and S.L. analyzed the data. D.L.R., L.R., and D.E.F. wrote the manuscript.

* Corresponding author. Correspondence should be addressed to: Daniel E. Feldman, dfeldman@berkeley.edu, 142 Weill Hall #3200, Dept. of Neuroscience, Univ. of California, Berkeley, Berkeley, CA 94720-3200
09 9 2024
2024.07.17.603927https://creativecommons.org/licenses/by-nc-nd/4.0/ This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which allows reusers to copy and distribute the material in any medium or format in unadapted form only, for noncommercial purposes only, and only so long as attribution is given to the creator.
nihpp-2024.07.17.603927.pdf
Prior reward is a potent cue for attentional capture, but the underlying neurobiology is largely unknown. In a novel whisker touch detection task, we show that mice flexibly shift attention between specific whiskers on a trial-by-trial timescale, guided by the recent history of stimulus-reward association. Two-photon calcium imaging and spike recordings revealed a robust neurobiological correlate of attention in the somatosensory cortex (S1), boosting sensory responses to the attended whisker in L2/3 and L5, but not L4. Attentional boosting in L2/3 pyramidal cells was topographically precise and whisker-specific, and shifted receptive fields toward the attended whisker. L2/3 VIP interneurons were broadly activated by whisker stimuli, motion, and arousal but did not carry a whisker-specific attentional signal, and thus did not mediate spatially focused tactile attention. Together, these findings establish a new model of focal attention in the mouse whisker tactile system, showing that the history of stimuli and rewards in the recent past can dynamically engage local modulation in cortical sensory maps to guide flexible shifts in ongoing behavior.

Attention
sensory coding
value
trial history
mouse
VIP interneuron
==== Body
pmcINTRODUCTION

Humans and other animals engage attention to prioritize processing of behaviorally relevant stimuli in complex environments, including for vision1, audition2 and touch3. Past experiences across different timescales play an important role in guiding attention4. In humans, prior stimulus-reward association is a highly robust cue for attentional capture5–17, such that perceptual detection or discrimination is selectively enhanced for previously rewarded stimuli, even when those stimuli are no longer important to current goals14. The effects of reward history on allocation of attentional priority have been termed experience-driven attention15, value-driven attention12–14, memory-guided attention16, or attentional bias by previous reward history17.

The neurobiology of attention has been studied extensively in non-human primates1,18, and has been shown to involve boosting of signal-to-noise ratio for neural encoding of attended sensory features across the cortical sensory hierarchy, including primary sensory cortex. However, the fine-scale organization of attentional boosting in sensory cortex, and the neural circuits that control it, remain unclear. Mice provide powerful cell-type specific tools to identify precise neural circuit mechanisms underlying attentional processing. Here, we developed a new model of focal attention based on reward history in the mouse whisker tactile system, and used it to investigate the neurobiological basis of attention in sensory cortex.

We studied attentional behavior in head-fixed mice during a Go/NoGo whisker touch detection task which included Go stimuli on many different whiskers. When the whisker location of a touch stimulus was unpredictable from trial to trial, mice naturally used the history of stimulus-reward association as a cue to guide attention to specific, recently rewarded whiskers. This task generates a rich set of trial histories to probe which stimulus and reward contingencies guide attention, and allows us to track the shifting locus of attention on a trial-by-trial time scale. Using this platform, we identified a robust, spatially focused neural correlate of attention in whisker somatosensory cortex (S1), and tested a major circuit model of attention involving VIP interneurons in sensory cortex. More broadly, our study shows that in addition to the well-known role of learned stimulus-reward associations in slowly driving cortical map plasticity, the recent history of stimuli and rewards acts on a fast time scale to dynamically and locally modify cortical sensory maps with corresponding stimulus-specific shifts in perceptual sensitivity.

RESULTS

To study attention, we developed a head-fixed whisker detection task. Mice had nine whiskers inserted in a piezo array, and on each trial, one randomly selected whisker (Go trials) or no whisker (NoGo trials) was deflected in a brief train. Mice were rewarded for licking during a response window on Go trials (Hits) but not on NoGo trials (False Alarms) (Fig. 1A–B). The task was performed in darkness and incorporated a delay period of 0, 0.5, or 1-sec in different mice to separate whisker stimuli from response licks. Go and NoGo trials were randomly intermixed with a variable inter-trial interval (ITI), and whisker identity on Go trials was randomly selected. Thus, mice could not anticipate the upcoming whisker or precise trial timing. Sequential trials could be one Go and one NoGo, two NoGo trials, two Go trials on different whiskers, or two Go trials on the same whisker (Fig. 1C).

Expert mice effectively distinguished Go from NoGo trials, quantified by d’ from signal detection theory (Fig. S1A). We analyzed behavior in each session during a continuous task-engaged phase that excluded early and late low-performance epochs (d’<0.5) that reflect motivational effects19. In expert mice, overall mean d’ was 0.923 ± 0.051 (n = 476 sessions, 22 mice), but local fluctuations in d’ regularly occurred over the time course of several trials, suggesting that the recent history of stimuli or rewards may dynamically alter sensory detection behavior (Fig. S1A). To examine this, we classified each current Go trial based on the history of whisker stimuli and reward on immediately preceding trials. We used trial history categories of: i) prior NoGo, ii) prior Hit to the same whisker as the current trial, iii) prior Hit to a different whisker, iv) prior Miss to the same whisker, and v) prior Miss to a different whisker. Current NoGo trials were classified into categories of prior NoGo, prior Hit (to any whisker), or prior Miss (to any whisker). We separately tracked trials preceded by a single Hit trial from those preceded by multiple sequential prior Hits (“prior >1 Hit”; Fig. S1B).

Reward history cues focal attention in the whisker system

Prior trial history strongly influenced detection on the current trial. When the prior trial was a NoGo, mice detected the current Go whisker with d’ = 1.13 ± 0.02 (termed d’NoGo, mean ± SEM across 476 sessions, 22 mice), which we consider baseline detection sensitivity. d’ for detecting a given Go whisker was elevated following 1 prior Hit to the same whisker, and even more so following multiple consecutive Hits to the same whisker (d’>1HitSame = 2.45 ± 0.11, p = 1.0e-4 vs d’NoGo, permutation test). In contrast, d’ was reduced following one or multiple prior Hits to a different whisker than the current trial (d’>1HitDiff = 0.82 ± 0.06, p = 1.0e-4 vs d’NoGo, permutation test) (Fig. 1D–E). d’ after multiple prior Hits to the same whisker (d’>1HitSame) was substantially greater than after multiple prior Hits to a different whisker (d’>1HitDiff) (Fig. 1D). Thus, recent Hits engage a whisker-specific boost in detection, evident as a whisker-specific increase in d’, termed Δ d’ (Fig. 1D). This effect was found across mice with 0, 0.5, or 1-sec delay period (p = 1.0e-4 for d’>1HitSame vs d’>1HitDiff in each case), so these data were combined for behavioral analyses (Fig. S1D).

Trial history-dependent boosting of detection required the conjunction of prior stimulus plus reward, because d’ did not increase after a prior Miss to the same whisker (d’MissSame = 0.90 ± 0.09, p = 4.2e-3 relative to d’NoGo). Boosting also failed to occur if the mouse licked to the prior Go but received no reward (i.e., unrewarded hits) or a very small reward (<4% of maximal reward volume (d’>1HitSame vs d’>1HitDiff, p = 0.49, permutation test) (Fig. 1F). Thus, boosting was not driven by prior whisker deflection alone, or stimulus-evoked licking, but by whisker-reward association on recent trials.

We interpret this effect as whisker-specific attention, because it shares characteristic features of stimulus-specific attention documented in primates during detection tasks20–24. This involves both increased sensitivity (d’) and a shift in decision criterion (c, also referred to as response bias), with the whisker-specific Δ d’ reflecting increased Hit rate following >1 prior Hit to the same whisker relative to >1 prior Hit to a different whisker (p = 1.0e-4, permutation test; Fig. 1G). The whisker-specific increase in d’ was observed across mice (Fig. 1H; p = 1.0e-4, permutation test). We also observed shifts in criterion20,22–23 (Δ c) which with a more modest whisker-specific component (c>1HitSame vs c>1HitDiff, p = 0.002, permutation test) (Fig. 1I). On average, the whisker-specific shift in sensitivity was larger than the whisker-specific shift in criterion (Fig. 1J; Δ d’>1HitSame = 1.4 ± 0.20, Δ d’>1HitDiff = −0.47 ± 0.11, p = 1.0e-4, permutation test; Δ c>1HitSame = −2.03 ± 0.18, Δ c>1HitDiff = −1.34 ± 0.08, p = 7.0e-4, permutation test).

Mice trained on all 3 delay periods showed the whisker-specific boost in d’ (Fig. S1C–D). Reaction times on Go trials were reduced by prior Hit trials, as expected for attention (assessed in mice without a delay period), but this effect was not whisker-specific, which may reflect a ceiling effect (Fig. S1E). The task interleaved whisker deflections of varying amplitude, and whisker-specific boosting of d’ was greatest for current Go trials with low-amplitude (weak) deflections, as expected for attention (Fig. S1F). Trial history effects were consistently observed across different variations of the task (used in different mice) in which we manipulated the stimulus probability of each whisker in blocks, or manipulated the probability of sequential same-whisker Go trials to make prior same Hit histories more likely (Fig. S1G–H). Trial history effects were driven by stimulus-reward association, not by stimulus salience or reward probability, because whisker deflections were physically identical and reward probability was always 100% for each whisker.

Behavioral shifts in d’ exhibited the hallmark effects of focused attention: spatial specificity, temporal specificity, and flexible targeting. Spatially, d’ was boosted most strongly by prior Hits to the same whisker, more weakly by prior Hits to an immediate same-row or same-arc neighbor, and not at all by prior Hits to a diagonally adjacent or more distant whisker (Fig. 1K). Thus, boosting is somatotopically organized. Temporally, attentional boosting fell off with the time interval between Go trials (which varied due to variable ITI and intervening NoGo trials), and subsided after ~10s (Fig. 1L). Enhancement of d’ was flexibly shifted to different whiskers in an interleaved manner, and had similar magnitude when cued by trial history to any of the 9 whiskers in rows B-D or arcs 1–3 (Fig. 1M). Thus, mice use recent history of stimulus-reward association to dynamically boost sensory detection of spatially specific whiskers on a rapid trial-by-trial timescale, consistent with attentional enhancement1,18,20–23, 25–26. These properties strongly resemble attentional capture guided by reward history in humans and non-human primates4–17,24.

Whisker-specific attention is not mediated by whisker or body movement

To test whether trial history effects involve whisker or body movement, we extracted these movements, plus pupillary dilations related to arousal27–28, from behavioral videos of 9 mice (74 sessions) using DeepLabCut29 (Fig. S2A). Reward retrieval at the end of Hit trials was associated with whisker movement, body movement (detected from platform motion), and pupil dilation that slowly subsided during the ITI before the next trial (Fig. 2A). The magnitude of movement and pupil dilation during the ITI was greater after 1 or >1 prior Hits, relative to prior NoGo, but was identical for prior same and prior different conditions (Fig. 2A, Fig. S2B). Thus, mice exhibited increased motion and arousal following prior Hits. During the subsequent trial, whisker stimulation evoked modest whisker and body motion during the stimulus period, and these were also heightened after prior Hits (p = 1e-4, permutation test), indicating that behavioral arousal and motion effects from prior Hits persisted into subsequent trials, but did not differ between prior same whisker and prior different conditions (p = 0.39, permutation test Fig. 2A–B).

To test whether active whisker movement contributed to the whisker-specific Δ d’ effect, we paralyzed whisker movements with Botulinum toxin B (Botox) injection in the vibrissal pad in 4 task-expert mice. Paralysis was maintained over 7–8 days, and history effects were compared between standard sessions prior to Botox and the Botox sessions. The average whisker-specific shift in behavioral d’ and c following >1 prior hit did not differ between Botox and non-Botox sessions (Fig. 2C–D) demonstrating that whisker-specific attentional effects do not require active whisker movement. Thus, although prior hits also engage increases in whisker motion, body motion and arousal, whisker-specific attentional effects were independent of these effects on global behavior state.

Neural correlates of focal attention in L2/3 PYR cells in S1

The somatotopic precision of attentional effects on behavior (Fig. 1K) suggests a neural basis in a somatotopically organized brain area like S1. We performed 2-photon imaging in S1, using Drd3-Cre;Ai162D mice that transgenically express GCaMP6s in L2/3 PYR cells. Mice performed the task with a delay period that separated whisker-evoked responses (analyzed 0–799 ms after stimulus onset) from later licks and rewards, and any trials with early licks were excluded. Imaging fields were localized in the S1 whisker map by post-hoc cytochrome oxidase staining for whisker barrel boundaries in L4 (Fig. 3A–B). 61% of L2/3 PYR neurons were whisker responsive, and trial history modulated whisker responses for many individual neurons. For example, in Fig. 3C, neurons increased their response to the C3 whisker after >1 prior hit to that same whisker, but not after >1 prior hit to a different whisker.

On average, L2/3 PYR cells responded to a given Go whisker when the prior trial was a NoGo, responded more strongly following 1 prior Hit to the same whisker (p = 1e-4), and even more after >1 prior Hit to the same whisker (p = 1e-4). This boosting of sensory responses did not occur when the prior trial was a Miss to the same whisker, or >1 Hit to a different whisker (n = 6 mice, 70 sessions, 6906 cells, Fig. 3D–E). Thus, recent stimulus-reward association modulated whisker-evoked ΔF/F in L2/3 PYR cells in a way that closely resembled the behavioral attention effect (Fig. 3E vs Fig. 1D). On current NoGo trials, no whisker stimulus was presented and ΔF/F traces were largely flat, except for NoGo trials following prior Hit trials, which exhibited a surprising rising ΔF/F signal. We interpret this as an expectation or arousal effect, which parallels the increased FA rate on these trials (Fig. 1G). These effects were evident in single example fields (Fig. S3A). History-dependent boosting of whisker responses was most evident in whisker-responsive cells (defined from trials after a prior NoGo), and did not occur in cells that were non-responsive after a prior NoGo (Fig. S3B–C). Boosting after prior Hits was whisker-specific in all 6/6 mice, but its magnitude varied across mice (Fig. 3F) and correlated with the magnitude of the behavioral attention effect measured by Δ d’ in each mouse (Fig. 3G).

We quantified the attention effect in individual cells using three attention modulation indices (AMI). AMI>1HitSame-NoGo and AMI>1HitDiff-NoGo quantify the change in whisker-evoked response observed after >1 prior Hit to the same (or different) whisker vs after a prior NoGo. Most cells showed positive AMI>1HitSame-NoGo values and negative AMI>1HitDiff-NoGo values, indicating up- and down-modulation of response magnitude by the identity of the prior whisker Hit. AMI>1HitSame->1HitDiff compares response magnitude after >1 prior hit to the same whisker vs >1 prior hit to a different whisker. This was shifted to positive values for most L2/3 PYR cells, indicating whisker-specific attentional modulation (Fig. 3H, I). This was reproducible across individual mice (Fig. S3D). Modulation of whisker-evoked responses for all cells sorted by AMI>1HitSame->1HitDiff is shown in Fig. S3E.

History-dependent boosting was somatotopically restricted in S1. After >1 prior hits to a given whisker, responses to an immediate same-row adjacent neighboring whisker were boosted strongly, those to an immediate same-arc neighbor were boosted less, and those to diagonal adjacent neighbors or further whiskers were boosted the least or not at all. This somatotopic profile of ΔF/F boosting strongly resembled the somatotopy of behavioral d’ boosting (Fig. 3J). Spatially within S1, >1 prior hits to a reference whisker boosted whisker-evoked ΔF/F to that whisker most strongly for L2/3 PYR cells within the reference whisker column and in the near half of the neighboring column, and weakly or not at all beyond that. This somatotopically constrained boosting was not observed for >1 prior hits to a different whisker, or following a miss to the reference whisker (Fig. 3K). Analysis of individual cell AMI confirmed somatotopically precise boosting (Fig. S4A). This defines the spatial profile of the trial history-based ‘attentional spotlight’ in S1 as boosting responses to the cued whisker within a region of 1.5 columns width in L2/3. At the center of the spotlight, the representation of the attended whisker is boosted within its own column (Fig. S4B–C).

Attention involves shifts in receptive fields toward the attended whisker

In primates, attention not only increases sensory response magnitude and signal-to-noise ratio in sensory cortex, but can also shift neural tuning toward attended stimuli30–32. We tested whether attention involves receptive field shifts by L2/3 PYR cells in S1. We analyzed all whisker-responsive cells relative to the boundaries of the nearest column (defined in the Prior NoGo condition). We calculated the mean receptive field across 9 whiskers centered on the columnar whisker (CW) when the prior trial was a NoGo (Fig. 4A, black traces in center panel). We then recalculated the receptive field, for these same cells, for >1 Prior Hit to each of the surround whiskers (purple traces in outside panels). For many whiskers, recent stimulus-reward association shifted the mean receptive field toward the attended whisker or nearby whiskers (Fig. 4A). To quantify this effect, we calculated the tuning center of mass (CoM) across the 3 by 3 whisker array for these mean receptive fields. History-based cueing to an attentional target whisker generally shifted tuning CoM towards that whisker (Fig. 4B). An exception was when attention was cued upward, which caused little upward CoM shift. This was not explained by known experimental factors, but could reflect a spatial asymmetry in attentional effects33 on whisker touch (Fig. 4B). Attentional cueing to the CW elevated responses to that whisker within its S1 column but caused very little tuning change (Fig. S5A).

To measure receptive field shifts in individual neurons, we quantified the shift in CoM for each responsive cell along an attentional axis from the CW to the attentional target whisker, which was one of the 8 surrounding whiskers (Fig. S5B–C). In the >1 prior Hit condition, the mean CoM shift along the attentional axis was 0.12 ± 0.04 (n = 173 cells with sufficient whisker sampling, p = 3.6e-3, one-sample permutation test vs. mean of 0), indicating a tuning shift toward the attended whisker. A range of tuning shifts were observed, with significantly more cells showing CoM shifts towards the attended whisker than away from it (62% vs. 38%, p = 2.3e-3, binomial exact test for difference from 0.5). Thus, attentional cueing involves receptive field shifts as well as modulation of whisker response magnitude.

Attention boosts population decoding of recently rewarded whiskers

Is the magnitude of attentional boosting of L2/3 PYR responses sufficient to improve neural coding of attended whisker stimuli on single trials? To test this, we built a simple neural decoder that uses logistic regression to predict the presence of any whisker stimulus from single-trial population mean ΔF/F calculated across all whisker-responsive cells in a single imaging field. Each field spanned ~1–1.5 columns within the 9-whisker region of S1, typically centered over the location of column corresponding to the center whisker on the piezo array. A separate decoder was trained on each session (n = 70 sessions, 6 mice), and performance was assessed from held-out trials (Fig. 5A). Because S1 is somatotopically organized, we observed modest, above-chance performance for detecting any of the 9 whiskers from mean activity in a single field (relative to NoGo trials or shuffled data), strong performance for detecting the field best whisker (fBW) that is topographically matched to the field location, and no ability to detect non-topographically aligned whiskers (non-fBWs) (Fig. 5B).

Population mean ΔF/F and decoder performance were related to current trial outcome, with greater ΔF/F and stimulus prediction on current Hit trials than Miss trials, and on current false alarm trials than correct rejection trials. This is consistent with the known representation of sensory decision in S134–35 (Fig. 5C–D). Next, we examined decoding as a function of prior trial history. Detection of any whisker stimulus from Go trials was improved following >1 Hit to the same whisker, relative to prior NoGo (p = 1e-0.4, permutation test) or to >1 prior Hit to a different whisker (p = 0.001, permutation test) (Fig. 5E). This boost in decoding performance did not occur when the decoder was tested only on fBW trials but did occur for non-fBW trials (Fig. 5F–G). This same effect was found when decoding was evaluated on current Hit trials only, so this is an effect of trial history, not current sensory decision (Fig. 5E–F).

Thus, attention to a whisker strengthens its encoding on single trials in S1. The lack of improvement in decoding attended fBWs likely reflects the strong coding of these whiskers under baseline conditions, such that attentional boosting of CW responses (Fig. S4) does not further improve detection.

Attentional modulation of neural coding with Neuropixels spike recordings

To examine attentional modulation of S1 neural coding at finer temporal resolution and across layers, we recorded extracellular spiking in S1 using Neuropixels probes36 that spanned L1–6 (Fig. 6A). Mice performed a modified version of the task in which Go stimuli were distributed over only 4 or 5 whiskers, rather than 9, to enable adequate sampling of each history condition per session. Behaviorally, mice performing the task with 4–5 whiskers showed the whisker-specific Δ d’ attention effect, but at lower magnitude due to the smaller number of whiskers (Fig. S6A). All sessions included the CW for the recording site plus 3 nearby whiskers. We spike sorted to identify single units, classified units as regular-spiking (RS) or fast-spiking (FS), and assigned laminar identity based on CSD analysis of local field potentials (Fig. S6B–F). Many single units showed history-dependent modulation of whisker-evoked spiking (Fig. 6B).

S1 units responded to each deflection in the stimulus train. On average for L2/3 RS units, whisker-evoked spiking was boosted in Go trials after >1 prior Hit to the same whisker, relative to >1 prior Hit to a different whisker (Fig. 6C–D). This effect did not occur after prior Miss trials, and firing on NoGo trials was not significantly regulated (Fig. 6D). L4 RS units showed only a slight trend for history-dependent modulation that did not reach significance (Fig. 6E–F). L5a/b RS units (grouped together) showed a whisker-specific attentional boost similar to L2/3 (Fig. 6G–H).

To examine heterogeneity across units we calculated AMI for each unit. Most L2/3 RS units responded more strongly after prior Hits to the same whisker relative to prior NoGo (as measured by AMI>1HitSame-NoGo) and more weakly after prior Hits to a different whisker relative to prior NoGo (as measured by AMI>1HitDiff-NoGo). This whisker-specific attentional effect, evident as the separation between AMI>1HitSame-NoGo and AMI>1HitDiff-NoGo distributions, was not present in L4, and was weaker in L5a/b (Fig. 6I). These laminar trends were also apparent in AMI>1HitSame->1HitDiff, which was shifted positively in L2/3 relative to L5a/b and L4 units (Fig. 6J). Calculating the mean AMI across neurons confirmed whisker-specific attentional shifts in L2/3 and L5, but not in L4 (Fig. 6K; (AMI>1HitSame-NoGo vs AMI>1HitDiff-NoGo, L2/3: p = 0.03, L4 p = 0.65, L5a/b: p = 0.01; AMI>1HitSame->1HitDiff, L2/3: p = 0.031, L4 p = 0.51, L5a/b: p = 0.06, permutation test). This suggests that history-dependent attentional modulation is not simply inherited from the thalamus, but has a cortical component.

VIP interneurons do not carry a simple “attend here” signal

We used reward history-based attention in S1 to investigate the candidate involvement of VIP interneurons in attentional control. VIP cells are known to disinhibit PYR cells to increase PYR sensory gain during arousal, locomotion, and whisking37–41. For attention, this same VIP circuit has been hypothesized to be activated by long-range (e.g., top-down) inputs, and to act to amplify local PYR responses to selected sensory features40–41 (Fig. 7A). Whether VIP cells mediate goal-directed attention is still unclear42–43, and their involvement in history-based attention has not been tested. We used 2-photon imaging from L2/3 VIP cells in VIP-Cre;Ai162 mice to ask whether VIP cells are activated when mice direct attention to a particular whisker column in S1 (Fig. 7B).

VIP cells in S1 are activated by arousal (indexed by pupil size), whisker and body movement, and goal-directed licking during this whisker detection task44. These behaviors all peak at the end of Hit trials, as mice retrieve rewards, and then systematically decline during the ITI, which ends with a 3-sec lick-free period that is required to initiate the next trial. As a result, VIP cell ΔF/F falls systematically during the ITI after Hit trials, correlated with these behavioral variables, and falls less after NoGo or Miss trials (Fig. 7C). VIP cells in S1 also show robust whisker stimulus-evoked ΔF/F transients, which ride on this declining baseline44. To test whether whisker-evoked VIP responses are greater when reward history cues attentional capture, we calculated mean ΔF/F traces (n = 7 mice, 103 sessions, 1843 VIP cells) as a function of trial history. Baseline (prestimulus) ΔF/F declined more steeply on trials following prior Hits than following prior NoGo or Miss, as expected. Superimposed on this, and clearest after detrending the baseline, whisker-evoked Δ F/F was also increased after prior Hit trials. However, this was not whisker-specific (Fig. 7D–F). AMI analysis confirmed increased responsiveness for most VIP cells in both prior >1 Hit same and prior >1 Hit different conditions, but no whisker-specific attention effect (the AMI>1HitSame->1HitDiff index was peaked at 0; Fig. 7G, for all cells individually see Fig. S7).

Together, these findings indicate that VIP cells as a population do not carry a whisker-specific attention signal, but do exhibit a general increase in activity with multiple prior Hits to any whisker that is consistent with global arousal and motion effects. Although we cannot rule out that a subpopulation of VIP cells may carry a whisker-specific “attend here” signal, the VIP population as a whole does not.

DISCUSSION

Attention captured by recent history cues4–17, including stimulus-, reward-, and choice-history, provides a powerful model to study mechanisms of attention. In our paradigm, enhanced detection (D d’) of whisker stimuli was driven by recent whisker stimulus-reward association, had the defining features of selective attention (spatially focused, flexibly allocated, and temporally constrained)4,20–23. Behavioral and neural effects in our study were whisker-specific, and thus did not correspond to a global arousal or motion effect45. They were not explained by priming, which occurs in response to stimulus presentation without reward association, does not require detection of the priming stimulus, and typically has short (<100 ms) duration46.

Attentional boosting in our study was not bottom-up attention, because it was not driven by physical stimulus salience (whisker stimuli had the same stimulus strength), and it was not top-down attention, because it was automatically engaged and not goal-directed (i.e., whisker stimuli were all equally rewarded, so it did not increase overall reward rate in the task). These results align well with automatic attentional capture by reward history in humans4–17, which is theorized to represent a category of attentional processes often called “selection history”8–10 distinct from classical top-down and bottom-up attention. Rodents are well known to exhibit history-dependent response biases (i.e., shifts in decision criterion for behavioral choices) in perceptual decision-making tasks47–51, including serial dependence47,51–52, contraction bias47–48,51, adaptation aftereffects53, win-stay/lose-switch strategies54–56 and choice alternation54. Our findings show that mice also use prior reward history to prioritize sensory processing, through stimulus-specific shifts in perceptual sensitivity (d’), in addition to shifts in decision criterion (c).

The neural correlates of attention have been primarily studied in non-human primates, and include increased sensory-evoked spike rate18,57–58, reduced variability18, neural synchrony modulation59, and changes in receptive fields18,30–32 including in receptive size, boosting of peak responses, and receptive field shifts. In primates, these effects are greatest at higher levels of the sensory hierarchy but also occur in primary sensory cortex18,60. The precise spatial organization of these coding effects in sensory cortex has been unknown, and has important implications for identifying the neural control circuits for attention. On the macroscopic scale, human brain imaging and focal pharmacological inactivation studies in non-human primates indicate that spatial attention in vision is retinotopically organized within visual cortical areas1,18,61–64. But the precise spatial organization of attentional modulation in sensory cortex (i.e., the spatial profile of the spotlight of attention) has not been known. Importantly, our task design (in which we track spontaneous behaviors, and separate stimulus and lick response windows with a delay period) allowed us to distinguish attentional signals from global arousal and motion signals, which dominate neural activity during behavior and are widespread across cortical areas45.

We took advantage of S1 whisker map topography65 to quantitatively define the precise spatial structure of the attentional spotlight relative to anatomical cortical columns in S1. Attentional capture boosted sensory responses to the attended whisker in a region comprising that whisker’s column plus the near half of surrounding columns. In this region, whisker-evoked spike rate and receptive field peak increased (in the central attended column) and receptive fields shifted toward the attended whisker (for cells in surrounding columns). Together, this increased total neural activity evoked by the attended whisker, both by increasing the number of PYR cells responding to that whisker, and by elevating the number of spikes per cell. This somatotopically restricted boosting61–66 is distinct from the spatially broad modulation of sensory responses that occurs across entire cortical areas (or multiple areas) in response to global behavioral state (e.g., arousal indexed by pupil size, active whisker movement for S1, or locomotion for V1)45. Our results show that attentional boosting can be flexibly targeted with a precision of ~300 μm in cortical space for stimulus-specific modulation of the neural code. Thus, neural control circuits for attention (which may involve feedforward, local, feedback, or neuromodulatory circuits) must operate with this spatial precision.

History signals in mouse cortex have not previously been described for attention, but have been identified in posterior parietal cortex (PPC) and orbitofrontal cortex (OFC) during decision making67–69 and in reversal learning70–73. S1 receives instructional signals from OFC that are necessary for reversal learning72, but whether this pathway plays a role in attention cued by recent reward history is unknown.

The neural mechanisms and control circuits for attention remain poorly understood, and likely differ between different forms of attention. We found that attention to touch modulates PYR sensory responses in L2/3 and L5a/b but not L4, suggesting either an intracortical origin, or a thalamic origin in secondary thalamic nuclei like the posterior medial nucleus (POm), which projects to L2/3 and L5a. Thalamic control of attention has been implicated in visual and tactile cross-modal attention tasks in mice74, as well as some non-human primate studies75. We tested one major circuit model39–43 for attentional boosting in sensory cortex, that long-range inputs amplify pyramidal (PYR) cell sensory responses by activating local VIP interneurons in sensory cortex39,76. L2/3 VIP interneurons receive local, feedforward, and feedback glutamatergic input, as well as by neuromodulatory input, and inhibit other cortical interneurons to disinhibit PYR cells, thus boosting PYR sensory responses39,76. VIP interneurons are activated by global behavioral state (e.g., locomotion and whisking), by spontaneous arousal during quiet wakefulness27,37–38, and by top-down contextual signals40. Top-down input from anterior cingulate cortex (ACC) to VIP cells in sensory cortex has been suggested to mediate top-down attentional effects on sensory processing39. However, recent studies have questioned this model42,77, finding that L2/3 VIP modulation of PYR activity is orthogonal to attentional effects in a cross-modal attention task42. We tested the potential involvement of VIP cells in focal attention by asking whether VIP cell activity is enhanced in attended columns, as required if these cells contribute to boosting of PYR cell responsiveness. We found that L2/3 VIP cells, at least as a full population, do not carry this whisker-specific “attend-here” signal, but instead show general, non-whisker-specific activation in response to any prior Hit, consistent with an arousal- or global behavior-related signal. This suggests VIP cells are more engaged in global modulation of whisker sensory responsiveness during arousal and motion, and not whisker-specific attention cued by reward history.

Attention to whisker touch cued by recent reward history in mice is a novel paradigm for studying the neurobiological mechanisms of focal attention. This model complements recent visual tasks that aim to study top-down79–82 and bottom-up83–84 attention in head-fixed mice. Together, these paradigms can reveal the extent to which common vs. distinct neurobiological mechanisms are engaged in different forms of attention, and across different sensory modalities.

METHODS

Animals

All methods followed NIH guidelines and were approved by the UC Berkeley Animal Care and Use Committee. The study used 22 mice. These included 7 Drd3-Cre;Ai162D mice and 10 VIP-Cre;Ai162D mice (used for behavior and 2-photon imaging), and 5 offspring from Drd3-Cre × Ai162D crosses (genotype not determined) used in extracellular recording experiments. VIP-Cre (JAX # 10908) and Ai162D mice (JAX # 031562) were from The Jackson Laboratory. Drd3-Cre mice were from Gensat MMRRC (strain number 034610).

Mice were kept in a reverse 12:12 light cycle, and were housed with littermates before surgery and individually after cranial window surgery. Mice were roughly evenly divided between male and female, and no sex differences were found for the results reported here. Behavioral, imaging, and analysis methods were as described in Ramamurthy et al., 202344, and are here described more briefly. Results reported for 7 VIP-Cre;Ai162D mice and 3 Drd3-Cre;Ai162D mice are new analyses which include data from the dataset reported in Ramamurthy et al., 2023.

Surgery for behavioral training and 2-photon imaging

Mice (2–3 months of age) were anesthetized with isoflurane (1–3%) and maintained at 37°C. Dexamethasone (2 mg/kg) was given to minimize inflammation, meloxicam (5–10 mg/kg) for analgesia, and enrofloxacin (10 mg/kg) to prevent infection. Using sterile technique, a lightweight (<3 g) metal head plate containing a 6 mm aperture was affixed to the skull using cyanoacrylate glue and Metabond (C&B Metabond, Parkell). The headplate allowed both head fixation and 2-photon imaging through the aperture. Intrinsic signal optical imaging (ISOI) was used to localize either C-row (C1, C2, C3) or D-row (D1, D2, D3) barrel columns in S185, and a 3 mm craniotomy was made within the aperture using a biopsy punch over either the C2 or D2 column. The craniotomy was covered with a 3 mm diameter glass coverslip (#1 thickness, CS-3R, Warner Instruments) over the dura, and sealed with Metabond to form a chronic cranial window. Mice were monitored on a heating pad until sternal recumbency was restored, given subcutaneous buprenorphine (0.05 mg/kg) to relieve post-operative pain and then returned to their cages. After the mice recovered for a week, behavioral training began.

Behavioral task

To motivate behavioral training, each mouse received 0.8–1.5 mL of water daily, calibrated to maintain 85% of pre-training body weight. Mice were weighed and observed daily. Behavioral training sessions took place 5–7 days per week. For behavioral training, the mouse was head-fixed and rested on a spring-mounted stage44,86. Nine whiskers were inserted in a 3 × 3 piezo array, typically centered on a D-row or C-row whisker. Piezo tips were located ~5 mm from the face, and each whisker was held in place by a small amount of rubber cement. A tenth piezo was present near the 3×3 array but did not hold any whisker (“dummy piezo”). A capacitive lick sensor (for imaging experiments) or an infrared (IR) lick sensor (for extracellular recording experiments) detected licks, and water reward (mean 4 μl) was delivered via a solenoid valve. Mice were transiently anesthetized with isoflurane (0.5–2.0%) at the start of each session to enable head-fixation and whisker insertion, after which isoflurane was discontinued and behavioral testing began after the effects of anesthesia had fully recovered. Behavior was performed in the dark with 850 nm IR illumination for video monitoring. Masking noise was presented from nearby speakers to mask piezo actuator sounds. Task control, user input and task monitoring were performed using custom Igor Pro (WaveMetrics) routines and an Arduino Mega 2560 microcontroller board.

Training stages

A series of training stages (1–5 days each) were used to shape behavior on the Go/NoGo detection task. In Stage 1, mice were habituated to the experimental rig and to handling. In Stage 2 mice were head-fixed and conditioned to lick for water reward at the port. In Stage 3, mice received a reward (cued by a blue light) for suppressing licks for at least ~3 seconds, termed the Interlick Interval (ILI) threshold. Stage 4 introduced whisker stimulation for the first time, with 50% Go trials (whisker stimulation) and 50% NoGo trials (no whisker stimulus), with automatic reward delivery in the response window on Go trials (i.e., classical conditioning). The dummy piezo was actuated on all NoGo trials, so that any unmasked piezo sounds did not provide cues for task performance. The Go whisker randomly chosen from among 9 possible whiskers. This stage ended when mice shifted licks in time to occur before reward delivery. In Stage 5, training switched to operant conditioning mode, and mice were required to lick in the response window (0 – 300 ms after whisker stimulus onset) to receive a reward. There was no delay period at this stage. Learning progress was tracked by divergence of Go/NoGo lick probability. In Stage 6, the delay period was introduced. To do this, we introduced a trial abort window in which licking during the stimulus presentation period caused a trial to be canceled without reward, to discourage licking during the stimulus presentation. Simultaneously, we implemented a ramp-plateau reward gradient within the response window, so that later licks resulted in a larger reward. Over the course of this stage, the trial abort window was gradually lengthened, and the time of reward plateau was gradually increased. Learning progress was tracked by the gradual increase in median trial first lick time. Stage 7 represented the final whisker detection task, which included completely randomized Go/NoGo trials, fixed trial abort window and reward plateau parameters, and eliminating the blue light that signaled reward delivery. Mice were deemed task experts when they exhibited stable performance at d’ > 1 (mean running d-prime) for three consecutive sessions. Expert mice performed 500 – 1000 trials daily.

Task structure

Task structure was identical to Ramamurthy et al., 202344. Briefly, on each trial, a whisker stimulus was applied to either one randomly selected whisker (Go trials, 50–60% of trials) or no whisker (NoGo trials, 40–50% of trials). The whisker stimulus consisted of a train of five deflections separated by 100 ms each. Every deflection was a 300 μm amplitude (6⁰ angular deflection) rostrocaudal ramp-and-return movement with 5 ms rise/fall time and 10 ms duration. NoGo trials presented the same stimulus on a dummy piezo that did not contact a whisker, so that any unmasked auditory cue from piezo movement was matched between Go and NoGo trials. Trial onset was irregular with an ITI of 3 ± 2 s. Mice had to restrict licking to greater than 3-sec interlick interval (ILI) to initiate the next trial. On Go trials, mice were rewarded for licking within the 2.0 s response window with ILI of <300 ms. Licking was not rewarded on NoGo trials. Each trial outcome was recorded as a Hit, Miss, False Alarm, or Correct Rejection.

Different delay periods were used in different mice (Fig. S3C–D). Mice used in imaging or extracellular recording experiments had a delay period of 500 ms (for all spike recording mice and 7/14 2p imaging mice) or 1000 ms (for the other seven 2p imaging mice). This was used to separate sensory-driven neural activity from action- and reward-related activity. Three mice used only for behavioral data collection were tested without a delay period, which enabled testing of attention effects on lick response latency (Fig. S3E).

Because reward size varied with lick time in the response window, reward volume varied across trials. In addition, for mice with the 1000 ms delay period, mice sometimes licked on Go trials after stimulus presentation but before the response window opened, and thus earned no reward. These represent unrewarded Hits, so that attention effects could be quantified based on absence of reward and reward size on prior Hit (Fig. 1F).

Task variations

All mice in this study were trained on an “equal probability” (EqP) version of the task in which whisker identity on each Go trial was randomly selected from nine possible whiskers with equal 1/9 probability (EqP). Two other task variations manipulated either the global or local probability of each specific whisker, while still randomly selecting whisker identity on each trial. In high probability (HiP) sessions, we manipulated the global stimulus probability of each whisker by presenting one whisker with higher probability (80% of Go trials) than the others. This was done in 300–400 trial blocks, interleaved with standard EqP blocks (13 mice) on the same day. In high probability of same whisker (HiPSame) sessions (run on separate days from EqP sessions) we manipulated the local probability of repeating the same whisker stimulus on consecutive trials, while maintaining the overall probability of each whisker at 1/9. This was done in blocks interleaved with EqP blocks (7 mice). Whisker-specific attentional cueing was observed in all 3 task variants (Fig. 1H), so data from all variations were combined for the rest of the analyses. All 22 mice were trained and tested on the EqP version and either the HiP or HiPSame version of the task, but EqP blocks/sessions contributed trials to history analysis only in a subset of animals, since multiple hits to a given whisker in a single session were adequately sampled only in sessions with a higher number of total trials.

A modified task version was used for extracellular recording experiments, in order to adequately sample trial history conditions when only 2–4 days of acute recording were possible per mouse. To do this, we reduced the number of whiskers sampled during Go trials from 9 whiskers to either 4 or 5 whiskers 5 (4–5-whisker task; Fig. 6A). The 4–5-whisker task was used in 3 mice for extracellular recordings, and was also applied in 2 mice that were used in PYR cell imaging, where performance could be compared to the standard 9-whisker task (Fig. S6A).

Behavioral movies & DeepLabCut tracking

Behavioral movies were acquired at 15–30 frames/sec using either a Logitech HD Pro Webcam C920 (modified for IR detection) or FLIR Blackfly S (BFS-U3–63S4M; used in video analyses). DeepLabCut29 was used to track spontaneous face and body movements. Movies were manually labeled to generate training datasets for tracking facial motion (snout tip, whisker pad and 2–3 whiskers), body motion (corner of the mouse stage, whose motion reflects limb and postural movements), pupil size (8 labels on the circumference of the pupil), eyelids (8 labels on the circumference of the eyelid), and licking (tongue and lickport). Three separate networks were trained (100,000–200,000 iterations) such that a good fit to training data was achieved (loss < 0.005). One network each was trained for face/body motion for the two camera setups (version 1 network: 1110 labeled frames from 37 video clips across 6 mice) and another for pupillometry (version 2 network: 2463 labeled frames from 27 video clips across 2 mice).

Behavioral movies from 74 sessions in 9 mice were analyzed for whisker motion (average across all whisker-related labels), body motion (stage corner) and pupil size (ellipse fit to the pupil markers). Blinking artifacts were removed using a one-dimensional moving median filter (40 frame window) applied to the trace of pupil size as a function of time. Pupil size measured on each frame was normalized to the mean pupil size over each individual session.

Whisker paralysis by Botox injection

Four mice (2 VIP-Cre;Ai162D mice and 2 Drd3-Cre;Ai162D mice) that were used for imaging experiments also underwent Botox injection to induce paralysis of whisking34,44,86–88. Both whisker pads were injected with Botox (Botulinum Neurotoxin Type A from Clostridium botulinum, List Labs #130B). A stock solution of 40 ng/μl Botox was prepared with 1mg/ml bovine serum albumin in distilled water. Each whisker pad was injected with 1 μl of a 10 pg/μl final dilution using a microliter syringe (Hamilton). Whisking stopped within 1 day, and gradually recovered in ~1 week. Following the initial dose, a 50% Botox supplement was injected once per week, as needed. Imaging was performed > 24 hours after any Botox injection.

Behavioral Analysis

Behavioral performance was assessed using the signal detection theory measures89 of detection sensitivity (d’) and criterion (c), calculated from Hit rates (HR) and False Alarm rates (FA), as per their standard definitions: d′=ZHR−ZFA

c=12(ZHR+ZFA)

where Z is the inverse cumulative of the normal distribution.

To assess the overall behavioral performance of mice on the whisker detection task, d’ was computed across trials over the entire behavioral session. For each session, a sliding d’ cutoff (calculated over a 50 trial sliding window) was applied to the start and end of the session, and analysis was restricted between the first and last trial that met the threshold, to minimize satiety effects. A standard sliding d’ cutoff of 0.5 was used for all behavioral and imaging analyses. We tested d’ cutoffs 0.5, 0.7 and 1 to ensure that choice of d’ cutoff did not affect key results. A sliding d’ cutoff of 1.2 was used for extracellular recording analyses.

Definition of trial histories

For each current trial, trial history conditions were defined based on outcome and stimulus on prior trials, as follows. These definitions are illustrated in Fig. S1B.

History conditions were defined for current Go trials as follows: Prior Miss On the same whisker: The current trial is a Go and the outcome on the previous trial was a Miss to the same whisker.

On a different whisker: The current trial is a Go and the outcome on the previous trial was a Miss to a different whisker.

Prior NoGo: The previous trial was a NoGo (any outcome).

Prior 1 Hit On the same whisker: The current trial is a Go and the outcome on the previous trial was a Hit to the same whisker.

On a different whisker: The current trial is a Go and the outcome on the previous trial was a Hit to a different whisker.

Prior >1 Hit: On the same whisker: The current trial is a Go and the outcomes on the previous two or more trials were Hits to the same whisker as the whisker presented on the current trial.

On a different whisker: The current trial is a Go and the outcomes on the previous two or more trials were Hits to a single consistent whisker that was a different identity than the whisker presented on the current trial (i.e., two or more Hits in a row to the same whisker that differed from the current Go trial).

For each history condition defined above for current Go trials, a matched condition was defined for current NoGo trials. This allowed us to compute behavioral d’ and c values within each history condition, and to compare neural signals on Go and NoGo trials within each history condition. History conditions for current NoGo trials were defined as follows: Prior Miss: The current trial is a NoGo and the outcome on the previous trial was a miss to any whisker. NoGo trials in this category were used for comparison with Go trials in both categories 1a and 1b above.

Prior NoGo: The previous trial was a NoGo (any outcome).

Prior 1 Hit: The current trial is a NoGo and the outcome on the previous trial was a hit to any whisker. NoGo trials in this category were used for comparison with Go trials in both categories 3a and 3b above.

Prior >1 Hit: The current trial is a NoGo and the outcomes on the previous two or more trials were Hits to any single repeated whisker. NoGo trials in this category were used for comparison with Go trials in categories 4a and 4b, above.

The history conditions above were defined based on sequences of consecutive trials, including both Go and NoGo. The history-dependent effects on detection behavior were maintained when history conditions were defined by ignoring NoGo trials, and categorizing history based solely on Go trials (data not shown). NoGo trials were ignored when characterizing the temporal profile of attention as a function of time since last Go trial (Fig. 1L).

2-photon calcium imaging

A Sutter Moveable-Objective Microscope with resonant-galvo scanning (RESSCAN-MOM, Sutter) was used to perform 2p imaging in expert mice. A Ti-Sapphire femtosecond pulsed laser (Coherent Chameleon Ultra II) tuned to 920 nm, or an ALCOR 920 nm fixed wavelength femtosecond fiber laser (Spark Lasers), was used for GCaMP6s excitation. A water-dipping objective (16x, 0.8 NA, Nikon) was used, and emission was band-pass filtered (HQ 575/50 filter, Chroma) and detected by GaAsP photomultiplier tubes (H10770PA-40, Hamamatsu). Single Z-plane images (512 × 512 pixels) were acquired serially at 7.5 Hz (30 Hz averaged every 4 frames) using ScanImage 5 software (Vidrio Technologies). Laser power measured at the objective was 60 – 90 mW. On average, 14 imaging fields (305 μm × 305 μm) were obtained per mouse at depths of 110 – 250 μm below the cortical surface90. If there was >25% XY overlap, imaging fields were required to be at least 20 μm apart in depth to avoid repeated imaging of the same cells. After completion of all imaging experiments, the mouse was euthanized and the brain was collected to perform histology.

Histological localization of imaging fields

The brain was extracted and fixed overnight in 4% paraformaldehyde. After flattening, the cortex was sunk in 30% sucrose and sectioned at 50–60 μm parallel to the surface. Cytochrome oxidase (CO) staining showed surface vasculature in the most superficial tangential section as well as boundaries of barrels in L4. Histological sections were manually aligned using Fiji91, and imaging fields were localized in the whisker map aided by the surface blood vessels imaged at the beginning of each session. The centroid of each of the nine anatomical barrels corresponding to whiskers stimulated in each session and the XY coordinates of all imaged cells were localized relative to barrel boundaries. A cell was located within a specific barrel column if >50% of its pixels were within its boundaries, and cells outside barrel boundaries were classified as septal cells. Major and minor axes of all barrels were averaged to calculate the mean barrel width.

Image processing, ROI selection, and ΔF/F calculation

Custom MATLAB pipeline code (Ramamurthy et al, 202344; adapted from LeMessurier, 201992) was used for image processing. Correction for slow XY drift was performed using dftregistration93. Regions-of-interest (ROIs) were manually drawn as ellipsoid regions over the somata of neurons visible in the average projection across the full imaging movie after registration. Mean fluorescence of the pixels in each ROI was calculated to obtain the raw fluorescence time series. For PYR cell imaging, neuropil masks were created as 10 pixel-wide rings beginning two pixels from the somatic ROI, excluding any pixels correlated with any somatic ROI (r>0.2). Mean fluorescence of neuropil masks was scaled by 0.3 and subtracted from raw somatic ROI fluorescence. Neuropil subtraction was not performed for VIP cells, which were spatially well-separated. The mean fluorescence time series was converted to ΔF/F for each ROI, defined as (Ft-F0)/F0, where F0 is the 20th percentile of fluorescence across the entire imaging movie and Ft is the fluorescence on each frame.

Quantification of whisker-evoked responses

Whisker-evoked ΔF/F signal on Go trials was quantified in a post-stimulus analysis window (7 frames, 0.799 s), relative to pre-event baseline window (2 frames, 0.270 s). Whisker responses were measured as (mean ΔF/F in the post-stimulus window – mean ΔF/F in the baseline window) for each ROI. On NoGo trials, the ΔF/F analysis was aligned to the NoGo stimulus (dummy piezo deflection). Whisker responses for each cell were normalized by z-scoring to prestimulus baseline activity. For some analyses, each ROI’s Go-NoGo response magnitude to every whisker was also calculated as (median whisker-evoked ΔF/F signal across Go trials – median ΔF/F signal across NoGo trials). Trials aborted due to licks occurring during the post-stimulus window (0 – 0.799 s) were excluded from analyses. If at least one whisker produced a significant response above baseline activity (permutation test), the cell was considered to be whisker-responsive. This was done by combining the whisker-evoked ΔF/F signal distribution on Go trials with the ΔF/F signal distribution on NoGo trials, randomly splitting the combined distribution into two groups and comparing the difference in their means to the true Go-NoGo distribution difference (10,000 iterations). Differences greater than the 95th percentile of the permuted distribution were assessed as significant. The nine whisker response p-values were corrected for multiple comparisons (False Discovery Rate correction94). Cells without a positive ΔF/F response to at least one whisker were considered non-responsive.

The standard method for assessing whisker-responsiveness used all trials belonging to each session (combining trials across all history conditions). In the analysis of attentional modulation of receptive fields, the significance of whisker responses was separately assessed using only trials in the Prior NoGo category and compared to trials in Prior >1 Hit condition, which allowed us to test whether there was history-dependent acquisition of whisker-evoked responses by previously non-responsive cells.

Definition of each cell’s columnar whisker (CW) and best whisker (BW)

Each cell’s anatomical home column was determined by histological localization of the cell relative to barrel column boundaries. For cells located within column boundaries, the CW was the whisker corresponding to its anatomical home location. For septa-related cells (i.e., those outside of column boundaries and above a L4 septum), the CW was the whisker corresponding to the nearest barrel column. The best whisker (BW) was defined for each cell as the whisker that evoked the numerically highest magnitude response.

Attention Modulation Index (AMI)

Multiple AMI metrics were used to quantify attentional modulation of whisker response magnitude in individual cells. The definitions were:

AMI (>1Hit-NoGo): AMI>1HitSame−NoGo=GoPrior>1HitSame−GoPriorNoGo|GoPrior>1HitSame+GoPriorNoGo|

AMI>1HitDiff−NoGo=GoPrior>1HitDiff−GoPriorNoGo|GoPrior>1HitDiff+GoPriorNoGo|

AMI(>1HitSame->1HitDiff): AMI>1HitSame−>1HitDiff=GoPrior>1HitSame−GoPrior>1HitDiff|GoPrior>1HitSame+GoPrior>1HitDiff|

where Go = mean whisker-evoked ΔF/F for current Go trials (on any whisker) with the specified trial history.

Attentional modulation of receptive fields

Population average 9-whisker receptive fields were constructed centered on the CW, and included the CW plus the 8 immediately adjacent whiskers. A separate population average receptive field was calculated for each trial history (Fig. 4A). These represent the average tuning of cells within each whisker column, following each trial history. To test for shifts in receptive fields by attention to specific whiskers, we first computed the center of mass (CoM) of the population average receptive fields for each history condition. CoM was calculated in a Cartesian CW-centered whisker space, as defined in Fig. S5B. The CW position is considered the origin in this space. Receptive field shifts associated with prior trial history were visualized as vectors from CoM measured after NoGo trials, to CoM measured after >1 prior hit to specific whiskers.

To quantify the receptive field shift (ΔRF CoM) for individual cells, we defined the attention axis as the axis connecting the CW position to the attended whisker position in the Cartesian CoM space. We projected the CoMPriorNoGo and CoMPrior>1Hit onto this axis, and computed the receptive field shift (ΔRF CoM) as the distance between these projected positions normalized to the distance from CW to attended whisker along the attention axis (Fig. S5C). Since not all whisker positions could be sampled for all cells across history conditions, ΔRF CoM was quantified only for the subset of cells for which at least 6 of the 9 whisker positions were sampled in both Prior >1 Hit and Prior NoGo conditions. Only response magnitudes at whisker positions sampled in both Prior >1 Hit and Prior NoGo conditions for any given cell contributed to the CoMs computed for that cell. The RF shift was computed separately for each attended whisker position that was sampled for a given cell, and then averaged across these attended whisker positions to generate a single RF shift metric for that cell.

Somatotopic profile of attentional modulation

To quantify the somatotopic profile of attentional modulation in S1 (Fig. 3K), we considered each of the 9 tested whiskers separately. For each whisker (termed the reference whisker), every cell was placed in a spatial bin representing its distance to the center of the reference whisker column in S1. Both columnar and septal-related cells were included. Mean whisker response magnitude was calculated, separated by trial history, in each bin. This was repeated for all 9 reference whiskers, and Fig. 3K shows the average response. Thus, the Prior NoGo trace reflects normal somatotopy, i.e., the normal point representation of an average whisker. The somatotopic profile of attentional modulation is evident as the difference between other history conditions and the Prior NoGo condition.

A similar binning procedure was used to calculate the somatotopic profile of AMI modulation across S1 columns (Fig. S4A).

Imaging analysis for VIP cells

Analysis of VIP cell responses was performed similarly to PYR cells, except that neuropil subtraction was not performed. To separate whisker-evoked VIP responses from slow trends in VIP baseline activity related to whisker motion, body motion and arousal in the ITI (Ramamurthy et al., 202344 and Fig. 2) we applied linear baseline detrending. For detrending, the median pre-stimulus baseline trace (in a 1.07-sec window) was calculated across Go trials (aligned to stimulus onset time) and NoGo trials (aligned to dummy piezo onset time). A line was fit to this median trace. This line was extrapolated and subtracted from each individual trial to yield the full peri-stimulus trace. This linear detrending was done separately for each history condition, due to the differences in pre-stimulus slopes for each condition. Note that linear detrending for prior same and prior different categories in each reward condition was identical. Analysis of history effects on VIP whisker responses (Fig. 7F–G) was performed after linear detrending. While VIP cells did not show whisker-specific attentional modulation (Fig. 7F–G), we verified that PYR cells still showed whisker-specific attentional boosting after detrending with the same methods (data not shown).

Neural decoding from population activity on single trials

We used a generalized linear model (GLM, Matlab ‘glmfit’) to predict the presence of a whisker stimulus from the trial-by-trial mean population activity of whisker-responsive L2/3 PYR cells in 2p imaging experiments. For each session, mean ΔF/F in the post-stimulus window (0 – 0.799 s) for each trial was calculated across all whisker-responsive ROIs that were simultaneously imaged in a single behavioral session. This single-trial population activity was used as the predictor of stimulus presence (current Go trial) or stimulus absence (current NoGo trial).

A separate decoder was fit from the data for each imaging session. We fitted a logistic regression model using leave-one-out cross-validation to predict the presence of any of the 9 whiskers, and tested it on hold-out data. Training/testing datasets were randomly re-sampled (majority class undersampled to match the minority class) to have identical numbers of trials within each response category (Hit, Miss, CR, FA), in order to remove bias. Decoder performance was measured as the average fraction of trials classified as containing a whisker stimulus (assessed over 25–50 iterations) and compared to performance for a decoder trained with shuffled trial labels.

For each field, we defined the fBW (field best whisker) as the whisker which evoked the numerically highest mean population ΔF/F. A single decoder was trained for each session to predict any whisker from training data containing Go trials from all whiskers, as well as NoGo trials. Decoder performance was assessed either for detecting any whisker, or just the fBW, or just non-fBW trials.

Extracellular recordings

For surgical preparation for mice used in extracellular recording experiments, methods were similar to that described above, except a lightweight chronic head post was affixed to the skull using cyanoacrylate glue and Metabond, and ISOI was performed to localize D-row (D1, D2, D3) barrel columns in S1. A 5-mm diameter glass coverslip (#1 thickness, CS-3R, Warner Instruments) was placed over the skull, sealed with Kwik-Cast silicone adhesive (World Precision Instruments) and dental cement. Mice recovered for 7 days prior to the start of behavioral training (~4 weeks before recording).

The day before recording, mice were anaesthetized with isoflurane, the protective coverslip was removed, and a craniotomy (~ 1.2 × 1.2 mm) was made over S1, centered over the D-row (D1, D2, D3) barrel columns localized by ISOI. A plastic ring was cemented around the craniotomy to create a recording chamber. During recording sessions, mice were anesthetized and positioned on the rig. A reference ground was attached inside the chamber. The craniotomy and reference wire were covered in a saline bath. Recordings were made with Neuropixels 1.0 probes using SpikeGLX software release v.20201024 (http://billkarsh.github.io/SpikeGLX/), Imec phase30 v3.31. Acute recordings were made in external reference mode with action potentials (AP) sampled at 30 kHz at 500x gain, and local field potential (LFP) sampled at 2.5 kHz at 250x gain. The AP band was common average referenced and band-pass filtered from 0.3 kHz to 6 kHz. The Neuropixels probe was mounted on a motorized stereotaxic micromanipulator (MP-285, Sutter Instruments) and advanced through the dura mater (except in cases where the dura had detached during the craniotomy). To reduce insertion-related mechanical tissue damage and to increase the single unit yield, the probe was lowered with a slow insertion speed of 1–2 μm/sec. The probe was first lowered to 700 μm, and a short 10-minute recording was conducted to map its location in S1. After identifying the columnar whisker for the recording penetration location, probe insertion continued until the final depth was reached. The probe was then left untouched for ~ 20 minutes. The craniotomy was sealed after probe insertion with silicone sealant (Kwik-Cast, World Precision Instruments) to prevent drying. Anesthesia was then discontinued, and mice were allowed to fully wake up before recording.

After recording was complete for the day, the probe was removed, the craniotomy was sealed with silicone sealant (Kwik-Cast, World Precision Instruments), and the recording chamber was sealed with a cover glass and a thin layer of dental cement. 3–4 sequential days of recording were performed in each mouse, with the probe located in a different whisker column in S1 on each day. On the final recording day, a Neuropixels probe was coated with red-fluorescent Dil (1,1′-Dioctadecyl-3,3,3′,3′-tetramethylindocarbocyanine perchlorate; Sigma-Aldrich) dissolved in 100% ethanol, 1–2mg/mL, which was allowed to partially dry on probe before probe insertion. The probe was briefly inserted to deposit DiI at recording and several fiducial sites. The mouse was euthanized and the brain was extracted, sectioned tangentially to the pial surface, and processed to stain for cytochrome oxidase (CO). DiI deposition sites in L2/3 were localized relative to column boundaries in CO from L4 (Fig. S6B).

The columnar location of each recording penetration was determined from DiI marks on the last recording day, plus relative locations of other penetrations based on reference images of surface vasculature and microdrive coordinates. The laminar depth of each recording was determined from current source density (CSD) and LFP power spectrum analysis, as described below.

Analysis of extracellular recording data

Spike Sorting

Spike sorting was performed by automatic clustering using Kilosort3 followed by manual curation using the ‘phy’ GUI (https://github.com/kwikteam/phy). Isolated units were manually inspected for mean spike waveform, stability over time, and inter-spike interval refractory period violations (we required that < 2% of intervals < 1.5 ms). Only well-isolated single units were analyzed. Single units were classified as regular-spiking or fast-spiking based on trough-to-peak duration of the spike waveform at the highest-amplitude recording channel, with a separation criterion of 0.45 ms95 (Fig. S6C). Only data from regular spiking units was analyzed here.

Layer assignment using CSD and LFP power spectrum analysis

To calculate the CSD for each recording penetration, the whisker stimulus-evoked local field potential (LFP, 500Hz low-pass) was calculated for each Neuropixels channel. LFP traces were normalized to correct for variations in channel impedance and were interpolated between channels (20 um site spacing, 1.6–2x interpolation) prior to calculating the second spatial derivative, which defines the CSD96. For visualization, CSDs were convolved with a 2D (depth × time) Gaussian, which revealed depth-restricted regions of current sources and sinks in response to each whisker stimulus. The L4-L5A boundary was defined from CSD as the zero-crossing between the most negative current sink (putative L4) and the next deeper current source (putative L5A) (Fig. S6F). To estimate the brain surface location (defining the top of L1), we computed the LFP power spectrum as a function of channel depth. A sharp increase in low-frequency LFP power marked the brain surface, which was used to verify correct selection of the L4-L5A boundary from CSD. Each cortical layer was then assigned boundary depths based on layer thicknesses reported in Lefort et al., 200997.

Whisker response quantification and attentional modulation for spike recordings

Firing rates for Go and NoGo trials were quantified in a 0.5 s window after stimulus onset (lick-free window). Units were classified as whisker-responsive or non-responsive by testing for greater firing rate on Go vs NoGo trials, using a permutation test, as for the 2p imaging data. To quantify whisker-evoked response magnitude, firing rate for each unit was z-scored relative to pre-stimulus baseline. Peristimulus time histograms (PSTHs) were constructed using 10-ms bins aligned to stimulus onset. AMI metrics were used to quantify attentional modulation of whisker-evoked spiking, and were calculated exactly as for 2p imaging data, but from z-scored whisker-evoked firing rate in the post-stimulus window.

Statistical analysis of summary data

Statistics were performed in Matlab. Sample size (n) and p-value for each analysis are reported in the figure panel, with the statistical test reported in the figure legend, or sometimes in the Results section. Permutation tests for the difference in means were used to assess differences in mean between two groups (referred to in brief as ‘permutation test’) or differences in the mean of a distribution relative to zero. Binomial exact test was used to assess the statistical significance of deviations from the expected distribution of observations into two categories. Summary data are reported as mean ± SEM, unless otherwise specified. Population means were compared using permutation tests, corrected for multiple comparisons (False Discovery Rate correction94) as needed. To compare population means with multiple subgroups, we first assessed whether a main effect was present using a permutation test, and then performed post hoc tests for pairwise means, correcting for multiple comparisons (with False Discovery Rate correction). We used a significance level (alpha) of 0.05. All statistical tests reported in the Results use n of cells, but we also verified that all major behavioral and neural effects were consistent across individual mice. We show individual mouse data for comparison with population data summaries. To account for inter-individual variability, we also verified all key results using a linear mixed effects model with the formula: Response Variable~Fixed Effect+(1∣Mouse Sex)+(1∣Mouse ID)

where mouse sex and mouse ID are modeled as random effects.

Supplementary Material

1

ACKNOWLEDGMENTS

This work was supported by NIH Grant 2 R01 NS092367 (D.E.F.) and NIH Grant 5 K99 NS129753-02 (D.L.R.).

DATA AVAILABILITY

The datasets generated during and/or analyzed during the current study are available from the corresponding author on reasonable request.

Figure 1. Recent reward history cues spatially specific attention for whisker touch.

A. Whisker detection task in head-fixed mice. B. Trial structure. Delay period was 0, 500, or 1000 ms in different mice (see Figure S1). Bottom, Trial types and trial outcomes. C. Example trial sequence showing interleaved Go and NoGo trials, with the identity of the deflected whisker on each Go trial chosen randomly. Water drop indicates successful Hit and reward. D. Mean effect of trial history on detection sensitivity (d-prime) on the current trial. Error bars are SEM across sessions. One or more prior Hits to the same whisker increased d’ for detecting that whisker on the current trial (‘attend toward’), while prior Hits to a different whisker decreased d’ (attend away). P-values: (1) from permutation test (prior same 1 Hit vs prior NoGo: p = 1e-4, prior same >1 Hit vs prior NoGo: p = 1e-4, prior different >1 Hit vs prior NoGo: p = 1e-4, prior same >1 Hit vs prior different >1 Hit: p = 1e-4), (2) from linear mixed effects model with fixed effect of prior history class (p = 8.7e-27). E. Same effect in a single example mouse. Left, the underlying effect on hit rate for 5 example sessions (thin lines) and mean ± SEM for all 34 sessions. Right, d’ for this mouse from 34 sessions. F. d’ did not increase when the prior trial was an unrewarded hit to the same whisker or earned very small reward (<4%: p = 0.49, >4%: p = 1e-4). P-values are for difference in d’ between prior same and prior different conditions (permutation test). G. Effect of >1 prior hits on Hit Rate and FA rate on the current trial. Connected symbols are the three history conditions for one mouse. Gray lines, receiver-operating characteristic (ROC) curves for different d’ levels. Prior hits to the same whisker drove a whisker-specific increase in hit rate over FA rate that improved d’. H. Behavioral shift in d’ for each individual mouse. I. Behavioral shifts in criterion (c) were also observed, but were less whisker-specific. J. Whisker-specific changes in d’ (Δ d’) and c (Δ c) for each mouse (number is mouse identity). We focus on Δ d’ to index whisker-specific attention. K. Spatial gradient of the attention effect. The facial position of the prior Hit whisker is plotted on the x-axis, relative to the current trial whisker. P-values are for difference from prior same (same row: p = 0.27, same arc: 1e-3, diagonal: 1e-4, further: 1e-4, permutation test). Bottom, each position shown schematically. L. Temporal profile of the whisker-specific attention effect (3–5 s: 5e-3, 5–7 s: 1e-4, 7–9 s: 1e-4, 9–11 s: 0.17, >11s: 0.15, permutation test), P-values are from permutation test. Bottom, Inter-Go-trial intervals sampled in the task. M. Flexible targeting of attention. x-axis represents identity of whisker in current Go trial, grouped into C, D, or E-row whiskers (left), or arcs 1, 2 or 3 (right). For all of these whiskers, prior reward history boosts detection in a whisker-specific way. Bottom, each row or arc shown schematically. P-values are for prior same vs prior different > 1 hit (p = 1e-4 for all whisker positions, permutation test). See also Figure S1.

Figure 2. Whisker motion, body motion, and arousal do not account for whisker-specific behavioral effects.

A. Mean whisker motion, platform motion (proxy for body motion), and pupil size traces, obtained using DeepLabCut from behavioral videos, across 43,647 trials, 74 sessions in 9 mice. Each panel shows the last 3 seconds of the intertrial interval (ITI) period after the prior trial, plus the stimulus period of the current trial. Traces and shading are mean ± SEM across all trials. Dashed line, stimulus onset (Go trials) or dummy piezo onset (NoGo trials). Prior trial identity is indicated by the Prior trial history. Δ whisker position and Δ platform position traces were zeroed to the mean of positions in a 0.25-sec window prior to stimulus onset (t = 0). Pupil size was normalized within each session to the mean pupil size across the whole session. Bottom row, lick histogram. B. Mean whisker movement, body movement, and pupil size change during the stimulus period. Prior Hits increased stimulus-evoked whisker and body motion on subsequent trials (Δ whisker motion, prior >1 hit same vs prior NoGo: p = 1e-4, prior >1 hit same vs >1 hit different: p = 0.39; Δ body motion, prior >1 hit same vs prior NoGo: p = 1e-4, prior >1 hit same vs >1 hit different: p = 0.39; Δ pupil area, prior >1 hit same vs prior NoGo: p = 0.49, prior >1 hit same vs >1 hit different: p = 0.53). p-values are for >1 Prior Hit vs Prior NoGo (top), and >1 Prior Hit Same vs > 1 Prior Hit Different (right) (permutation test). C. Design of Botox experiment. Behavior was assayed on an average of 12 sessions prior to Botox injection, and 7 sessions after Botox whisker paralysis. D. Reward history-dependent attention effect in each of the 4 mice tested, for standard sessions (before Botox, open symbols) and Botox sessions (filled symbols). M, Mouse numbers as in Fig. 1. Large points are mean ± SEM across mice. Conventions as in Fig. 1J. Whisker paralysis did not alter the mean whisker-specific d-prime effect or criterion effect (p = 0.88, comparing same vs. different shift in Δ d-prime for standard and Botox sessions). See also Figure S2.

Figure 3. Neural correlates of attentional capture in L2/3 pyramidal cells in S1.

A-B. Example imaging field in S1 centered on the D3 column. L2/3 PYR cells are expressing GCaMP6s. Scale bar = 100 μm. C. Example trials showing responses to the columnar whisker D3, but only weak responses to the surround whisker C3 that were strongly modulated by prior trial history. C3 responses increased following multiple prior hits to the C3 whisker (left), but not following other history conditions (right). D. Mean ΔF/F traces across all neurons in 70 sessions, separated by trial history. Each trace is the mean response to any Go whisker (solid) or on NoGo trials (dash). Shading shows SEM, which is often thinner than the mean trace. E. Quantification of whisker-evoked ΔF/F by current trial type (Go or NoGo) and by prior trial condition (prior same 1 Hit vs prior NoGo: p = 1e-4, prior same >1 Hit vs prior NoGo: p = 1e-4, prior different >1 Hit vs prior NoGo: p = 0.1, prior same >1 Hit vs prior different >1 Hit: p = 1e-4, permutation test). Bars are SEM. F. Mean modulation of ΔF/F response magnitude calculated by individual mouse. Bars are SEM. G. Correlation between history-based modulation of PYR whisker responses and behavioral d-prime, by mouse (r = 0.837, p = 6.9e-4). H. AMI>1HitSame-NoGo and AMI>1HitDiff-NoGo for each cell. Positive values denote greater response compared to the prior NoGo condition. I. AMI>1HitSame->1HitDiff for each cell. J. Boosting of whisker-evoked ΔF/F responses as a function of somatotopic offset between prior >1Hit whisker and the current trial whisker. Bars show d-prime (data from Fig. 1K, for the 6 mice used in PYR imaging) for comparison. K. Somatotopic organization of attentional capture. Left, the y-axis shows mean ΔF/F evoked by a reference whisker, as a function of cell position in S1 relative to the center of the reference whisker column. When calculated from Prior NoGo trials (thick black trace), this defines the classic point representation of a single whisker. This is boosted in prior same >1 hit trials, but not in prior different >1 hit trials, or prior same miss trials. Right, same data presented as difference from the Prior NoGo condition, to plot magnitude of modulation by trial history. See also Figure S3 and Figure S4.

Figure 4. Attentional cueing involves receptive field shifts toward attended whiskers.

A. Mean whisker-evoked ΔF/F traces for all whisker-responsive cells in all imaged columns. Center, mean ± SEM (across N = 5399 cells) for trials when prior trial was NoGo, separated by the identity of the current trial whisker. This reports the average whisker tuning curve for these neurons, in the absence of attentional cueing. Outer flanks, the whisker responses measured when prior trial history was >1Hit to the indicated attentional target whisker (thick purple trace is mean, thin traces show ± SEM). Purple fill is drawn between mean traces to aid visualization. r, u, c, d denote rostral, up, caudal, or down from CW. B. Mean tuning center-of-mass (CoM) when the prior trial was NoGo (black circle) vs after prior >1 Hit to each of the indicated whiskers as defined in panel A. CoM coordinate system is shown in Figure S4B. Vectors are color-coded for whether the target whisker was rostral, caudal, up, or down from the CW. C. Magnitude of CoM shift along the attention axis, as defined in Figure S4C. Negative values are shifts away from the attended whisker. The mean CoM shift was significantly greater than zero (p = 3.6e-3, permutation test). See also Figure S5.

Figure 5. Attentional cueing improves neural decoding of attended whiskers on single trials.

A. Neural decoder design. Left, example field showing L2/3 PYR cells tuned to different best whiskers intermixed in each column, consistent with prior studies86. Right, on each trial, population mean ΔF/F was calculated across all responsive cells. Leave-one-out cross-validation was used to fit a logistic regression predicting stimulus presence (any whisker Go trial) or absence (NoGo trial). B. Decoder performance determined from held-out trials. Each dot is one session. Decoder performance is low for any whisker, because many trials are Go trials for whiskers that are not strongly represented in the imaging field. Decoder performance for fBW trials is high, because this whisker is strongly represented in the imaging field. C-D. Mean ΔF/F (C) and fraction of trials decoded as a stimulus (Go) trial (D), separated by current trial type and outcome. Each dot is a session. All prior history conditions are combined in C-D panels (ΔF/F, HR vs Miss: p = 1e-4, FA vs CR: p = 1e-4; decoder, HR vs Miss: p = 1e-4, FA vs CR: p = 9e-4, permutation test). E. Mean decoder performance separated by trial history type, when current trial is any Go whisker (maroon or yellow) or a NoGo trial (gray dash). Solid lines with circle markers show decoder performance across all current trial outcomes, and dotted lines with asterisk markers show decoder performance on current hit trials only. (All trials, prior same >1 Hit vs prior NoGo: p = 1e-4, prior different >1 Hit vs prior NoGo: p = 0.21, prior same >1 Hit vs prior different >1 Hit: p = 1e-4; Hit trials only, prior same >1 Hit vs prior NoGo: p = 7e-3, prior different >1 Hit vs prior NoGo: p = 0.05, prior same >1 Hit vs prior different >1 Hit: p = 2e-3, permutation test). F. Same as (E), but for decoder performance when current trial is an fBW Go trial or a NoGo trial. G. Same as (E), but for decoder performance when current trial is a non-fBW Go trial or a NoGo trial. H. Summary of attentional modulation of decoding accuracy for the prior >1 Hit trial history condition. >1 prior Hit to a whisker improves single-trial decoding for non-fBW whiskers, but not for the fBW, in each field (non-fBW: p = 1e-4, fBW: p = 0.648). P-values are for prior same vs prior different > 1 hit (permutation test).

Figure 6. Attentional effects on extracellular single-unit spiking in S1.

A. Neuropixels recording during the whisker detection task. B. Mean PSTH for an example L2/3 RS unit recorded in the D1 whisker column, across all Go trials (D1, C2, delta, and gamma whiskers, top) and NoGo trials (bottom). Whisker-evoked responses were boosted when prior trial history was 1 or >1 Hit to the same whisker, and reduced when prior trial was >1 Hit to a different whisker. C, E, G. Layer-specific analysis for RS units of population mean PSTH (10 ms bins) for Go trials, by trial history. D, F, H. For the same units, population average whisker-evoked response on current Go trials (500-ms window) by trial history, or for current NoGo trials in an equivalent window (prior same >1 Hit vs prior different >1 Hit, L2/3: p = 6.9e-4, L4: 0.08, L5a/b: 1e-4, permutation test). I. Cumulative histogram of AMI>1HitSame-NoGo and AMI>1HitDiff-NoGo values for RS units by layer. J. Cumulative histogram of AMI>1HitSame->1HitDiff by layer. K. Mean AMI values by layer (AMI>1HitSame-NoGo vs AMI>1HitDiff-NoGo, L2/3: 0.03, L4: 0.65, L5a/b: 0.01; AMI>1HitSame->1HitDiff, L2/3: 0.01, L4:0.51, L5a/b: 0.06)). See also Figure S6.

Figure 7. L2/3 VIP cells carry a general arousal signal, but not a whisker-specific attentional signal.

A. Circuit model for potential VIP cell role in arousal, movement, and attentional modulation of sensory responses in sensory cortex. B. Example imaging field in L2/3 in a VIP-Cre;Ai162D mouse. Scale bar = 50 μm. C. Mean ΔF/F traces from L2/3 VIP cells across all mice and sessions during the last 3 seconds of the intertrial interval (ITI) period after the prior trial for prior rewarded (prior 1 Hit or prior >1 Hit trials) and prior unrewarded trials (prior NoGo or prior Miss). Traces are zeroed to the last 2 frames prior to stimulus onset. The declining baseline, evident on both prior rewarded (blue) and unrewarded (gray) trials, is due to VIP cell encoding of arousal, whisker motion, and body motion which all decline in the inter-trial interval as described in a previous study44, and shown in Fig. 2A. Prior rewarded trials show a steeper peak and steeper decline due to licking and reward consumption at the end of the prior trial. D. Mean ΔF/F traces across all VIP neurons in 103 sessions, separated by trial history. Each trace is the mean response to any Go whisker (solid) or on NoGo trials (dash). Shading shows SEM. Conventions as in Fig. 3D. E. Same data as in (D) after subtracting an extrapolation of the pre-stimulus baseline trend (see Methods). Whisker-evoked responses are modulated by prior hits on either the same or different whiskers, which is consistent with a global arousal effect, but not whisker-specific attentional capture. F. Quantification of trial history effects for L2/3 VIP cells (prior same 1 Hit vs prior NoGo: p = 0.11, prior same >1 Hit vs prior NoGo: p = 1e-3, prior different >1 Hit vs prior NoGo: p = 1e-3, prior same >1 Hit vs prior different >1 Hit: p = 0.54, permutation test). Conventions as in Fig. 3E. G. AMI>1HitSame-NoGo and AMI>1HitDiff-NoGo for each cell. Positive values denote greater response compared to the prior NoGo condition. H. AMI>1HitSame->1HitDiff for each cell. See also Figure S7.

COMPETING INTERESTS

The authors declare no competing interests.
==== Refs
REFERENCES

1. Maunsell J.H.R. Neuronal Mechanisms of Visual Attention. Annu Rev Vis Sci 1 , 373–391 (2015).28532368
2. Hromádka T. & Zador A.M. Toward the mechanisms of auditory attention. Hear Res 229 , 180–185 (2007).17307316
3. Gallace A. & Spence C. 147 Tactile attention. in In Touch with the Future: The sense of touch from cognitive neuroscience to virtual reality 0 (Oxford University Press, 2014).
4. Nobre A.C. & Stokes M.G. Premembering Experience: A Hierarchy of Time-Scales for Proactive Attention. Neuron 104 , 132–146 (2019).31600510
5. Awh E. , Belopolsky A.V. & Theeuwes J. Top-down versus bottom-up attentional control: a failed theoretical dichotomy. Trends in cognitive sciences 16 , 437–443 (2012).22795563
6. Theeuwes J. & Failing M. Attentional Selection: Top-Down, Bottom-Up and History-Based Biases. (Cambridge University Press, Cambridge, 2020).
7. Meyer K.N. , Sheridan M.A. & Hopfinger J.B. Reward history impacts attentional orienting and inhibitory control on untrained tasks. Atten Percept Psychophys 82 , 3842–3862 (2020).32935290
8. Failing M. & Theeuwes J. Selection history: How reward modulates selectivity of visual attention. Psychonomic Bulletin & Review 25 , 514–538 (2018).28986770
9. Anderson B.A. & Britton M.K. Selection history in context: Evidence for the role of reinforcement learning in biasing attention. Atten Percept Psychophys 81 , 2666–2672 (2019).31309530
10. Anderson B.A. , Kim H. , Kim A.J. , Liao M.-R. , Mrkonja L. , Clement A. , and Grégoire L. The past, present, and future of selection history. Neuroscience & Biobehavioral Reviews 130 , 326–350. 10.1016/j.neubiorev.2021.09.004. (2021).34499927
11. Chelazzi L. , Perlato A. , Santandrea E. & Della Libera C. Rewards teach visual selective attention. Vision Research 85 , 58–72 (2013).23262054
12. Anderson B.A. A value-driven mechanism of attentional selection. J Vis 13 (2013).
13. Anderson B.A. Neurobiology of value-driven attention. Current Opinion in Psychology 29 , 27–33 (2019).30472540
14. Anderson B.A. , Laurent P.A. & Yantis S. Value-driven attentional capture. Proceedings of the National Academy of Sciences 108 , 10367 (2011).
15. Addleman D.A. & Jiang Y.V. Experience-Driven Auditory Attention. Trends in Cognitive Sciences 23 , 927–937 (2019).31521482
16. Chen D. & Hutchinson J.B. What Is Memory-Guided Attention? How Past Experiences Shape Selective Visuospatial Attention in the Present. Curr Top Behav Neurosci 41 , 185–212 (2019).30584646
17. Lynn J. , and Shin M. Strategic top-down control versus attentional bias by previous reward history. Atten Percept Psychophys 77 , 2207–2216. 10.3758/s13414-015-0939-9. (2015).26041273
18. Noudoost B. , Chang M.H. , Steinmetz N.A. & Moore T. Top-down control of visual attention. Curr Opin Neurobiol 20 , 183–190 (2010).20303256
19. Matteucci G. , Cortical sensory processing across motivational states during goal-directed behavior. Neuron 110 , 4176–4193.e4110 (2022).36240769
20. Luo T.Z. & Maunsell J.H.R. Attention can be subdivided into neurobiological components corresponding to distinct behavioral effects. Proceedings of the National Academy of Sciences 116 , 26187–26194 (2019).
21. Krauzlis R.J. , Wang L. , Yu G. & Katz L.N. What is attention? Wiley Interdiscip Rev Cogn Sci 14 , e1570 (2023).34169668
22. Luo T.Z. & Maunsell J.H. Neuronal Modulations in Visual Cortex Are Associated with Only One of Multiple Components of Attention. Neuron 86 , 1182–1188 (2015).26050038
23. Luo T.Z. & Maunsell J.H.R. Attentional Changes in Either Criterion or Sensitivity Are Associated with Robust Modulations in Lateral Prefrontal Cortex. Neuron 97 , 1382–1393.e1387 (2018).29503191
24. Peck C.J. , Jangraw D.C. , Suzuki M. , Efem R. & Gottlieb J. Reward modulates attention independently of action value in posterior parietal cortex. J Neurosci 29 , 11182–11191 (2009).19741125
25. Desimone R. & Duncan J. Neural mechanisms of selective visual attention. Annu Rev Neurosci 18 , 193–222 (1995).7605061
26. Cohen M.R. & Maunsell J.H.R. A Neuronal Population Measure of Attention Predicts Behavioral Performance on Individual Trials. The Journal of Neuroscience 30 , 15241 (2010).21068329
27. Reimer J. , Pupil Fluctuations Track Fast Switching of Cortical States during Quiet Wakefulness. Neuron 84 , 355–362 (2014).25374359
28. Reimer J. , Pupil fluctuations track rapid changes in adrenergic and cholinergic activity in cortex. Nature Communications 7 , 13289 (2016).
29. Mathis A. , DeepLabCut: markerless pose estimation of user-defined body parts with deep learning. Nature Neuroscience 21 , 1281–1289 (2018).30127430
30. Womelsdorf T. , Anton-Erxleben K. , Pieper F. & Treue S. Dynamic shifts of visual receptive fields in cortical area MT by spatial attention. Nature Neuroscience 9 , 1156–1160 (2006).16906153
31. Womelsdorf T. , Anton-Erxleben K. & Treue S. Receptive field shift and shrinkage in macaque middle temporal area through attentional gain modulation. J Neurosci 28 , 8934–8944 (2008).18768687
32. Anton-Erxleben K. , Stephan V.M. & Treue S. Attention reshapes center-surround receptive field structure in macaque cortical area MT. Cereb Cortex 19 , 2466–2478 (2009).19211660
33. He S. , Cavanagh P. & Intriligator J. Attentional resolution and the locus of visual awareness. Nature 383 , 334–337 (1996).8848045
34. Yang H. , Kwon S.E. , Severson K.S. & O’Connor D.H. Origins of choice-related activity in mouse somatosensory cortex. Nat Neurosci 19 , 127–134 (2016).26642088
35. Yamashita T. & Petersen C. Target-specific membrane potential dynamics of neocortical projection neurons during goal-directed behavior. Elife 5 (2016).
36. Jun J.J. , Fully integrated silicon probes for high-density recording of neural activity. Nature 551 , 232–236 (2017).29120427
37. Lee S. , Kruglikov I. , Huang Z.J. , Fishell G. & Rudy B. A disinhibitory circuit mediates motor integration in the somatosensory cortex. Nat Neurosci 16 , 1662–1670 (2013).24097044
38. Fu Y. , A cortical circuit for gain control by behavioral state. Cell 156 , 1139–1152 (2014).24630718
39. Zhang S. , Selective attention. Long-range and local circuits for top-down modulation of visual cortex processing. Science 345 , 660–665 (2014).25104383
40. Batista-Brito R. , Zagha E. , Ratliff J.M. & Vinck M. Modulation of cortical circuits by top-down processing and arousal state in health and disease. Current opinion in neurobiology 52 , 172–181 (2018).30064117
41. Xia R. , Chen X. , Engel T.A. & Moore T. Common and distinct neural mechanisms of attention. Trends Cogn Sci 28 , 554–567 (2024).38388258
42. Myers-Joseph D. , Wilmes K.A. , Fernandez-Otero M. , Clopath C. & Khan A.G. Disinhibition by VIP interneurons is orthogonal to cross-modal attentional modulation in primary visual cortex. Neuron 112 , 628–645.e627 (2024).38070500
43. Rahmatullah N. , Hypersensitivity to Distractors in Fragile X Syndrome from Loss of Modulation of Cortical VIP Interneurons. The Journal of Neuroscience 43 , 8172 (2023).37816596
44. Ramamurthy D.L. , VIP interneurons in sensory cortex encode sensory and action signals but not direct reward signals. Curr Biol 33 , 3398–3408.e3397 (2023).37499665
45. Musall S. , Kaufman M.T. , Juavinett A.L. , Gluf S. & Churchland A.K. Single-trial neural dynamics are dominated by richly varied movements. Nature Neuroscience 22 , 1677–1686 (2019).31551604
46. Schmidt F. , Haberkamp A. & Schmidt T. Dos and don’ts in response priming research. Adv Cogn Psychol 7 , 120–131 (2011).22253674
47. Akrami A. , Kopec C.D. , Diamond M.E. & Brody C.D. Posterior parietal cortex represents sensory history and mediates its effects on behaviour. Nature 554 , 368–372 (2018).29414944
48. Hachen I. , Reinartz S. , Brasselet R. , Stroligo A. & Diamond M.E. Dynamics of history-dependent perceptual judgment. Nature Communications 12 , 6036 (2021).
49. Waiblinger C. , McDonnell M.E. , Borden P.Y. & Stanley G.B. Emerging experience-dependent dynamics in primary somatosensory cortex reflect behavioral adaptation. bioRxiv, 2021.2001.2029.428886 (2021).
50. Lak A. , Reinforcement biases subsequent perceptual decisions when confidence is low, a widespread behavioral phenomenon. eLife 9 , e49834 (2020).32286227
51. Boboeva V. , Pezzotta A. , Clopath C. & Akrami A. Unifying network model links recency and central tendency biases in working memory. eLife 12 , RP86725 (2024).
52. Cicchini G.M. , Mikellidou K. & Burr D.C. Serial Dependence in Perception. Annual Review of Psychology 75 , 129–154 (2024).
53. Samonds J.M. , Lieberman S. & Priebe N.J. Motion Discrimination and the Motion Aftereffect in Mouse Vision. eNeuro 5 (2018).
54. Urai A.E. , de Gee J.W. , Tsetsos K. & Donner T.H. Choice history biases subsequent evidence accumulation. eLife 8 , e46331 (2019).31264959
55. Ashwood Z.C. , Mice alternate between discrete strategies during perceptual decision-making. Nat Neurosci 25 , 201–212 (2022).35132235
56. Roy N.A. , Bak J.H. , Akrami A. , Brody C.D. & Pillow J.W. Extracting the dynamics of behavior in sensory decision-making experiments. Neuron 109 , 597–610.e596 (2021).33412101
57. Reynolds J.H. , Pasternak T. & Desimone R. Attention increases sensitivity of V4 neurons. Neuron 26 , 703–714 (2000).10896165
58. Briggs F. , Mangun G.R. & Usrey W.M. Attention enhances synaptic efficacy and the signal-to-noise ratio in neural circuits. Nature 499 , 476–480 (2013).23803766
59. Cohen M.R. & Maunsell J.H. Attention improves performance primarily by reducing interneuronal correlations. Nature neuroscience 12 , 1594–1600 (2009).19915566
60. Luck S.J. , Chelazzi L. , Hillyard S.A. & Desimone R. Neural Mechanisms of Spatial Selective Attention in Areas V1, V2, and V4 of Macaque Visual Cortex. Journal of Neurophysiology 77 , 24–42 (1997).9120566
61. Silver M.A. , Ress D. & Heeger D.J. Topographic maps of visual spatial attention in human parietal cortex. J Neurophysiol 94 , 1358–1371 (2005).15817643
62. Tootell R.B. , The retinotopy of visual spatial attention. Neuron 21 , 1409–1422 (1998).9883733
63. Woldorff M.G. , Retinotopic organization of early visual spatial attention effects as revealed by PET and ERPs. Hum Brain Mapp 5 , 280–286 (1997).20408229
64. Bollimunta A. , Bogadhi A.R. & Krauzlis R.J. Comparing frontal eye field and superior colliculus contributions to covert spatial attention. Nat Commun 9 , 3553 (2018).30177726
65. Staiger J.F. & Petersen C.C.H. Neuronal Circuits in Barrel Cortex for Whisker Sensory Perception. Physiol Rev 101 , 353–415 (2021).32816652
66. Engel T.A. , Selective modulation of cortical state during spatial attention. Science 354 , 1140–1144 (2016).27934763
67. Hattori R. , Danskin B. , Babic Z. , Mlynaryk N. & Komiyama T. Area-Specificity and Plasticity of History-Dependent Value Coding During Learning. Cell 177 , 1858–1872.e1815 (2019).31080067
68. Hwang E.J. , Dahlen J.E. , Mukundan M. & Komiyama T. History-based action selection bias in posterior parietal cortex. Nature Communications 8 , 1242 (2017).
69. Findling C. , Brain-wide representations of prior information in mouse decision-making. bioRxiv, 2023.2007.2004.547684 (2023).
70. Miller K.J. , Botvinick M.M. & Brody C.D. Value representations in the rodent orbitofrontal cortex drive learning, not choice. eLife 11 , e64575 (2022).35975792
71. Marmor O. , Pollak Y. , Doron C. , Helmchen F. & Gilad A. History information emerges in the cortex during learning. eLife 12 , e83702 (2023).37921842
72. Banerjee A. , Value-guided remapping of sensory cortex by lateral orbitofrontal cortex. Nature 585 , 245–250 (2020).32884146
73. Chéreau R. , Dynamic perceptual feature selectivity in primary somatosensory cortex upon reversal learning. Nature Communications 11 , 3245 (2020).
74. Petty G.H. & Bruno R.M. Attentional modulation of secondary somatosensory and visual thalamus of mice. eLife. 13 :RP97188 (2024).
75. Kastner S. & Arcaro M.J. The Thalamus in Attention. in The Thalamus (ed. Halassa M.M. ) 324–339 (Cambridge University Press, Cambridge, 2022).
76. Pi H.J. , Cortical interneurons that specialize in disinhibitory control. Nature 503 , 521–524 (2013).24097352
77. Hu F. & Dan Y. An inferior-superior colliculus circuit controls auditory cue-directed visual spatial attention. Neuron 110 , 109–119.e103 (2022).34699777
78. Speed A. & Haider B. Probing mechanisms of visual spatial attention in mice. Trends in Neurosciences (2021).
79. Speed A. , Del Rosario J. , Mikail N. & Haider B. Spatial attention enhances network, cellular and subthreshold responses in mouse visual cortex. Nature Communications 11 , 505 (2020).
80. Kanamori T. & Mrsic-Flogel T.D. Independent response modulation of visual cortical neurons by attentional and behavioral states. Neuron 110 , 3907–3918.e3906 (2022).36137550
81. McBride E.G. , Lee S.J. & Callaway E.M. Local and Global Influences of Visual Spatial Selection and Locomotion in Mouse Primary Visual Cortex. Curr Biol 29 , 1592–1605.e1595 (2019).31056388
82. Lehnert J. , Visual attention to features and space in mice using reverse correlation. Current Biology 33 , 3690–3701.e3694 (2023).37611588
83. You W.-K. & Mysore S.P. Endogenous and exogenous control of visuospatial selective attention in freely behaving mice. Nature Communications 11 , 1986 (2020).
84. Goldstein S. , Wang L. , McAlonan K. , Torres-Cruz M. & Krauzlis R.J. Stimulus-driven visual attention in mice. Journal of Vision 22 , 11–11 (2022).
85. Drew P.J. & Feldman D.E. Intrinsic Signal Imaging of Deprivation-Induced Contraction of Whisker Representations in Rat Somatosensory Cortex. Cerebral Cortex 19 , 331–348 (2009).18515797
86. Wang H.C. , LeMessurier A.M. & Feldman D.E. Tuning instability of non-columnar neurons in the salt-and-pepper whisker map in somatosensory cortex. Nature Communications 13 , 6611 (2022).
87. Landers M. , Pytte C. & Zeigler H.P. Reversible blockade of rodent whisking: Botulinum toxin as a tool for developmental studies. Somatosensory & Motor Research 19 , 358–363 (2002).12590837
88. Yang H. , Kwon S.E. , Severson K.S. & O’Connor D.H. Origins of choice-related activity in mouse somatosensory cortex. Nature Neuroscience 19 , 127–134 (2016).26642088
89. Green D.M. & Swets J.A. Signal detection theory and psychophysics (John Wiley, Oxford, England, 1966).
90. Pronneke A. , Characterizing VIP Neurons in the Barrel Cortex of VIPcre/tdTomato Mice Reveals Layer-Specific Differences. Cereb Cortex 25 , 4854–4868 (2015).26420784
91. Schindelin J. , Fiji: an open-source platform for biological-image analysis. Nature methods 9 , 676–682 (2012).22743772
92. LeMessurier A.M. (2019). Imaging_analysis_pipeline. GitHub 5a21c9f. https://github.com/alemessurier/imaging_analysis_pipeline
93. Guizar-Sicairos M. , Thurman S.T. & Fienup J.R. Efficient subpixel image registration algorithms. Opt. Lett. 33 , 156–158 (2008).18197224
94. Benjamini Y. & Hochberg Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B (Methodological) 57 , 289–300 (1995).
95. Barthó P. , Characterization of neocortical principal cells and interneurons by network interactions and extracellular features. J Neurophysiol 92 , 600–608 (2004).15056678
96. Freeman J.A. & Nicholson C. Experimental optimization of current source-density technique for anuran cerebellum. J Neurophysiol 38 , 369–382 (1975).165272
97. Lefort S. , Tomm C. , Floyd Sarria J.C. & Petersen C.C.H. The Excitatory Neuronal Network of the C2 Barrel Column in Mouse Primary Somatosensory Cortex. Neuron 61 , 301–316 (2009).19186171
