Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 5:00 PM ET until 8:00 PM ET on Friday, September 11 due to maintenance. We apologize for the inconvenience.


Title: Emergence of Emotion Selectivity in Deep Neural Networks Trained to Recognize Visual Objects
Recent neuroimaging studies have shown that the visual cortex plays an important role in representing the affective significance of visual input. The origin of these affect-specific visual representations is debated: they are intrinsic to the visual system versus they arise through reentry from frontal emotion processing structures such as the amygdala. We examined this problem by combining convolutional neural network (CNN) models of the human ventral visual cortex pre-trained on ImageNet with two datasets of affective images. Our results show that in all layers of the CNN models, there were artificial neurons that responded consistently and selectively to neutral, pleasant, or unpleasant images and lesioning these neurons by setting their output to zero or enhancing these neurons by increasing their gain led to decreased or increased emotion recognition performance respectively. These results support the idea that the visual system may have the intrinsic ability to represent the affective significance of visual input and suggest that CNNs offer a fruitful platform for testing neuroscientific theories.  more » « less
Award ID(s):
2318984 1908299
PAR ID:
10517490
Author(s) / Creator(s):
; ; ;
Editor(s):
Wei, Xue-Xin
Publisher / Repository:
Public Library of Science
Date Published:
Journal Name:
PLOS Computational Biology
Volume:
20
Issue:
3
ISSN:
1553-7358
Page Range / eLocation ID:
e1011943
Subject(s) / Keyword(s):
Emotion Selectivity, Deep Neural Network
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Abstract The perception of opportunities and threats in complex visual scenes represents one of the main functions of the human visual system. The underlying neurophysiology is often studied by having observers view pictures varying in affective content. While deep neural networks (DNNs) have shown promise in modeling visual recognition of objects, their capacity to model visual affective processing remains to be better understood. In this study, we proposed a biologically inspired deep neural network model, referred to as the Visual Cortex Amygdala (VCA) model, for this purpose. The model integrates a vision transformer module for visual encoding and an amygdala-mimetic module that incorporates an anatomical hierarchy and self-attention-based computational mechanisms for affective decoding. We evaluated the model along three dimensions: (1) predictive accuracy for emotional valence and arousal, (2) representational alignment with human amygdala activity, and (3) internal organization of emotion representation within the model. The results showed that (1) the model can predict with high accuracy human emotion ratings on 1,182 images from the International Affective Picture System (IAPS) dataset (valence: r ≈ 0.9; arousal: r ≈ 0.7), (2) the model’s internal representations aligned with functional Magnetic Resonance Imaging (fMRI) data from the human amygdala, and (3) at the single model neuron level, the amygdala module evolved emotion selectivity, and at the model neural population level, deeper layers of the amygdala module developed representational geometry progressively more aligned with affective dimensions. We also explored the effect of visual encoding and the effect of structure and computational mechanisms on emotional assessment. 
    more » « less
  2. Recent fMRI studies in human subjects have found affect-specific neural representations of emotional scenes in early visual cortex. The origin of these representations is debated. One group of hypotheses suggests that these representations result from reentrant feedback from anterior emotion-modulating structures (e.g., the amygdala), whereas another group of hypotheses states that sensory cortex, including retinotopic visual cortex, may itself code for the emotional qualities of visual stimuli, without the necessity for feedback processing. We examined this problem by employing a neural encoding model that can generate synthetic fMRI responses to natural images in early visual cortex. The model works by linearly mapping features extracted by convolutional neural networks onto voxel-wise BOLD responses in different visual areas and is trained on the Natural Scenes Dataset. Dividing the images in the International Affective Picture System into three broad categories: pleasant, neutral and unpleasant, we found that in early visual cortex, the neural patterns evoked by the emotional images cannot be decoded from that evoked by the neutral images, in contrast with the findings from recent fMRI studies in human subjects. Because the model-generated responses are free from emotion-modulated reentrant feedback, this finding can be seen as lending support to the reentry hypothesis. Interestingly, when face stimuli from the AffectNet were shown to the model, the neural patterns evoked by emotional faces in early visual cortex can be significantly decoded from that evoked by neutral faces, suggesting that the early visual cortex may contribute differently to the emotional processing of faces versus scenes. 
    more » « less
  3. Hedonic valence, the intrinsic pleasantness or unpleasantness of an experience, is fundamental to human psychological functioning. Yet, how valence is represented in the brain remains an open question. Functional MRI studies have demonstrated that the brain encodes both positive and negative valence, but this evidence largely stems from experiments using simplified, controlled stimuli, such as images, sounds, or words. As a result, it remains unclear how valence is processed during rich, naturalistic experiences that more closely reflect real life. In addition, most studies adopt a single statistical model, raising concerns about the robustness of their findings. This study used a formal voxel-wise Bayesian model selection approach to test alternative statistical models supporting Bipolarity, Valence-General, and Bivalence hypotheses to identify the most optimal model of valence representation during narrative listening. Our results provide evidence for the Bipolar model. We identified distributed brain regions that selectively encode valence as a bipolar continuum (negative to positive) during narrative comprehension, including classical emotion-related hubs such as ventromedial prefrontal cortex, as well as regions not traditionally associated with emotion processing, such as inferior occipital cortex, supramarginal cortex, inferior frontal cortex, and middle cingulate. Regions selectively encoding arousal and those broadly responsive to both valence and arousal were also identified. These findings highlight the importance of using formal model comparison and naturalistic paradigms in affective neuroscience, advancing our understanding of how valence is represented in the brain during real-world experiences. 
    more » « less
  4. Abstract Visual neurons respond to a vast range of images, from textures to objects, but the rules linking these responses remain unclear. Although tuning to simple features is well established in the primary visual cortex, this framework breaks down in higher areas, where neurons encode diverse and unpredictable features. To ask what features neurons prioritize, we used generative models (deep networks that synthesize new images from a learned latent space), allowing neurons in V1, V4 and the posterior inferotemporal cortex (PIT) to guide image synthesis through closed-loop optimization. We compared models that emphasize texture versus those that emphasize object structure. Although V1 and V4 aligned more strongly with texture-based spaces, many PIT neurons responded equally well to both types of optimized images, revealing a focus on shared local motifs rather than whole-object templates, and this alignment to objects emerged later in their response. These findings reveal coding principles across the ventral stream and clarify the limits of current vision models. 
    more » « less
  5. Humans and other primates can robustly report whether they've seen specific images before, even when those images are extremely similar to ones they've previously seen. Multiple lines of evidence suggest that pattern separation computations in the hippocampus (HC) contribute to this behavior by shaping the fidelity of visual memory. However, unclear is whether HC uniquely determines memory fidelity or whether computations in other brain areas also contribute. To investigate, we recorded neural signals from inferotemporal cortex (ITC) and HC of two rhesus monkeys (1 male, 1 female) as they performed a memory task in which they judged whether images were novel or exactly repeated in the presence of visually similar lure images with a range of visual similarities. We found behavioral evidence for sharpening, reflected as memory performance that was nonlinearly transformed relative to a benchmark defined by visual representations in ITC. As expected, we found that behavioral sharpening aligned with visual memory representations in HC. Surprisingly, and unaccounted for by HC pattern separation proposals, we also found neural correlates of behavioral sharpening reflected in ITC. These results, coupled with further analysis of the data, suggest that ITC contributes to shaping the fidelity of visual memory in the transformation from visual processing to memory storage and signaling. Significance StatementVisual recognition memories are stored with remarkable visual fidelity, allowing humans and other primates to distinguish images they have encountered from visually similar images they have not. This fidelity has long been attributed to computations in the hippocampus that sharpen visual representations before memory storage (pattern separation). Unclear is how this proposal aligns with other evidence that visual memories are stored within high-level visual cortex itself, before signals reach the hippocampus. Here we demonstrate that, like the hippocampus, inferotemporal cortex also reflects sharpened visual memory representations, suggesting that visual cortex contributes to shaping the visual fidelity of visual memory. 
    more » « less