Abstract Though the right hemisphere has been implicated in talker processing, it is thought to play a minimal role in phonetic processing, at least relative to the left hemisphere. Recent evidence suggests that the right posterior temporal cortex may support learning of phonetic variation associated with a specific talker. In the current study, listeners heard a male talker and a female talker, one of whom produced an ambiguous fricative in /s/-biased lexical contexts (e.g., epi?ode) and one who produced it in /∫/-biased contexts (e.g., friend?ip). Listeners in a behavioral experiment (Experiment 1) showed evidence of lexically guided perceptual learning, categorizing ambiguous fricatives in line with their previous experience. Listeners in an fMRI experiment (Experiment 2) showed differential phonetic categorization as a function of talker, allowing for an investigation of the neural basis of talker-specific phonetic processing, though they did not exhibit perceptual learning (likely due to characteristics of our in-scanner headphones). Searchlight analyses revealed that the patterns of activation in the right superior temporal sulcus (STS) contained information about who was talking and what phoneme they produced. We take this as evidence that talker information and phonetic information are integrated in the right STS. Functional connectivity analyses suggested that the process of conditioning phonetic identity on talker information depends on the coordinated activity of a left-lateralized phonetic processing system and a right-lateralized talker processing system. Overall, these results clarify the mechanisms through which the right hemisphere supports talker-specific phonetic processing.
more »
« less
Revisiting the left ear advantage for phonetic cues to talker identification
Previous research suggests that learning to use a phonetic property [e.g., voice-onset-time, (VOT)] for talker identity supports a left ear processing advantage. Specifically, listeners trained to identify two “talkers” who only differed in characteristic VOTs showed faster talker identification for stimuli presented to the left ear compared to that presented to the right ear, which is interpreted as evidence of hemispheric lateralization consistent with task demands. Experiment 1 ( n = 97) aimed to replicate this finding and identify predictors of performance; experiment 2 ( n = 79) aimed to replicate this finding under conditions that better facilitate observation of laterality effects. Listeners completed a talker identification task during pretest, training, and posttest phases. Inhibition, category identification, and auditory acuity were also assessed in experiment 1. Listeners learned to use VOT for talker identity, which was positively associated with auditory acuity. Talker identification was not influenced by ear of presentation, and Bayes factors indicated strong support for the null. These results suggest that talker-specific phonetic variation is not sufficient to induce a left ear advantage for talker identification; together with the extant literature, this instead suggests that hemispheric lateralization for talker-specific phonetic variation requires phonetic variation to be conditioned on talker differences in source characteristics.
more »
« less
- Award ID(s):
- 1827591
- PAR ID:
- 10387943
- Date Published:
- Journal Name:
- The Journal of the Acoustical Society of America
- Volume:
- 152
- Issue:
- 5
- ISSN:
- 0001-4966
- Page Range / eLocation ID:
- 3107 to 3123
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
Speech categories are defined by multiple acoustic dimensions and their boundaries are generally fuzzy and ambiguous in part because listeners often give differential weighting to these cue dimensions during phonetic categorization. This study explored how a listener's perception of a speaker's socio-indexical and personality characteristics influences the listener's perceptual cue weighting. In a matched-guise study, three groups of listeners classified a series of gender-neutral /b/-/p/ continua that vary in VOT and F0 at the onset of the following vowel. Listeners were assigned to one of three prompt conditions (i.e., a visually male talker, a visually female talker, or audio-only) and rated the talker in terms of vocal (and facial, in the visual prompt conditions) gender prototypicality, attractiveness, friendliness, confidence, trustworthiness, and gayness. Male listeners and listeners who saw a male face showed less reliance on VOT compared to listeners in the other conditions. Listeners' visual evaluation of the talker also affected their weighting of VOT and onset F0 cues, although the effects of facial impressions differ depending on the gender of the listener. The results demonstrate that individual differences in perceptual cue weighting are modulated by the listener's gender and his/her subjective evaluation of the talker. These findings lend support for exemplar-based models of speech perception and production where socio-indexical features are encoded as a part of the episodic traces in the listeners' mental lexicon. This study also shed light on the relationship between individual variation in cue weighting and community-level sound change by demonstrating that VOT and onset F0 co-variation in North American English has acquired a certain degree of socio-indexical significance.more » « less
-
Normal-hearing older listeners are as accurate as younger listeners when perceiving native English words in quiet despite challenges in temporal processing. Older listeners may compensate for the declined use of fine-grained temporal cues by reducing the weight of temporal cues (VOT) and increase the reliance on other acoustic correlates (F0) of the sound contrast. In Experiment 1, younger (age 18–25) and older (age 55–65) normal-hearing listeners participate in an online 2AFC identification task with /d/-/t/ contrast varying in both VOT and F0. We predict that, while both younger and older listeners rely more on VOT than on F0, older listeners, because of their reduced temporal processing abilities, rely on F0 to a larger degree than younger listeners. Temporal processing not only involves local durational cues of the target segments, but also global contextual cues such as speaking rate. In Experiment 2, the same listeners complete another online 2AFC identification task with /dɑ/-/tɑ/ syllables that vary in VOT and vowel duration (short versus long). We predict that older listeners exhibit a smaller shift in the /d/-/t/ category boundary between the long and short vowel durations than younger listeners since older adults are less sensitive to contextual temporal information.more » « less
-
Holliday, Jeff; Lee-KIm, Sang-Im; Cho, Taehong (Ed.)Listeners use their knowledge about a talker to guide speech perception. In the present study, we manipulated both the familiarity and predictability of this talker-specific knowledge. Twenty-two listeners from western Pennsylvania completed an audio-visual lexical decision task while EEG was recorded. Critically, listeners were introduced to the talkers beforehand, with short videos establishing the talker’s U.S. English accent identity: Mainstream, more similar to the participants themselves; Southern, a relatively less familiar variety; and Unpredictable, which switched between Mainstream and Southern accents. During the task, listeners watched these talkers produce tokens with accents that either aligned with or violated their accent identity. Behaviorally, alignment between talker and token accent influenced performance on Southern-accented tokens, while accuracy on Mainstream-accented tokens was nearly at ceiling regardless of talker accent identity. Neurally, accent familiarity had the largest effect pre-speech (in response to the visual presentation of the talker), while both predictability and familiarity impacted processing of the speech signal itself. Overall, our results suggest that listeners use information about a talker’s accent, even when it is unfamiliar or unpredictable, to alter their listening strategies for more successful speech recognition.more » « less
-
This dissertation compared speech perception across younger and older normal-hearing adults. We ask four research questions to assess acoustic cue weighting and the role of contextual information (lexical information and speaking rate) in speech perception. Experiment 1 tested for age-related changes in cue-weighting. The absolute weights showed that older listeners relied on both VOT and F0 more than younger listeners, and some listeners’ reliance on VOT correlated with inhibitory control when perceiving the /d/-/t/ contrast. The relative weights suggested that older listeners relied on VOT less and F0 more than younger listeners. Experiment 2 tested whether younger and older listeners used different cue-weighting adjustment strategies in distributional learning. Both older and younger listeners adjusted their reliance on acoustic cues when the primary acoustic cue (VOT) became ambiguous. Older listeners adjusted F0 to a greater degree than younger listeners, while younger listeners adjusted VOT more than older listeners. With a lexically-guided learning paradigm, Experiment 3 explored if younger and older adults differed in their use of lexical information when learning to map acoustic tokens that were ambiguous. Older and younger listeners utilized the lexical context to the same extent. In Experiment 4, the contextual effect of speaking rate was examined by embedding voicing contrasts in short and long syllables and presenting these syllables to younger and older listeners. Older and younger listeners compensated for variation in speaking rate in a similar manner as younger listeners. The findings in perceptual learning demonstrate perceptual flexibility among normal-hearing older listeners, despite an assumed decline in temporal processing.more » « less
An official website of the United States government

