Search for: All records

Creators/Authors contains: "Lu, Yijing"

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. Free, publicly-accessible full text available August 17, 2026
  2. Purpose:Articulatory behaviors during moments of stuttering have been understudied, largely due to the technical difficulty of collecting such data. Tracking moving articulators during stuttering requires advanced instrumentation, and eliciting stuttering in a lab setting poses challenges for experimental design. To address these difficulties, we present a novel methodology that combines real-time vocal tract magnetic resonance imaging (MRI) with a suite of connected speech tasks to elicit stuttering. Method:A high-performance 0.55 T MRI system, with a custom eight-channel upper airway coil and a spiral balanced steady-state free precession pulse sequence, was used to acquire real-time MRI speech production data from seven adults who stutter. During scans, participants performed three connected speech tasks that incorporate stuttering-inducing factors: (a) passage reading, (b) short interviews with the experimenter, and (c) picture description within a time limit. Speech tasks were interleaved with one another. Results:Each participant produced over 100 stuttered words, covering various disfluency types and linguistic features. Fluent and disfluent productions of the same words were elicited, enabling direct articulatory comparisons. Participants did not show a significant decrease in the percentage of syllables stuttered (%SS) inside the scanner compared to outside, suggesting that our protocol effectively mitigated fluency-enhancing factors during scanning. %SS in each speech task varied substantially across participants, justifying the inclusion of multiple task types. Interleaving different tasks helped maintain a stable %SS throughout. The collected real-time MRI vocal tract videos reveal meaningful articulatory behaviors during stuttering that are not detectable via acoustics alone. Conclusions:The suite of specially designed speech tasks was effective in eliciting stuttering during real-time MRI data collection. Combining these speech tasks with dynamic MRI technology offers a powerful approach to studying the articulatory mechanisms of stuttering. In addition to real-time MRI, these speech tasks have the potential to be combined with other experimental instrumentation to facilitate collecting data specifically during stuttered speech. 
    more » « less
    Free, publicly-accessible full text available September 10, 2026
  3. Free, publicly-accessible full text available August 17, 2026
  4. Variability in speech pronunciation is widely observed across different linguistic backgrounds, which impacts modern automatic speech recognition performance. Here, we evaluate the performance of a self-supervised speech model in phoneme recognition using direct articulatory evidence. Findings indicate significant differences in phoneme recognition, especially in front vowels, between American English and Indian English speakers. To gain a deeper understanding of these differences, we conduct real-time MRI-based articulatory analysis, revealing distinct velar region patterns during the production of specific front vowels. This underscores the need to deepen the scientific understanding of self-supervised speech model variances to advance robust and inclusive speech technology. 
    more » « less
  5. Individuals who have undergone treatment for oral cancer oftentimes exhibit compensatory behavior in consonant production. This pilot study investigates whether compensatory mechanisms utilized in the production of speech sounds with a given target constriction location vary systematically depending on target manner of articulation. The data reveal that compensatory strategies used to produce target alveolar segments vary systematically as a function of target manner of articulation in subtle yet meaningful ways. When target constriction degree at a particular constriction location cannot be preserved, individuals may leverage their ability to finely modulate constriction degree at multiple constriction locations along the vocal tract. 
    more » « less
  6. There is a lack of general agreement among previous studies (e.g., Bakst, 2016; Dediu & Moisik, 2019; Westbury et al., 1998) on whether measurements of vocal tract morphology are robust predictors of inter-speaker variation in tongue shaping for American English /ɹ/. One possible reason is the different quantifications of /ɹ/ tongue shapes that were employed. The current study compares the relationships between a single set of anatomical measurements and three different measures of lingual articulation for /ɹ/ in /ɑɹɑ/ in midsagittal real-time MRI data. A novel method was developed to quantify the palatal constriction location and length, which served as the first two measures of tongue shape. A linear Support Vector Machine divided the constriction location and length measures into regions that approximate the visually identified categories of “retroflex” and “bunched.” The third shape measurement is the signed distance of each token of /ɹ/ to the division boundary, representing the degree of “retroflexion” or “bunchedness” based on palatal constriction properties. These three measures showed marginally to moderately significant linear relationships with two specific measures of individual speakers’ vocal tract anatomy: the degree of mandibular inclination and the length of the oral cavity roof. Overall, the effect of anatomy on the lingual articulation of /ɹ/ is not strong. [Work supported by NSF, Grant 1908865.] 
    more » « less
  7. The theory of Task Dynamics provides a method of predicting articulatory kinematics from a discrete phonologically-relevant representation (“gestural score”). However, because the implementations of that model (e.g., Nam et al., 2004) have generally used a simplified articulatory geometry (Mermelstein et al., 1981) whose forward model (from articulator to constriction coordinates) can be analytically derived, quantitative predictions of the model for individual human vocal tracts have not been possible. Recently, methods of deriving individual speaker forward models from real-time MRI data have been developed (Sorensen et al., 2019). This has further allowed development of task dynamic models for individual speakers, which make quantitative predictions. Thus far, however, these models (Alexander et al., 2019) could only synthesize limited types of utterances due to their inability to model temporally overlapping gestures. An updated implementation is presented, which can accommodate overlapping gestures and incorporates an optimization loop to improve the fit of modeled articulatory trajectories to the observed ones. Using an analysis-by-synthesis approach, the updated implementation can be utilized: (1) to refine the hypothesized speaker-general gestural parameters (target, stiffness) for individual speakers; (2) to test different degrees of temporal overlapping among multiple gestures such as a CCVC syllable. [Work supported by NSF, Grant 1908865.] 
    more » « less