Abstract PurposeTo demonstrate speech‐production real‐time MRI (RT‐MRI) using a contemporary 0.55T system, and to identify opportunities for improved performance compared with conventional field strengths. MethodsExperiments were performed on healthy adult volunteers using a 0.55T MRI system with high‐performance gradients and a custom 8‐channel upper airway coil. Imaging was performed using spiral‐based balancedSSFPand gradient‐recalled echo (GRE) pulse sequences using a temporal finite‐difference constrained reconstruction. Speech‐production RT‐MRI was performed with three spiral readout durations (8.90, 5.58, and 3.48 ms) to determine trade‐offs with respect to articulator contrast, blurring, banding artifacts, and overall image quality. ResultsBoth spiral GRE and bSSFP captured tongue boundary dynamics during rapid consonant‐vowel syllables. Although bSSFP provided substantially higher SNR in all vocal tract articulators than GRE, it suffered from banding artifacts at TR > 10.9 ms. Spiral bSSFP with the shortest readout duration (3.48 ms, TR = 5.30 ms) had the best image quality, with a 1.54‐times boost in SNR compared with an equivalent GRE sequence. Longer readout durations led to increased SNR efficiency and blurring in both bSSFP and GRE. ConclusionHigh‐performance 0.55T MRI systems can be used for speech‐production RT‐MRI. Spiral bSSFP can be used without suffering from banding artifacts in vocal tract articulators, provide better SNR efficiency, and have better image quality than what is typically achieved at 1.5 T or 3 T.
more »
« less
This content will become publicly available on September 10, 2026
Effective Elicitation of Stuttering in Magnetic Resonance Imaging Data Collection Using a Suite of Connected Speech Tasks
Purpose:Articulatory behaviors during moments of stuttering have been understudied, largely due to the technical difficulty of collecting such data. Tracking moving articulators during stuttering requires advanced instrumentation, and eliciting stuttering in a lab setting poses challenges for experimental design. To address these difficulties, we present a novel methodology that combines real-time vocal tract magnetic resonance imaging (MRI) with a suite of connected speech tasks to elicit stuttering. Method:A high-performance 0.55 T MRI system, with a custom eight-channel upper airway coil and a spiral balanced steady-state free precession pulse sequence, was used to acquire real-time MRI speech production data from seven adults who stutter. During scans, participants performed three connected speech tasks that incorporate stuttering-inducing factors: (a) passage reading, (b) short interviews with the experimenter, and (c) picture description within a time limit. Speech tasks were interleaved with one another. Results:Each participant produced over 100 stuttered words, covering various disfluency types and linguistic features. Fluent and disfluent productions of the same words were elicited, enabling direct articulatory comparisons. Participants did not show a significant decrease in the percentage of syllables stuttered (%SS) inside the scanner compared to outside, suggesting that our protocol effectively mitigated fluency-enhancing factors during scanning. %SS in each speech task varied substantially across participants, justifying the inclusion of multiple task types. Interleaving different tasks helped maintain a stable %SS throughout. The collected real-time MRI vocal tract videos reveal meaningful articulatory behaviors during stuttering that are not detectable via acoustics alone. Conclusions:The suite of specially designed speech tasks was effective in eliciting stuttering during real-time MRI data collection. Combining these speech tasks with dynamic MRI technology offers a powerful approach to studying the articulatory mechanisms of stuttering. In addition to real-time MRI, these speech tasks have the potential to be combined with other experimental instrumentation to facilitate collecting data specifically during stuttered speech.
more »
« less
- Award ID(s):
- 2106930
- PAR ID:
- 10685097
- Publisher / Repository:
- https://pubs.asha.org
- Date Published:
- Journal Name:
- Journal of Speech, Language, and Hearing Research
- Volume:
- 68
- Issue:
- 9
- ISSN:
- 1092-4388
- Page Range / eLocation ID:
- 4275 to 4289
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
Abstract The significance of respiratory droplet transmission in spreading respiratory diseases such as COVID-19 has been identified by researchers. Although one cough or sneeze generates a large number of respiratory droplets, they are usually infrequent. In comparison, speaking and singing generate fewer droplets, but occur much more often, highlighting their potential as a vector for airborne transmission. However, the flow dynamics of speech and the transmission of speech droplets have not been fully investigated. To shed light on this topic, two-dimensional geometries of a vocal tract for a labiodental fricative [f] were generated based on real-time MRI of a subject during pronouncing [f]. In these models, two different curvatures were considered for the tip tongue shape and the lower lip to highlight the effects of the articulator geometries on transmission dynamics. The commercial ANSYS-Fluent CFD software was used to solve the complex expiratory speech airflow trajectories. Simultaneously, the discrete phase model of the software was used to track submicron and large size respiratory droplets exhaled during [f] utterance. The simulations were performed for high, normal, and low lung pressures to explore the influence of loud, normal, and soft utterances, respectively, on the airflow dynamics. The presented results demonstrate the variability of the airflow and droplet propagation as a function of the vocal tract geometrical characteristics and loudness.more » « less
-
This study investigates the speech articulatory coordination in schizophrenia subjects exhibiting strong positive symptoms (e.g. hallucinations and delusions), using two distinct channel-delay correlation methods. We show that the schizophrenic subjects with strong positive symptoms and who are markedly ill pose complex articulatory coordination pattern in facial and speech gestures than what is observed in healthy subjects. This distinction in speech coordination pattern is used to train a multimodal convolutional neural network (CNN) which uses video and audio data during speech to distinguish schizophrenic patients with strong positive symptoms from healthy subjects. We also show that the vocal tract variables (TVs) which correspond to place of articulation and glottal source outperform the Mel-frequency Cepstral Coefficients (MFCCs) when fused with Facial Action Units (FAUs) in the proposed multimodal network. For the clinical dataset we collected, our best performing multimodal network improves the mean F1 score for detecting schizophrenia by around 18% with respect to the full vocal tract coordination (FVTC) baseline method implemented with fusing FAUs and MFCCs.more » « less
-
Variability in speech pronunciation is widely observed across different linguistic backgrounds, which impacts modern automatic speech recognition performance. Here, we evaluate the performance of a self-supervised speech model in phoneme recognition using direct articulatory evidence. Findings indicate significant differences in phoneme recognition, especially in front vowels, between American English and Indian English speakers. To gain a deeper understanding of these differences, we conduct real-time MRI-based articulatory analysis, revealing distinct velar region patterns during the production of specific front vowels. This underscores the need to deepen the scientific understanding of self-supervised speech model variances to advance robust and inclusive speech technology.more » « less
-
Abstract BackgroundElectrocorticography (ECoG) language mapping is often performed extraoperatively, frequently involves offline processing, and relationships with direct cortical stimulation (DCS) remain variable. We sought to determine the feasibility and preliminary utility of an intraoperative language mapping approach guided by real-time visualization of electrocorticograms. MethodsA patient with astrocytoma underwent awake craniotomy with intraoperative language mapping, utilizing a dual iPad stimulus presentation system coupled to a real-time neural signal processing platform capable of both ECoG recording and delivery of DCS. Gamma band modulations in response to 4 language tasks at each electrode were visualized in real-time. Next, DCS was conducted for each neighboring electrode pair during language tasks. ResultsAll language tasks resulted in strongest heat map activation at an electrode pair in the anterior to mid superior temporal gyrus. Consistent speech arrest during DCS was observed for Object and Action naming tasks at these same electrodes, indicating good correspondence with ECoG heat map recordings. This region corresponded well with posterior language representation via preoperative functional MRI. ConclusionsIntraoperative real-time visualization of language task-based ECoG gamma band modulation is feasible and may help identify targets for DCS. If validated, this may improve the efficiency and accuracy of intraoperative language mapping.more » « less
An official website of the United States government
