Attention:NSF PAR will be unavailable due to scheduled facility maintenance July 10th, 5 PM EDT through July 13th. We apologize for any inconvenience.


Search for: All records

Creators/Authors contains: "Wang, Z"

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. Speech foundation models often struggle in low-resource domains due to domain mismatch and data scarcity. We propose Gumbel-BEARD, a domain adaptation framework that automates Whisper encoder layer selection via an end-to-end trainable hard Gumbel-Softmax selector. It enables self-supervised adaptation with a BEST-RQ objective that dynamically adapts to target acoustic characteristics without manual tuning. Experiments on the MyST child speech corpus demonstrate efficiency and scalability: with 10 h of labeled data for fine-tuning, our method matches a fully supervised baseline trained on the complete 133 h labeled set. We establish new state-of-the-art word error rates (WERs) of 8.21% using Whisper-medium on MyST and 11.06% using Whisper-small on the OGI Spontaneous dataset. Evaluation on CORAAL further confirms robustness to adult dialectal domain shifts, with up to 6% relative WER reduction, highlighting the generalizability of our approach to diverse low-resource conditions. 
    more » « less
    Free, publicly-accessible full text available September 27, 2027
  2. Transformer-based Speech Foundation Models excel in most Automatic Speech Recognition tasks but often suffer performance degradation when applied to domains with mismatched acoustic characteristics. While Parameter Efficient Fine-Tuning (PEFT) methods, such as Low-Rank Adaptation (LoRA), adjust global attention, they lack the local context modeling crucial for capturing domain-specific variations. We propose GC-LoRA, a novel adapter architecture that injects Conformer-style local convolutional processing into pretrained Transformer encoders. By integrating a lightweight adapter to encoder attention output projections, our method efficiently captures local acoustic dependencies without disrupting pretrained global representations. Experiments across diverse datasets (acoustically-degraded, bandlimited, dialectal, child) demonstrate the efficacy of our approach, achieving Word Error Rate (WER) reductions of up to 10.9% compared to baselines while adding minimal trainable parameters. 
    more » « less
    Free, publicly-accessible full text available September 27, 2027
  3. While Speech Large Language Models (Speech-LLMs) have achieved strong performance on adult Automatic Speech Recognition (ASR), their effectiveness on child speech remains under-explored, and single models often struggle to handle diverse adult and child age groups simultaneously. This paper proposes a Mixture-of-Experts (MoE) Speech-LLM for unified ASR across adult and child speech spanning diverse environments and age groups. The framework employs a Classifier-based Domain Router (C-DR) with a coarse-to-fine strategy and integrates both a Mixture-of-Projectors (MoP) and a Mixture-of-LoRAs (MoL) to model domain-specific variations. To address routing uncertainty near domain boundaries, an Entropy-Aware Routing (EAR) mechanism is introduced to dynamically incorporate a shared expert. Experiments on public child corpora demonstrate consistent improvements over baselines while preserving adult ASR performance. To our knowledge, this is the first work leveraging Speech-LLMs for unified, multi-domain ASR encompassing both children and adults. 
    more » « less
    Free, publicly-accessible full text available September 27, 2027
  4. Free, publicly-accessible full text available July 12, 2027
  5. Free, publicly-accessible full text available June 12, 2027
  6. Free, publicly-accessible full text available June 3, 2027
  7. Free, publicly-accessible full text available January 15, 2027
  8. Free, publicly-accessible full text available February 14, 2027
  9. Free, publicly-accessible full text available July 31, 2026
  10. Low-income households (LIH), exposed to the uncertain modern grid, bear greater energy burdens and face inequitable access to reliable power compared to high-income households (HIH). This paper proposes a two-stage stochastic community-based microgrid planning (CMP) framework to boost energy justice within the system. To reduce the negative impact of income levels, a weighted energy cost model for households within the microgrid (MG) is designed. To address the multisource uncertainty during the operation period, a two-stage stochastic framework is developed. Moreover, to assess the proposed method, the unbalanced IEEE 123 node system is employed and modified as an isolated MG. The analysis reveals the proposed model can achieve a risk-averse solution while economic optimality is guaranteed. Additionally, the designed weighted method improves the LIH’s impact rate to 67.95% and decreases the total planning cost by 22.43%. 
    more » « less
    Free, publicly-accessible full text available August 27, 2026