CORAAL QA: A Dataset and Framework for Open Domain Spontaneous Speech Question Answering from Long Audio Files

Shankar, Natarajan Balaji; Johnson, Alexander; Chance, Christina; Veeramani, Hariram; Alwan, Abeer

doi:10.1109/ICASSP48485.2024.10447109

Citation Details

CORAAL QA: A Dataset and Framework for Open Domain Spontaneous Speech Question Answering from Long Audio Files

This paper presents a novel dataset (CORAAL QA) and framework for audio question-answering from long audio recordings contain- ing spontaneous speech. The dataset introduced here provides sets of questions that can be factually answered from short spans of a long audio files (typically 30min to 1hr) from the Corpus of Re- gional African American Language. Using this dataset, we divide the audio recordings into 60 second segments, automatically tran- scribe each segment, and use PLDA scoring of BERT-based seman- tic embeddings to rank the relevance of ASR transcript segments in answering the target question. In order to improve this framework through data augmentation, we use large language models including ChatGPT and Llama 2 to automatically generate further training ex- amples and show how prompt engineering can be optimized for this process. By creatively leveraging knowledge from large-language models, we achieve state-of-the-art question-answering performance in this information retrieval task. more »

Award ID(s):: 2202585

PAR ID:: 10506580

Author(s) / Creator(s):: Shankar, Natarajan Balaji; Johnson, Alexander; Chance, Christina; Veeramani, Hariram; Alwan, Abeer

Publisher / Repository:: IEEE

Date Published:: 2024-04-14

Journal Name:: Proceedings of the IEEE International Conference on Acoustics Speech and Signal Processing

ISSN:: 2379-190X

ISBN:: 979-8-3503-4485-1

Page Range / eLocation ID:: 13371 to 13375

Format(s):: Medium: X

Location:: Seoul, Korea, Republic of

Sponsoring Org:: National Science Foundation

Conference Paper:
https://doi.org/10.1109/ICASSP48485.2024.10447109

More Like this