GooAQ: Open Question Answering with Diverse Answer Types

Khashabi, Daniel; Ng, Amos; Khot, Tushar; Sabharwal, Ashish; Hajishirzi, Hannaneh; Callison-Burch, Chris

doi:10.18653/v1/2021.findings-emnlp.38

Citation Details

GooAQ: Open Question Answering with Diverse Answer Types

While day-to-day questions come with a variety of answer types, the current question-answering (QA) literature has failed to adequately address the answer diversity of questions. To this end, we present GooAQ, a large-scale dataset with a variety of answer types. This dataset contains over 5 million questions and 3 million answers collected from Google. GooAQ questions are collected semi-automatically from the Google search engine using its autocomplete feature. This results in naturalistic questions of practical interest that are nonetheless short and expressed using simple language. GooAQ answers are mined from Google’s responses to our collected questions, specifically from the answer boxes in the search results. This yields a rich space of answer types, containing both textual answers (short and long) as well as more structured ones such as collections. We benchmark T5 models on GooAQ and observe that: (a) in line with recent work, LM’s strong performance on GooAQ’s short-answer questions heavily benefit from annotated data; however, (b) their quality in generating coherent and accurate responses for questions requiring long responses (such as ‘how’ and ‘why’ questions) is less reliant on observing annotated data and mainly supported by their pre-training. We release GooAQ to facilitate further research on improving QA with diverse response types. more »

Award ID(s):: 1928474

PAR ID:: 10344228

Author(s) / Creator(s):: Khashabi, Daniel; Ng, Amos; Khot, Tushar; Sabharwal, Ashish; Hajishirzi, Hannaneh; Callison-Burch, Chris

Date Published:: 2021-01-01

Journal Name:: Findings of the Association for Computational Linguistics: EMNLP 2021

Page Range / eLocation ID:: 421 to 433

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
https://doi.org/10.18653/v1/2021.findings-emnlp.38

More Like this