Papers
arxiv:2104.08727

GooAQ: Open Question Answering with Diverse Answer Types

Published on Apr 18, 2021
Authors:
,
,
,

Abstract

GooAQ, a large-scale dataset with diverse answer types, benchmarks T5 models, highlighting their strong performance on short answers and pre-training support for long responses.

AI-generated summary

While day-to-day questions come with a variety of answer types, the current question-answering (QA) literature has failed to adequately address the answer diversity of questions. To this end, we present GooAQ, a large-scale dataset with a variety of answer types. This dataset contains over 5 million questions and 3 million answers collected from Google. GooAQ questions are collected semi-automatically from the Google search engine using its autocomplete feature. This results in naturalistic questions of practical interest that are nonetheless short and expressed using simple language. GooAQ answers are mined from Google's responses to our collected questions, specifically from the answer boxes in the search results. This yields a rich space of answer types, containing both textual answers (short and long) as well as more structured ones such as collections. We benchmarkT5 models on GooAQ and observe that: (a) in line with recent work, LM's strong performance on GooAQ's short-answer questions heavily benefit from annotated data; however, (b) their quality in generating coherent and accurate responses for questions requiring long responses (such as 'how' and 'why' questions) is less reliant on observing annotated data and mainly supported by their pre-training. We release GooAQ to facilitate further research on improving QA with diverse response types.

Community

Sign up or log in to comment

Models citing this paper 216

Browse 216 models citing this paper

Datasets citing this paper 1

Spaces citing this paper 10,397

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.