X-RiSAWOZ: High-Quality End-to-End Multilingual Dialogue Datasets and Few-shot Agents

Moradshahi, Mehrad; Shen, Tianhao; Bali, Kalika; Choudhury, Monojit; de Chalendar, Gaël; Goel, Anmol; Kim, Sungkyun; Kodali, Prashant; Kumaraguru, Ponnurangam; Semmar, Nasredine; Semnani, Sina J.; Seo, Jiwon; Seshadri, Vivek; Shrivastava, Manish; Sun, Michael; Yadavalli, Aditya; You, Chaobin; Xiong, Deyi; Lam, Monica S.

Citation Details

Task-oriented dialogue research has mainly focused on a few popular languages like English and Chinese, due to the high dataset creation cost for a new language. To reduce the cost, we apply manual editing to automatically translated data. We create a new multilingual benchmark, X-RiSAWOZ, by translating the Chinese RiSAWOZ to 4 languages: English, French, Hindi, Korean; and a code-mixed English- Hindi language. X-RiSAWOZ has more than 18,000 human-verified dialogue utterances for each language, and unlike most multilingual prior work, is an end-to-end dataset for building fully-functioning agents. The many difficulties we encountered in creating X-RiSAWOZ led us to develop a toolset to accelerate the post-editing of a new language dataset after translation. This toolset improves machine translation with a hybrid entity alignment technique that combines neural with dictionary-based methods, along with many automated and semi-automated validation checks. We establish strong baselines for X-RiSAWOZ by training dialogue agents in the zero- and few-shot settings where limited gold data is available in the target language. Our results suggest that our translation and post-editing methodology and toolset can be used to create new high-quality multilingual dialogue agents cost-effectively. Our dataset, more »

Award ID(s):: 1900638

PAR ID:: 10427016

Author(s) / Creator(s):: Moradshahi, Mehrad; Shen, Tianhao; Bali, Kalika; Choudhury, Monojit; de Chalendar, Gaël; Goel, Anmol; Kim, Sungkyun; Kodali, Prashant; Kumaraguru, Ponnurangam; Semmar, Nasredine; Semnani, Sina J.; Seo, Jiwon; Seshadri, Vivek; Shrivastava, Manish; Sun, Michael; Yadavalli, Aditya; You, Chaobin; Xiong, Deyi; Lam, Monica S.

Date Published:: 2023-07-01

Journal Name:: Findings of the Association for Computational Linguistics (ACL), Toronto, Canada, 2023

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this