Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection

Xu, Weijia; Agrawal, Sweta; Briakou, Eleftheria; Martindale, Marianna J; Carpuat, Marine

doi:10.1162/tacl_a_00563

Citation Details

Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection

Neural sequence generation models are known to “hallucinate”, by producing outputs that are unrelated to the source text. These hallucinations are potentially harmful, yet it remains unclear in what conditions they arise and how to mitigate their impact. In this work, we first identify internal model symptoms of hallucinations by analyzing the relative token contributions to the generation in contrastive hallucinated vs. non-hallucinated outputs generated via source perturbations. We then show that these symptoms are reliable indicators of natural hallucinations, by using them to design a lightweight hallucination detector which outperforms both model-free baselines and strong classifiers based on quality estimation or large pre-trained models on manually annotated English-Chinese and German-English translation test beds. more »

Award ID(s):: 1750695

PAR ID:: 10520446

Author(s) / Creator(s):: Xu, Weijia; Agrawal, Sweta; Briakou, Eleftheria; Martindale, Marianna J; Carpuat, Marine

Publisher / Repository:: Association for Computational Linguistics

Date Published:: 2023-01-01

Journal Name:: Transactions of the Association for Computational Linguistics

Volume:: 11

ISSN:: 2307-387X

Page Range / eLocation ID:: 546 to 564

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Journal Article:
https://doi.org/10.1162/tacl_a_00563

More Like this