Improving Authorship Verification using Linguistic Divergence

Zhang, Yifan; Boumber, Dainis; Hosseinia, Marjan; Yang, Fan Yang; Mukherjee, Arjun

Citation Details

We propose an unsupervised solution to the Authorship Verification task that utilizes pre-trained deep language models to compute a new metric called DV-Distance. The proposed metric is a measure of the difference between the two authors comparing against pre-trained language models. Our design addresses the problem of non-comparability in authorship verification, frequently encountered in small or cross-domain corpora. To the best of our knowledge, this paper is the first one to introduce a method designed with non-comparability in mind from the ground up, rather than indirectly. It is also one of the first to use Deep Language Models in this setting. The approach is intuitive, and it is easy to understand and interpret through visualization. Experiments on four datasets show our methods matching or surpassing current state-of-the-art and strong baselines in most tasks. more »

Award ID(s):: 1838147

PAR ID:: 10292218

Author(s) / Creator(s):: Zhang, Yifan; Boumber, Dainis; Hosseinia, Marjan; Yang, Fan Yang; Mukherjee, Arjun

Date Published:: 2021-03-01

Journal Name:: Workshop on Reducing Online Misinformation through Credible Information Retrieval

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this