Quantifying the Effects of Text Duplication on Semantic Models

Schofield, Alexandra; Magnusson, Mans; Mimno, David

Citation Details

Duplicate documents are a pervasive problem in text datasets and can have a strong effect on unsupervised models. Methods to remove duplicate texts are typically heuristic or very expensive, so it is vital to know when and why they are needed. We measure the sensitivity of two latent semantic methods to the presence of different levels of document repetition. By artificially creating different forms of duplicate text we confirm several hypotheses about how repeated text impacts models. While a small amount of duplication is tolerable, substantial over-representation of subsets of the text may overwhelm meaningful topical patterns. more »

Award ID(s):: 1652536

PAR ID:: 10057833

Author(s) / Creator(s):: Schofield, Alexandra; Magnusson, Mans; Mimno, David

Date Published:: 2018-01-01

Journal Name:: Empirical Methods in Natural Language Processing

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this