From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models

Feng, Shangbin; Park, Chan Young; Liu, Yuhan; Tsvetkov, Yulia

Citation Details

Language models (LMs) are pretrained on diverse data sources—news, discussion forums, books, online encyclopedias. A significant portion of this data includes facts and opinions which, on one hand, celebrate democracy and diversity of ideas, and on the other hand are inherently socially biased. Our work develops new methods to (1) measure media biases in LMs trained on such corpora, along social and economic axes, and (2) measure the fairness of downstream NLP models trained on top of politically biased LMs. We focus on hate speech and misinformation detection, aiming to empirically quantify the effects of political (social, economic) biases in pretraining data on the fairness of high-stakes social-oriented tasks. Our findings reveal that pretrained LMs do have political leanings which reinforce the polarization present in pretraining corpora, propagating social biases into hate speech predictions and media biases into misinformation detectors. We discuss the implications of our findings for NLP research and propose future directions to mitigate unfairness. more »

Award ID(s):: 2142739 2203097 2040926 2125201

PAR ID:: 10433148

Author(s) / Creator(s):: Feng, Shangbin; Park, Chan Young; Liu, Yuhan; Tsvetkov, Yulia

Date Published:: 2023-07-01

Journal Name:: ACL: Annual Meeting of the Association for Computational Linguistics

Page Range / eLocation ID:: 11737–11762

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this