ConfliBERT: A Pre-trained Language Model for Political Conflict and Violence

Hu, Yibo; Hosseini, MohammadSaleh; Skorupa Parolin, Erick; Osorio, Javier; Khan, Latifur; Brandt, Patrick; D’Orazio, Vito

doi:10.18653/v1/2022.naacl-main.400

Citation Details

ConfliBERT: A Pre-trained Language Model for Political Conflict and Violence

Analyzing conflicts and political violence around the world is a persistent challenge in the political science and policy communities due in large part to the vast volumes of specialized text needed to monitor conflict and violence on a global scale. To help advance research in political science, we introduce ConfliBERT, a domain-specific pre-trained language model for conflict and political violence. We first gather a large domain-specific text corpus for language modeling from various sources. We then build ConfliBERT using two approaches: pre-training from scratch and continual pre-training. To evaluate ConfliBERT, we collect 12 datasets and implement 18 tasks to assess the models’ practical application in conflict research. Finally, we evaluate several versions of ConfliBERT in multiple experiments. Results consistently show that ConfliBERT outperforms BERT when analyzing political violence and conflict. more »

Award ID(s):: 1931541

NSF-PAR ID:: 10470310

Author(s) / Creator(s):: Hu, Yibo; Hosseini, MohammadSaleh; Skorupa Parolin, Erick; Osorio, Javier; Khan, Latifur; Brandt, Patrick; D’Orazio, Vito

Publisher / Repository:: Association for Computational Linguistics

Date Published:: 2022-07-01

Page Range / eLocation ID:: 5469 to 5482

Format(s):: Medium: X

Location:: Seattle, United States

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
https://doi.org/10.18653/v1/2022.naacl-main.400

More Like this