skip to main content
US FlagAn official website of the United States government
dot gov icon
Official websites use .gov
A .gov website belongs to an official government organization in the United States.
https lock icon
Secure .gov websites use HTTPS
A lock ( lock ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites.


Title: A Corpus for Modeling User and Language Effects in Argumentation on Online Debating
Existing argumentation datasets have succeeded in allowing researchers to develop computational methods for analyzing the content, structure and linguistic features of argumentative text. They have been much less successful in fostering studies of the effect of “user” traits — characteristics and beliefs of the participants — on the debate/argument outcome as this type of user information is generally not available. This paper presents a dataset of 78,376 debates generated over a 10-year period along with surprisingly comprehensive participant profiles. We also complete an example study using the dataset to analyze the effect of selected user traits on the debate outcome in comparison to the linguistic features typically employed in studies of this kind.  more » « less
Award ID(s):
1741441
PAR ID:
10113368
Author(s) / Creator(s):
;
Date Published:
Journal Name:
Proceedings of the 57th Conference of the Association for Computational Linguistics (ACL)
Page Range / eLocation ID:
602-607
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. This paper examines the factors that govern persuasion for a priori UNDECIDED versus DECIDED audience members in the context of on-line debates. We separately study two types of influences: linguistic factors — features of the language of the debate itself; and audience factors — features of an audience member encoding demographic information, prior beliefs, and debate platform behavior. In a study of users of a popular debate platform, we find first that different combinations of linguistic features are critical for predicting persuasion outcomes for UNDECIDED versus DECIDED members of the audience. We additionally find that audience factors have more influence on predicting the side (PRO/CON) that persuaded UNDECIDED users than for DECIDED users that flip their stance to the opposing side. Our results emphasize the importance of considering the undecided and decided audiences separately when studying linguistic factors of persuasion. 
    more » « less
  2. Abstract Plant–insect interactions are ubiquitous, and have been studied intensely because of their relevance to damage and pollination in agricultural plants, and to the ecology and evolution of biodiversity. Variation within species can affect the outcome of these interactions. Specific genes and chemicals that mediate these interactions have been identified, but genome‐ or metabolome‐scale studies might be necessary to better understand the ecological and evolutionary consequences of intraspecific variation for plant–insect interactions. Here, we present such a study. Specifically, we assess the consequences of genome‐wide genetic variation in the model plantMedicago truncatulaforLycaeides melissacaterpillar growth and survival (larval performance). Using a rearing experiment and a whole‐genome SNP data set (>5 million SNPs), we found that polygenic variation inM. truncatulaexplains 9%–41% of the observed variation in caterpillar growth and survival. Genetic correlations among caterpillar performance and other plant traits, including structural defences and some anonymous chemical features, suggest that multipleM. truncatulaalleles have pleiotropic effects on plant traits and caterpillar performance (or that substantial linkage disequilibrium exists among distinct loci affecting subsets of these traits). A moderate proportion of the genetic effect ofM. truncatulaalleles onL. melissaperformance can be explained by the effect of these alleles on the plant traits we measured, especially leaf toughness. Taken together, our results show that intraspecific genetic variation inM. truncatulahas a substantial effect on the successful development ofL. melissacaterpillars (i.e., on a plant–insect interaction), and further point toward traits potentially mediating this genetic effect. 
    more » « less
  3. Identifying persuasive speakers in an adversarial environment is a critical task. In a national election, politicians would like to have persuasive speakers campaign on their behalf. When a company faces adverse publicity, they would like to engage persuasive advocates for their position in the presence of adversaries who are critical of them. Debates represent a common platform for these forms of adversarial persuasion. This paper solves two problems: the Debate Outcome Prediction (DOP) problem predicts who wins a debate while the Intensity of Persuasion Prediction (IPP) problem predicts the change in the number of votes before and after a speaker speaks. Though DOP has been previously studied, we are the first to study IPP. Past studies on DOP fail to leverage two important aspects of multimodal data: 1) multiple modalities are often semantically aligned, and 2) different modalities may provide diverse information for prediction. Our M2P2 (Multimodal Persuasion Prediction) framework is the first to use multimodal (acoustic, visual, language) data to solve the IPP problem. To leverage the alignment of different modalities while maintaining the diversity of the cues they provide, M2P2 devises a novel adaptive fusion learning framework which fuses embeddings obtained from two modules -- an alignment module that extracts shared information between modalities and a heterogeneity module that learns the weights of different modalities with guidance from three separately trained unimodal reference models. We test M2P2 on the popular IQ2US dataset designed for DOP. We also introduce a new dataset called QPS (from Qipashuo, a popular Chinese debate TV show) for IPP - we plan to release this dataset when the paper is published. M2P2 significantly outperforms 3 recent baselines on both datasets. 
    more » « less
  4. We report on the high success rates of our new, scalable, computational approach for sign recognition from monocular video, exploiting linguistically annotated ASL datasets with multiple signers. We recognize signs using a hybrid framework combining state-of-the-art learning methods with features based on what is known about the linguistic composition of lexical signs. We model and recognize the sub-components of sign production, with attention to hand shape, orientation, location, motion trajectories, plus non-manual features, and we combine these within a CRF framework. The effect is to make the sign recognition problem robust, scalable, and feasible with relatively smaller datasets than are required for purely data-driven methods. From a 350-sign vocabulary of isolated, citation-form lexical signs from the American Sign Language Lexicon Video Dataset (ASLLVD), including both 1- and 2-handed signs, we achieve a top-1 accuracy of 93.3% and a top-5 accuracy of 97.9%. The high probability with which we can produce 5 sign candidates that contain the correct result opens the door to potential applications, as it is reasonable to provide a sign lookup functionality that offers the user 5 possible signs, in decreasing order of likelihood, with the user then asked to select the desired sign. 
    more » « less
  5. We report on the high success rates of our new, scalable, computational approach for sign recognition from monocular video, exploiting linguistically annotated ASL datasets with multiple signers. We recognize signs using a hybrid framework combining state-of-the-art learning methods with features based on what is known about the linguistic composition of lexical signs. We model and recognize the sub-components of sign production, with attention to hand shape, orientation, location, motion trajectories, plus non-manual features, and we combine these within a CRF framework. The effect is to make the sign recognition problem robust, scalable, and feasible with relatively smaller datasets than are required for purely data-driven methods. From a 350-sign vocabulary of isolated, citation-form lexical signs from the American Sign Language Lexicon Video Dataset (ASLLVD), including both 1- and 2-handed signs, we achieve a top-1 accuracy of 93.3% and a top-5 accuracy of 97.9%. The high probability with which we can produce 5 sign candidates that contain the correct result opens the door to potential applications, as it is reasonable to provide a sign lookup functionality that offers the user 5 possible signs, in decreasing order of likelihood, with the user then asked to select the desired sign. 
    more » « less