Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 5:00 PM ET until 8:00 PM ET on Friday, September 11 due to maintenance. We apologize for the inconvenience.


Title: ConceptCarve: Dynamic Realization of Evidence
Finding evidence for human opinion and behavior at scale is a challenging task, often requiring an understanding of sophisticated thought patterns among vast online communities found on social media. For example, studying how ‘gun ownership’ is related to the perception of ‘Freedom’, requires a retrieval system that can operate at scale over social media posts, while dealing with two key challenges: (1) identifying abstract concept instances, (2) which can be instantiated differently across different communities. To address these, we introduce ConceptCarve, an evidence retrieval framework that utilizes traditional retrievers and LLMs to dynamically characterize the search space during retrieval. Our experiments show that ConceptCarve surpasses traditional retrieval systems in finding evidence within a social media community. It also produces an interpretable representation of the evidence for that community, which we use to qualitatively analyze complex thought patterns that manifest differently across the communities.  more » « less
Award ID(s):
2048001
PAR ID:
10688167
Author(s) / Creator(s):
;
Publisher / Repository:
Association for Computational Linguistics
Date Published:
Page Range / eLocation ID:
20792 to 20809
Format(s):
Medium: X
Location:
Vienna, Austria
Sponsoring Org:
National Science Foundation
More Like this
  1. This article seeks to go beyond traditional GIS methods used in creating maps for disaster response that commonly look at the disaster extent. Instead, a slightly different approach is taken using social media data collected from Twitter to explore how people communicate during disaster events, how online communities form and evolve, and how communication methods can improve. This study collected the Twitter data during the 2015 Nepal earthquake disaster and applied a spatiotemporal analysis to find any patterns that show shadows or gaps in communication channels in local communities’ communication. Linkages in social media can be used to understand how people communicate, how quickly they diffuse information, and how social networks form online during disasters. These can improve communication throughout disaster phases. This study offers a deeper understanding of the kinds of spatiotemporal patterns and spatial social networks that can be observed during disaster events. The need for better communication during disaster events is imperative for better disaster management, increasing community resilience, and saving lives. 
    more » « less
  2. Twitter bot detection is vital in combating misinformation and safeguarding the integrity of social media discourse. While malicious bots are becoming more and more sophisticated and personalized, standard bot detection approaches are still agnostic to social environments (henceforth, communities) the bots operate at. In this work, we introduce community-specific bot detection, estimating the percentage of bots given the context of a community. Our method{---}BotPercent{---}is an amalgamation of Twitter bot detection datasets and feature-, text-, and graph-based models, adjusted to a particular community on Twitter. We introduce an approach that performs confidence calibration across bot detection models, which addresses generalization issues in existing community-agnostic models targeting individual bots and leads to more accurate community-level bot estimations. Experiments demonstrate that BotPercent achieves state-of-the-art performance in community-level Twitter bot detection across both balanced and imbalanced class distribution settings, presenting a less biased estimator of Twitter bot populations within the communities we analyze. We then analyze bot rates in several Twitter groups, including users who engage with partisan news media, political communities in different countries, and more. Our results reveal that the presence of Twitter bots is not homogeneous, but exhibiting a spatial-temporal distribution with considerable heterogeneity that should be taken into account for content moderation and social media policy making. 
    more » « less
  3. Introduction. When examining moral foundations in immigration discourse on social media, most studies focus on ideology-based groups rather than across specific immigrant groups. This ignores moral framings that depict some communities with empathy and others as threats. This research explores how moral foundations vary in Twitter conversations about five immigrant groups. Method. Tweets were sorted into five categories (based on immigrant groups being discussed: African, Asian, European, Latin American, or Middle Eastern) utilising keyword searches and AI LLM modeling. GPT-3.5 Turbo was employed and achieved a satisfactory performance (0.83) compared to manual human labeling. Analysis. Scores for foundation variables (care/harm, fairness/cheating, loyalty/betrayal, authority/subversion, and purity/sanctity) were analysed using enhMFD1 dictionary. One-way ANOVA tested overall differences between groups and Tukey’s HSD post-hoc test identified specific patterns. Results. Latin American immigrant discourse emphasised care and authority. European-focused tweets featured stronger loyalty. African immigrant discourse highlighted loyalty with moderate authority, discourse about Middle Eastern immigrants showed elevated harm and betrayal, and Asian immigrant discourse portrayed higher fairness-vice. Conclusion(s). Immigrant groups are framed differently through moral language in social media conversations, which may influence perceptions and can inform strategies for addressing harmful narratives on social media. 
    more » « less
  4. Abstract Social media data and computational tools have become increasingly powerful alternatives to traditional methods of data collection and annotation for sociolinguistic research. While not without their own drawbacks, these resources can help address critical problems faced by researchers, such as sparse data, the observer's paradox, and annotating corpora at scale. This article presents a Twitter dataset of 227 million conversational messages across the U.S. and uses deep learning natural language processing (NLP) methods to analyze how 18 morphosyntactic features used by speakers of African American Language (AAL) vary geographically and across 12 demographic factors. Results demonstrate more frequent usage of AAL features in the rural South and in Mexican American communities, both of which are underrepresented in the literature. This work constitutes the first national-level description and analysis of overall morphosyntactic variation in AAL, and demonstrates how NLP tools enable the study of large-scale data to gain a more representative understanding of speech in marginalized communities. 
    more » « less
  5. The diffusion of information about open-source projects is a key factor influencing the adoption of projects and the allocation of developer efforts. Developers learn about new projects, and evaluate their quality and importance by accessing the related information. Social media is an important channel for information diffusion about open-source projects, with previous research suggesting the existence of a social media ecosystem that consists of multiple platforms and collectively supports information diffusion in open source. With different features supporting information diffusion, the same piece of information likely reaches different developer communities on different platforms, which attracts the attention and contribution of different developers and thus influences the success of open-source projects. Despite its importance, few works looked at the identity of the developer community that projectrelated information reaches on social media platforms and its associated impact on the discussed project. In this work, we track social media discussions on open-source projects on three different platforms: Twitter, HackerNews, and Reddit. We first describe the dynamics of project-related information diffusion across platforms, and we analyze the association between the number of posts on each platform, and the number of developers attracted to the discussed project from different communities. We find that posts about open-source projects first appear on Twitter and HackerNews, then move more towards Reddit. The number of project-related posts on Twitter mostly associate with the attracted developers from communities that are close to the project’s main contributor, while posts on other platforms associate more with the attention from remote communities. 
    more » « less