Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 5:00 PM ET until 8:00 PM ET on Friday, September 11 due to maintenance. We apologize for the inconvenience.


This content will become publicly available on July 6, 2027

Title: Splits! Flexible Sociocultural Linguistic Investigation at Scale
Variation in language use, shaped by speakers' sociocultural background and specific context of use, offers a rich lens into cultural perspectives, values, and opinions. For example, Chinese students discuss healthy eating with words like timing, regularity, and digestion, whereas Americans use vocabulary like balancing food groups and avoiding fat and sugar, reflecting distinct cultural models of nutrition (Banna et al., 2016). The computational study of these Sociocultural Linguistic Phenomena (SLP) has traditionally been done in NLP via tailored analyses of specific groups or topics, requiring specialized data collection and experimental operationalization—a process not well-suited to quick hypothesis exploration and prototyping. To address this, we propose constructing a "sandbox" designed for systematic and flexible sociolinguistic research. Using our method, we construct a demographically/topically split Reddit dataset, Splits!, validated by self-identification and by replicating several known SLPs from existing literature. We showcase the sandbox's utility with a scalable, two-stage process that filters large collections of potential SLPs (PSLPs) to surface the most promising candidates for deeper, qualitative investigation.  more » « less
Award ID(s):
2048001
PAR ID:
10688183
Author(s) / Creator(s):
; ;
Publisher / Repository:
Association for Computational Linguistics (ACL 2026)
Date Published:
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Context: Supervised learning-based projects (SLPs), i.e., software projects that use supervised learning algorithms, such as decision trees are useful for performing classification-related tasks. Yet, security weaknesses, such as the use of hard-coded passwords in SLPs, can make SLPs susceptible to security attacks. A characterization of security weaknesses in SLPs can help practitioners understand the security weaknesses that are frequent in SLPs and adopt adequate mitigation strategies. Objective: The goal of this paper is to help practitioners securely develop supervised learning-based projects by conducting an empirical study of security weaknesses in supervised learning-based projects. Methodology: We conduct an empirical study by quantifying the frequency of security weaknesses in 278 open source SLPs. Results: We identify 22 types of security weaknesses that occur in SLPs. We observe ‘use of potentially dangerous function’ to be the most frequently occurring security weakness in SLPs. Of the identified 3,964 security weaknesses, 23.79% and 40.49% respectively, appear for source code files used to train and test models. We also observe evidence of co-location, e.g., instances of command injection co-locates with instances of potentially dangerous function. Conclusion: Based on our findings, we advocate for a shift left approach for SLP development with security-focused code reviews, and application of security static analysis. 
    more » « less
  2. Population genetics has been successful at identifying the relationships between human groups and their interconnected histories. However, the link between genetic demography inferred at large scales and the individual human behaviours that ultimately generate that demography is not always clear. While anthropological and historical context are routinely presented as adjuncts in population genetic studies to help describe the past, determining how underlying patterns of human sociocultural behaviour impact genetics still remains challenging. Here, we analyse patterns of genetic variation in village-scale samples from two islands in eastern Indonesia, patrilocal Sumba and a matrilocal region of Timor. Adopting a ‘process modelling’ approach, we iteratively explore combinations of structurally different models as a thinking tool. We find interconnected socio-genetic interactions involving sex-biased migration, lineage-focused founder effects, and on Sumba, heritable social dominance. Strikingly, founder ideology, a cultural model derived from anthropological and archaeological studies at larger regional scales, has both its origins and impact at the scale of villages. Process modelling lets us explore these complex interactions, first by circumventing the complexity of formal inference when studying large datasets with many interacting parts, and then by explicitly testing complex anthropological hypotheses about sociocultural behaviour from a more familiar population genetic standpoint. 
    more » « less
  3. Augmentative and alternative communication (AAC) devices are used by many people around the world who experience difficulties in communicating verbally. One form of AAC device which is especially useful for minimally verbal autistic children in developing language and communication skills is the visual scene display (VSD). VSDs use images with interactive hotspots embedded in them to directly connect language to real-world contexts which are meaningful to the AAC user. While VSDs can effectively support emergent communicators (i.e., those who are beginning to learn how to use symbolic communication), their widespread adoption is impacted by how difficult these devices are to configure. We developed a prototype that uses generative AI to automatically suggest initial hotspots on an image to help non-experts efficiently create visual scene displays (VSDs). We conducted a within-subjects user study to understand how effective our prototype is in supporting non-expert users, specifically pre-service speech-language pathologists (SLPs) (N=16) who are not familiar with VSDs as an AAC intervention. Pre-service SLPs are actively studying to become clinically certified SLPs and have domain-specific knowledge about language and communication skill development. We evaluated the effectiveness of our prototype based on creation time, quality, and user confidence. We also analyzed the relevance and developmental appropriateness of the automatically generated hotspots and how often users interacted with (e.g., editing or deleting) the generated hotspots. Our results were mixed with SLPs becoming more efficient and confident. However, there were multiple negative impacts as well, including over-reliance and homogenization of communication options. The implications of these findings reach beyond the domain of AAC, especially as generative AI becomes more prevalent across domains, including assistive technology. Future work is needed to further identify and address these risks associated with integrating generative AI into assistive technology. 
    more » « less
  4. Abstract Dyslexia and dysgraphia are two specific learning disabilities (SLDs) that are prevalent among children. To minimize the negative impact these SLDs have on a child’s academic and social-emotional development, it is crucial to identify dyslexia and dysgraphia at an early age, enabling timely and effective intervention. The first step in this process is screening, which helps determine if a child requires further instruction or a more in-depth assessment. Current screening tools are expensive, require additional administration time beyond regular classroom activities, and are designed to screen exclusively for one condition, not for both dyslexia and dysgraphia, which often share some common behavioral characteristics. Most dyslexia screeners focus on speech and oral tasks and exclude writing activities. However, analyzing children’s writing samples for behavioral signs of dyslexia and dysgraphia can offer valuable insights into the screening process, which can be time-consuming. As a solution, we propose a co-designed framework for building artificial intelligence (AI) tools that could boost the efficiency of screening and aid practitioners such as speech-language pathologists (SLPs), occupational therapists, general educators, and special educators by simplifying their tasks. This paper reviews current screening methods employed by practitioners, the use of AI-based systems in identifying dyslexia and dysgraphia, and the handwriting datasets available to train such systems. The paper also outlines a framework for developing an AI-integrated screening tool that can identify writing-based behavioral indicators of dyslexia and dysgraphia in children’s handwriting. This framework can be used in conjunction with current screening tools like the Dysgraphia and Dyslexia Behavioral Indicator Checklist (DDBIC). The paper also proposes a methodology for collecting children’s offline and online handwriting samples to build a valuable dataset for developing AI solutions. The proposed framework and data collection methodology are co-designed with SLPs, occupational therapists (OTs), special educators, and general educators to ensure the tool can provide explainable, actionable information that would be invaluable in a practical setting. 
    more » « less
  5. null (Ed.)
    We report direct measurements of spatially resolved surface stresses of a dense suspension during large amplitude oscillatory shear (LAOS) in the discontinuous shear thickening regime using boundary stress microscopy. Consistent with previous studies, bulk rheology shows a dramatic increase in the complex viscosity above a frequency-dependent critical strain. We find that the viscosity increase is coincident with that appearance of large heterogeneous boundary stresses, indicative of the formation of transient solid-like phases (SLPs) on spatial scales large compared to the particle size. The critical strain for the appearance of SLPs is largely determined by the peak oscillatory stress, which depends on the peak shear rate and the frequency-dependent suspension viscosity. The SLPs dissipate and reform on each cycle, with a spatial pattern that is highly variable at low frequencies but remarkably persistent at the highest frequency measured ( ω = 10 rad s −1 ). 
    more » « less