NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

ezCoref: Towards Unifying Annotation Guidelines for Coreference Resolution

Gupta, Ankita; Karpinska, Marzena; Zhao, Wenlong; Krishna, Kalpesh; Merullo, Jack; Yeh, Luke; Iyyer, Mohit; O'Connor, Brendan (May 2023, Findings of the Association for Computational Linguistics: EACL 2023)
Vlachos, Andreas; Augenstein, Isabelle (Ed.)
Large-scale, high-quality corpora are critical for advancing research in coreference resolution. However, existing datasets vary in their definition of coreferences and have been collected via complex and lengthy guidelines that are curated for linguistic experts. These concerns have sparked a growing interest among researchers to curate a unified set of guidelines suitable for annotators with various backgrounds. In this work, we develop a crowdsourcing-friendly coreference annotation methodology, ezCoref, consisting of an annotation tool and an interactive tutorial. We use ezCoref to re-annotate 240 passages from seven existing English coreference datasets (spanning fiction, news, and multiple other domains) while teaching annotators only cases that are treated similarly across these datasets. Surprisingly, we find that reasonable quality annotations were already achievable (90% agreement between the crowd and expert annotations) even without extensive training. On carefully analyzing the remaining disagreements, we identify the presence of linguistic cases that our annotators unanimously agree upon but lack unified treatments (e.g., generic pronouns, appositives) in existing datasets. We propose the research community should revisit these phenomena when curating future unified annotation guidelines.
more » « less
Full Text Available
Evaluating Zero-Shot Event Structures: Recommendations for Automatic Content Extraction (ACE) Annotations

Cai, Erica; O’Connor, Brendan (January 2023, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers))

https://aclanthology.org/2023.acl-short.142/
more » « less
Full Text Available
"Let Me Just Interrupt You": Estimating Gender Effects in Supreme Court Oral Arguments

Cai, Erica; Gupta, Ankita; Keith, Katherine; O'Connor, Brendan; Rice, Douglas R. (January 2023, SocArxiv)

Full Text Available
Investigating Morphosyntactic Variation in African American English on Twitter

Masis, Tessa; Eggleston, Chloe; Green, Lisa J.; Jones, Taylor; Armstrong, Meghan; and O'Connor, Brendan (January 2023, Proceedings of the Society for Computation in Linguistics)

https://doi.org/10.7275/zdg0-0914
more » « less
Full Text Available
Cross-Dialect Social Media Dependency Parsing for Social Scientific Entity Attribute Analysis

Eggleston, Chloe; O’Connor, Brendan (January 2022, Proceedings of the Eighth Workshop on Noisy User-generated Text (W-NUT 2022))

In this paper, we utilize recent advancements in social media natural language processing to obtain state-of-the-art syntactic dependency parsing results for social media English. We observe performance gains of 3.4 UAS and 4.0 LAS against the previous state-of-the-art as well as less disparity between African-American and Mainstream American English dialects. We demonstrate the computational social scientific utility of this parser for the task of socially embedded entity attribute analysis: for a specified entity, derive its semantic relationships from parses’ rich syntax, and accumulate and compare them across social variables. We conduct a case study on politicized views of U.S. official Anthony Fauci during the COVID-19 pandemic.
more » « less
Full Text Available
Corpus-Guided Contrast Sets for Morphosyntactic Feature Detection in Low-Resource English Varieties

Masis, Tessa; Neal, Anissa; Green, Lisa; O'Connor, Brendan (January 2022, Proceedings of the first workshop on NLP applications to field linguistics)

The study of language variation examines how language varies between and within different groups of speakers, shedding light on how we use language to construct identities and how social contexts affect language use. A common method is to identify instances of a certain linguistic feature - say, the zero copula construction - in a corpus, and analyze the feature’s distribution across speakers, topics, and other variables, to either gain a qualitative understanding of the feature’s function or systematically measure variation. In this paper, we explore the challenging task of automatic morphosyntactic feature detection in low-resource English varieties. We present a human-in-the-loop approach to generate and filter effective contrast sets via corpus-guided edits. We show that our approach improves feature detection for both Indian English and African American English, demonstrate how it can assist linguistic research, and release our fine-tuned models for use by other researchers.
more » « less
Full Text Available
Corpus-Guided Contrast Sets for Morphosyntactic Feature Detection in Low-Resource English Varieties

Masis, Tessa; Neal, Anissa; Green, Lisa; O'Connor, Brendan (January 2022, Proceedings of the 1st Field Matters Workshop on NLP Applications to Field Linguistics, at COLING 2022)

Full Text Available
Examining Political Rhetoric with Epistemic Stance Detection

Gupta, Ankita; Blodgett, Su Lin; Gross, Justin; O'Connor, Brendan (January 2022, Proceedings of the Fifth Workshop on Natural Language Processing and Computational Social Science (NLP+CSS))

https://aclanthology.org/2022.nlpcss-1.11
more » « less
Full Text Available
Paying Attention to the Algorithm Behind the Curtain: Bringing Transparency to YouTube’s Demonetization Algorithms

Dunna, Arun; Keith, Katherine A.; Zuckerman, Ethan; Vallina-Rodriguez, Narseo; O'Connor, Brendan; Nithyanand, Rishab (January 2022, Proceedings of the 2022 ACM Conference on Computer Supported Cooperative Work)

Full Text Available
ClioQuery: Interactive Query-Oriented Text Analytics for Comprehensive Investigation of Historical News Archives

https://doi.org/10.1145/3524025

Handler, Abram; Mahyar, Narges; O'Connor, Brendan (January 2022, ACM transactions on interactive intelligent systems)

Full Text Available

« Prev Next »

Search for: All records