NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

No Simple Answer to Data Complexity: An Examination of Instance-Level Complexity Metrics for Classification Tasks

Cook, Ryan A; Lalor, John P; Abbasi, Ahmed (April 2025, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics)

Free, publicly-accessible full text available April 30, 2026
Hierarchical Deep Document Model

https://doi.org/10.1109/TKDE.2024.3487523

Yang, Yi; Lalor, John P; Abbasi, Ahmed; Zeng, Daniel Dajun (January 2025, IEEE Transactions on Knowledge and Data Engineering)

Full Text Available
Constructing a Psychometric Testbed for Fair Natural Language Processing

https://doi.org/10.18653/v1/2021.emnlp-main.304

Abbasi, Ahmed; Dobolyi, David; Lalor, John P.; Netemeyer, Richard G.; Smith, Kendall; Yang, Yi (November 2021, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing)

Full Text Available
Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?

https://doi.org/10.18653/v1/2021.acl-long.346

Rodriguez, Pedro; Barrow, Joe; Hoyle, Alexander Miserlis; Lalor, John P.; Jia, Robin; Boyd-Graber, Jordan (January 2021, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers))

Leaderboards are widely used in NLP and push the field forward. While leaderboards are a straightforward ranking of NLP models, this simplicity can mask nuances in evaluation items (examples) and subjects (NLP models). Rather than replace leaderboards, we advocate a re-imagining so that they better highlight if and where progress is made. Building on educational testing, we create a Bayesian leaderboard model where latent subject skill and latent item difficulty predict correct responses. Using this model, we analyze the ranking reliability of leaderboards. Afterwards, we show the model can guide what to annotate, identify annotation errors, detect overfitting, and identify informative examples. We conclude with recommendations for future benchmark tasks.
more » « less
Full Text Available

Search for: All records