Tweeki: Linking Named Entities on Twitter to a Knowledge Graph

Harandizadeh, Bahareh; Singh, Sameer

doi:10.18653/v1/2020.wnut-1.29

Citation Details

Tweeki: Linking Named Entities on Twitter to a Knowledge Graph

To identify what entities are being talked about in tweets, we need to automatically link named entities that appear in tweets to structured KBs like WikiData. Existing approaches often struggle with such short, noisy texts, or their complex design and reliance on supervision make them brittle, difficult to use and maintain, and lose significance over time. Further, there is a lack of a large, linked corpus of tweets to aid researchers, along with lack of gold dataset to evaluate the accuracy of entity linking. In this paper, we introduce (1) Tweeki, an unsupervised, modular entity linking system for Twitter, (2) TweekiData, a large, automatically-annotated corpus of Tweets linked to entities in WikiData, and (3) TweekiGold, a gold dataset for entity linking evaluation. Through comprehensive analysis, we show that Tweeki is comparable to the performance of recent state-of-the-art entity linkers models, the dataset is of high quality, and a use case of how the dataset can be used to improve downstream tasks in social media analysis (geolocation prediction). more »

Award ID(s):: 1817183

PAR ID:: 10291541

Author(s) / Creator(s):: Harandizadeh, Bahareh; Singh, Sameer

Date Published:: 2020-10-01

Journal Name:: Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020)

Page Range / eLocation ID:: 222 to 231

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
https://doi.org/10.18653/v1/2020.wnut-1.29

More Like this