A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces

Barbosa, Juliana Silva (ORCID:0009000442915070); Gondhali, Ulhas (ORCID:0000000255416906); Petrossian, Gohar (ORCID:0000000207082833); Sharma, Kinshuk (ORCID:0009000241735380); Chakraborty, Sunandan (ORCID:0000000233316082); Jacquet, Jennifer (ORCID:0000000238078060); Freire, Juliana (ORCID:0000000339157075)

doi:10.1145/3725256

Citation Details

This content will become publicly available on June 17, 2026

A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces

Wildlife trafficking remains a critical global issue, significantly impacting biodiversity, ecological stability, and public health. Despite efforts to combat this illicit trade, the rise of e-commerce platforms has made it easier to sell wildlife products, putting new pressure on wild populations of endangered and threatened species. The use of these platforms also opens a new opportunity: as criminals sell wildlife products online, they leave digital traces of their activity that can provide insights into trafficking activities as well as how they can be disrupted. The challenge lies in finding these traces. Online marketplaces publish ads for a plethora of products, and identifying ads for wildlife-related products is like finding a needle in a haystack. Learning classifiers can automate ad identification, but creating them requires costly, time-consuming data labeling that hinders support for diverse ads and research questions. This paper addresses a critical challenge in the data science pipeline for wildlife trafficking analytics: generating quality labeled data for classifiers that select relevant data. While large language models (LLMs) can directly label advertisements, doing so at scale is prohibitively expensive. We propose a cost-effective strategy that leverages LLMs to generate pseudo labels for a small sample of the data and uses these labels to create specialized classification models. Our novel method automatically gathers diverse and representative samples to be labeled while minimizing the labeling costs. Our experimental evaluation shows that our classifiers achieve up to 95% F1 score, outperforming LLMs at a lower cost. We present real use cases that demonstrate the effectiveness of our approach in enabling analyses of different aspects of wildlife trafficking. more »

Award ID(s):: 2146306 2106888

PAR ID:: 10653924

Author(s) / Creator(s):: Barbosa, Juliana Silva; Gondhali, Ulhas; Petrossian, Gohar; Sharma, Kinshuk; Chakraborty, Sunandan; Jacquet, Jennifer; Freire, Juliana

Publisher / Repository:: ACM

Date Published:: 2025-06-17

Journal Name:: Proceedings of the ACM on Management of Data

Volume:: 3

Issue:: 3

ISSN:: 2836-6573

Page Range / eLocation ID:: 1 to 23

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
This content will become publicly available on June 17, 2026
Journal Article:
https://doi.org/10.1145/3725256

More Like this