Robust Multi-Agent Bandits Over Undirected Graphs

Vial, Daniel; Shakkottai, Sanjay; Srikant, R.

doi:10.1145/3570614

Citation Details

Robust Multi-Agent Bandits Over Undirected Graphs

We consider a multi-agent multi-armed bandit setting in which n honest agents collaborate over a network to minimize regret but m malicious agents can disrupt learning arbitrarily. Assuming the network is the complete graph, existing algorithms incur O((m + K/n) łog (T) / Δ ) regret in this setting, where K is the number of arms and Δ is the arm gap. For m łl K, this improves over the single-agent baseline regret of O(Kłog(T)/Δ). In this work, we show the situation is murkier beyond the case of a complete graph. In particular, we prove that if the state-of-the-art algorithm is used on the undirected line graph, honest agents can suffer (nearly) linear regret until time is doubly exponential in K and n . In light of this negative result, we propose a new algorithm for which the i -th agent has regret O(( dmal (i) + K/n) łog(T)/Δ) on any connected and undirected graph, where dmal(i) is the number of i 's neighbors who are malicious. Thus, we generalize existing regret bounds beyond the complete graph (where dmal(i) = m), and show the effect of malicious agents is entirely local (in the sense that only the dmal (i) malicious agents directly connected to i affect its long-term regret). more »

Award ID(s):: 2207547 2106801 1934986

PAR ID:: 10410558

Author(s) / Creator(s):: Vial, Daniel; Shakkottai, Sanjay; Srikant, R.

Date Published:: 2022-12-01

Journal Name:: Proceedings of the ACM on Measurement and Analysis of Computing Systems

Volume:: 6

Issue:: 3

ISSN:: 2476-1249

Page Range / eLocation ID:: 1 to 57

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Journal Article:
https://doi.org/10.1145/3570614

More Like this