Distinct Elements in Streams: An Algorithm for the (Text) Book

Chakraborty, Sourav; Vinodchandran, N. V.; Meel, Kuldeep S.

Citation Details

Given a data stream 𝒟 = ⟨ a₁, a₂, …, a_m ⟩ of m elements where each a_i ∈ [n], the Distinct Elements problem is to estimate the number of distinct elements in 𝒟. Distinct Elements has been a subject of theoretical and empirical investigations over the past four decades resulting in space optimal algorithms for it. All the current state-of-the-art algorithms are, however, beyond the reach of an undergraduate textbook owing to their reliance on the usage of notions such as pairwise independence and universal hash functions. We present a simple, intuitive, sampling-based space-efficient algorithm whose description and the proof are accessible to undergraduates with the knowledge of basic probability theory. more »

Award ID(s):: 2130608

PAR ID:: 10461957

Author(s) / Creator(s):: Chakraborty, Sourav; Vinodchandran, N. V.; Meel, Kuldeep S.

Date Published:: 2022-09-01

Journal Name:: 30th Annual European Symposium on Algorithms (ESA 2022)

Volume:: 244

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this