C-DIEGO: An Algorithm with Near-Optimal Sample Complexity for Distributed, Streaming PCA

Zulqarnain, Muhammad; Gang, Arpita; Bajwa, Waheed U.

doi:10.1109/CISS56502.2023.10089668

Citation Details

C-DIEGO: An Algorithm with Near-Optimal Sample Complexity for Distributed, Streaming PCA

The accuracy of many downstream machine learning algorithms is tied to the training data having uncorrelated features. With the modern-day data often being streaming in nature, geographically distributed, and having large dimensions, it is paramount to apply both uncorrelated feature learning and dimensionality reduction techniques in this scenario. Principal Component Analysis (PCA) is a state-of-the-art tool that simultaneously yields uncorrelated features and reduces data dimensions by projecting data onto the eigenvectors of the population covariance matrix. This paper introduces a novel algorithm called Consensus-DIstributEd Generalized Oja (C-DIEGO), which is based on Oja's method, to estimate the dominant eigenvector of a population covariance matrix in a distributed, streaming setting. The algorithm considers a distributed network of arbitrarily connected nodes without a central coordinator and assumes data samples continuously arrive at the individual nodes in a streaming manner. It is established in the paper that C-DIEGO can achieve an order-optimal convergence rate if nodes in the network are allowed to have enough consensus rounds per algorithmic iteration. Numerical results are also reported in the paper that showcase the efficacy of the proposed algorithm. more »

Award ID(s):: 2148104 1940074 1907658

PAR ID:: 10409012

Author(s) / Creator(s):: Zulqarnain, Muhammad; Gang, Arpita; Bajwa, Waheed U.

Date Published:: 2023-03-22

Journal Name:: 2023 57th Annual Conference on Information Sciences and Systems (CISS)

Page Range / eLocation ID:: 1 to 6

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
https://doi.org/10.1109/CISS56502.2023.10089668

More Like this