Title: A Poincare Inequality and Consistency Results for Signal Sampling on Large Graphs
Large-scale graph machine learning is challenging as the complexity of learning models scales with the graph size. Subsampling the graph is a viable alternative, but sampling on graphs is nontrivial as graphs are non-Euclidean. Existing graph sampling techniques require not only computing the spectra of large matrices but also repeating these computations when the graph changes, e.g., grows. In this pa- per, we introduce a signal sampling theory for a type of graph limit—the graphon. We prove a Poincare ́ inequality for graphon signals and show that complements of node subsets satisfying this inequality are unique sampling sets for Paley-Wiener spaces of graphon signals. Exploiting connections with spectral clustering and Gaussian elimination, we prove that such sampling sets are consistent in the sense that unique sampling sets on a convergent graph sequence converge to unique sampling sets on the graphon. We then propose a related graphon signal sampling algorithm for large graphs, and demonstrate its good empirical performance on graph machine learning tasks.  more » « less
Award ID(s):
2134108
PAR ID:
10568532
Author(s) / Creator(s):
; ;
Publisher / Repository:
International Conference on Learning Representations (ICLR)
Date Published:
Format(s):
Medium: X
Location:
Vienna
Sponsoring Org:
National Science Foundation
More Like this
  1. The 𝑊 -random graphs provide a flexible framework for modeling large random networks. Using the Large Deviation Principle (LDP) for 𝑊 -random graphs from [19], we prove the LDP for the corresponding class of random symmetric Hilbert-Schmidt integral operators. Our main result describes how the eigenvalues and the eigenspaces of the integral operator are affected by large deviations in the underlying random graphon. To prove the LDP, we demonstrate continuous dependence of the spectral measures associated with integral operators on the corresponding graphons and use the Contraction Principle. To illustrate our results, we obtain leading order asymptotics of the eigenvalues of small-world and bipartite random graphs conditioned on atypical edge counts. These examples suggest several representative scenarios of how the eigenvalues and the eigenspaces are affected by large deviations. We discuss the implications of these observations for bifurcation analysis of Dynamical Systems and Graph Signal Processing. 
    more » « less
  2. This paper studies stochastic games on large graphs and their graphon limits. We propose a new formulation of graphon games based on a single typical player’s label-state distribution. In contrast, other recently proposed models of graphon games work directly with a continuum of players, which involves serious measure-theoretic technicalities. In fact, by viewing the label as a component of the state process, we show in our formulation that graphon games are a special case of mean field games, albeit with certain inevitable degeneracies and discontinuities that make most existing results on mean field games inapplicable. Nonetheless, we prove the existence of Markovian graphon equilibria under fairly general assumptions as well as uniqueness under a monotonicity condition. Most importantly, we show how our notion of graphon equilibrium can be used to construct approximate equilibria for large finite games set on any (weighted, directed) graph that converges in cut norm. The lack of players’ exchangeability necessitates a careful definition of approximate equilibrium, allowing heterogeneity among the players’ approximation errors, and we show how various regularity properties of the model inputs and underlying graphon lead naturally to different strengths of approximation. Funding: D. Lacker was partially supported by the Air Force Office of Scientific Research [Grant FA9550-19-1-0291] and the National Science Foundation [Award DMS-2045328]. 
    more » « less
  3. null (Ed.)
    Graphs are nowadays ubiquitous in the fields of signal processing and machine learning. As a tool used to express relationships between objects, graphs can be deployed to various ends: (i) clustering of vertices, (ii) semi-supervised classification of vertices, (iii) supervised classification of graph signals, and (iv) denoising of graph signals. However, in many practical cases graphs are not explicitly available and must therefore be inferred from data. Validation is a challenging endeavor that naturally depends on the downstream task for which the graph is learnt. Accordingly, it has often been difficult to compare the efficacy of different algorithms. In this work, we introduce several ease-to-use and publicly released benchmarks specifically designed to reveal the relative merits and limitations of graph inference methods. We also contrast some of the most prominent techniques in the literature. 
    more » « less
  4. In modern relational machine learning it is common to encounter large graphs that arise via interactions or similarities between observations in many domains. Further, in many cases the target entities for analysis are actually signals on such graphs. We propose to compare and organize such datasets of graph signals by using an earth mover’s distance (EMD) with a geodesic cost over the underlying graph. Typically, EMD is computed by optimizing over the cost of transporting one probability distribution to another over an underlying metric space. However, this is inefficient when computing the EMD between many signals. Here, we propose an unbalanced graph EMD that efficiently embeds the unbalanced EMD on an underlying graph into an L1 space, whose metric we call unbalanced diffusion earth mover’s distance (UDEMD). Next, we show how this gives distances between graph signals that are robust to noise. Finally, we apply this to organizing patients based on clinical notes, embedding cells modeled as signals on a gene graph, and organizing genes modeled as signals over a large cell graph. In each case, we show that UDEMD-based embeddings find accurate distances that are highly efficient compared to other methods. 
    more » « less
  5. Graph Signal Processing (GSP) offers a structured way to model and analyze complex data networks. However, a consistent challenge in real world applications is that the underlying graph topology is often unknown. While most existing graph learning methods focus on undirected graphs, specific relationships in network data, such as diffusion, weather data and social networks, sometimes necessitate a directed graph model. In this work, we propose a novel method for learning directed graph structures from observed signals by leveraging both signal smoothness and directional flow. We introduce a Dirichlet energy formulation specifically tailored to directed graphs, favoring smoothness and downwards flow of signals over the graph. We then develop the Perseus Measure to quantify how much the learned graphs obey the directional flow constraint. In addition, we present a novel random graph algorithm that builds a signal matrix alongside the graph, and we derive certain properties about the resulting Hierarchy Graphs. Experiments on both synthetic and real-world datasets show that our directed graph learning approach effectively captures directional relationships. 
    more » « less