Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 5:00 PM ET until 8:00 PM ET on Friday, September 11 due to maintenance. We apologize for the inconvenience.


Search for: All records

Creators/Authors contains: "Cloninger, Alexander"

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. Free, publicly-accessible full text available April 15, 2027
  2. Free, publicly-accessible full text available April 20, 2027
  3. Free, publicly-accessible full text available January 26, 2027
  4. Free, publicly-accessible full text available October 15, 2026
  5. Ensuring Conditional Independence (CI) constraints is pivotal for the development of fair and trustworthy machine learning models. In this paper, we introduce OTClean, a framework that harnesses optimal transport theory for data repair under CI constraints. Optimal transport theory provides a rigorous framework for measuring the discrepancy between probability distributions, thereby ensuring control over data utility. We formulate the data repair problem concerning CIs as a Quadratically Constrained Linear Program (QCLP) and propose an alternating method for its solution. However, this approach faces scalability issues due to the computational cost associated with computing optimal transport distances, such as the Wasserstein distance. To overcome these scalability challenges, we reframe our problem as a regularized optimization problem, enabling us to develop an iterative algorithm inspired by Sinkhorn's matrix scaling algorithm, which efficiently addresses high-dimensional and large-scale data. Through extensive experiments, we demonstrate the efficacy and efficiency of our proposed methods, showcasing their practical utility in real-world data cleaning and preprocessing tasks. Furthermore, we provide comparisons with traditional approaches, highlighting the superiority of our techniques in terms of preserving data utility while ensuring adherence to the desired CI constraints. 
    more » « less
  6. Abstract We propose the use of low bit-depth Sigma-Delta and distributed noise-shaping methods for quantizing the random Fourier features (RFFs) associated with shift-invariant kernels. We prove that our quantized RFFs—even in the case of $$1$$-bit quantization—allow a high-accuracy approximation of the underlying kernels, and the approximation error decays at least polynomially fast as the dimension of the RFFs increases. We also show that the quantized RFFs can be further compressed, yielding an excellent trade-off between memory use and accuracy. Namely, the approximation error now decays exponentially as a function of the bits used. The quantization algorithms we propose are intended for digitizing RFFs without explicit knowledge of the application for which they will be used. Nevertheless, as we empirically show by testing the performance of our methods on several machine learning tasks, our method compares favourably with other state-of-the-art quantization methods. 
    more » « less
  7. Abstract Orientations of active antithetic faults can provide useful constraints on in situ strength of the seismogenic crust. We use LINSCAN, a new unsupervised learning algorithm for identifying quasi‐linear clusters of earthquakes, to map small‐scale strike‐slip faults in the Anza‐Borrego shear zone, Southern California. We identify 332 right‐ and left‐lateral faults having lengths between 0.1 and 3 km. The dihedral angles between all possible pairs of conjugate faults are nearly normally distributed around 70°, with a standard deviation of ∼30°. The observed dihedral angles are larger than those expected assuming optimal fault orientations and the coefficient of friction of 0.6–0.8, but similar to the distribution previously reported for the Ridgecrest area in the Eastern California Shear Zone. We show that the observed fault orientations can be explained by fault rotation away from the principal shortening axis due to a cumulated tectonic strain. 
    more » « less
  8. Abstract In this paper we study supervised learning tasks on the space of probability measures. We approach this problem by embedding the space of probability measures into $$L^2$$ L 2 spaces using the optimal transport framework. In the embedding spaces, regular machine learning techniques are used to achieve linear separability. This idea has proved successful in applications and when the classes to be separated are generated by shifts and scalings of a fixed measure. This paper extends the class of elementary transformations suitable for the framework to families of shearings, describing conditions under which two classes of sheared distributions can be linearly separated. We furthermore give necessary bounds on the transformations to achieve a pre-specified separation level, and show how multiple embeddings can be used to allow for larger families of transformations. We demonstrate our results on image classification tasks. 
    more » « less