Zeroth-Order Optimization Finds Flat Minima

Zhang, Liang; Li, Bingcong; Thekumparampil, Kiran_Koshy; Oh, Sewoong; Muehlebach, Michael; He, Niao

Citation Details

Zeroth-order methods are extensively used in machine learning applications where gradients are infeasible or expensive to compute, such as black-box attacks, reinforcement learning, and language model fine-tuning. Existing optimization theory focuses on convergence to an arbitrary stationary point, but less is known about the implicit regularization that provides a fine-grained characterization of which particular solutions are reached. This paper shows that zeroth-order optimization with the standard two-point estimator favors solutions with small trace of Hessian, a measure widely used to distinguish between sharp and flat minima. The authors provide convergence rates of zeroth-order optimization to approximate flat minima for convex and sufficiently smooth functions, defining flat minima as minimizers that achieve the smallest trace of Hessian among all optimal solutions. Experiments on binary classification tasks with convex losses and language model fine-tuning support the theoretical findings. more »

Award ID(s):: 2505865

PAR ID:: 10631828

Author(s) / Creator(s):: Zhang, Liang; Li, Bingcong; Thekumparampil, Kiran_Koshy; Oh, Sewoong; Muehlebach, Michael; He, Niao

Publisher / Repository:: https://doi.org/10.48550/arXiv.2506.05454

Date Published:: 2025-06-05

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript
Conference Paper:
The DOI is not currently available.

More Like this