NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Zeroth-Order Optimization Finds Flat Minima

Zhang, Liang; Li, Bingcong; Thekumparampil, Kiran_Koshy; Oh, Sewoong; Muehlebach, Michael; He, Niao (June 2025, https://doi.org/10.48550/arXiv.2506.05454)

Zeroth-order methods are extensively used in machine learning applications where gradients are infeasible or expensive to compute, such as black-box attacks, reinforcement learning, and language model fine-tuning. Existing optimization theory focuses on convergence to an arbitrary stationary point, but less is known about the implicit regularization that provides a fine-grained characterization of which particular solutions are reached. This paper shows that zeroth-order optimization with the standard two-point estimator favors solutions with small trace of Hessian, a measure widely used to distinguish between sharp and flat minima. The authors provide convergence rates of zeroth-order optimization to approximate flat minima for convex and sufficiently smooth functions, defining flat minima as minimizers that achieve the smallest trace of Hessian among all optimal solutions. Experiments on binary classification tasks with convex losses and language model fine-tuning support the theoretical findings.
more » « less
Free, publicly-accessible full text available June 5, 2026
Preconditioned Sharpness-Aware Minimization: Unifying Analysis and a Novel Learning Algorithm

Zhang, Yilang; Li, Bingcong; Giannakis, GB (April 2025, IEEE)

Free, publicly-accessible full text available April 12, 2026
Preconditioned Sharpness-Aware Minimization: Unifying Analysis and a Novel Learning Algorithm

https://doi.org/10.1109/ICASSP49660.2025.10889586

Zhang, Yilang; Li, Bingcong; Giannakis, Georgios B (April 2025, The Illinois labor letter)

Free, publicly-accessible full text available April 6, 2026
Meta-Learning with Versatile Loss Geometries for Fast Adaptation Using Mirror Descent

Zhang, Yilang; Li, Bingcong; Giannakis, Georgios B (April 2024, IEEE International Conference on Acoustics, Speech, and Signal Processing)

Full Text Available
Meta-Learning With Versatile Loss Geometries for Fast Adaptation Using Mirror Descent

https://doi.org/10.1109/ICASSP48485.2024.10448144

Zhang, Yilang; Li, Bingcong; Giannakis, Georgios B (April 2024, IEEE)

Utilizing task-invariant prior knowledge extracted from related tasks, meta-learning is a principled framework that empowers learning a new task especially when data records are limited. A fundamental challenge in meta-learning is how to quickly "adapt" the extracted prior in order to train a task-specific model within a few optimization steps. Existing approaches deal with this challenge using a preconditioner that enhances convergence of the per-task training process. Though effective in representing locally a quadratic training loss, these simple linear preconditioners can hardly capture complex loss geometries. The present contribution addresses this limitation by learning a nonlinear mirror map, which induces a versatile distance metric to enable capturing and optimizing a wide range of loss geometries, hence facilitating the per-task training. Numerical tests on few-shot learning datasets demonstrate the superior expressiveness and convergence of the advocated approach.
more » « less
Full Text Available
Enhancing Sharpness-Aware Optimization Through Variance Suppression

Li, Bingcong; Giannakis, Georgios B (December 2023, Proceedings of Neural Information Processing Systems (NeurIPS))

Full Text Available
Enhancing Sharpness-Aware Optimization Through Variance Suppression

Li, Bingcong; Giannakis, Georgios B (November 2023, Proceedings of Neural Information Processing Systems (NeurIPS))

Full Text Available
Conic Descent Redux for Memory-Efficient Optimization

https://doi.org/10.1109/IEEECONF59524.2023.10476894

Li, Bingcong; Giannakis, Georgios B (October 2023, IEEE)

Conic programming has well-documented merits in a gamut of signal processing and machine learning tasks. This contribution revisits a recently developed first-order conic descent (CD) solver, and advances it in three aspects: intuition, theory, and algorithmic implementation. It is found that CD can afford an intuitive geometric derivation that originates from the dual problem. This opens the door to novel algorithmic designs, with a momentum variant of CD, momentum conic descent (MOCO) exemplified. Diving deeper into the dual behavior CD and MOCO reveals: i) an analytically justified stopping criterion; and, ii) the potential to design preconditioners to speed up dual convergence. Lastly, to scale semidefinite programming (SDP) especially for low-rank solutions, a memory efficient MOCO variant is developed and numerically validated.
more » « less
Full Text Available
Surrogate modeling for Bayesian optimization beyond a single Gaussian process

https://doi.org/10.1109/TPAMI.2023.3264741

Lu, Qin; Polyzos, Konstantinos D.; Li, Bingcong; Giannakis, Georgios B. (September 2023, IEEE Transactions on Pattern Analysis and Machine Intelligence)

Bayesian optimization (BO) has well-documented merits for optimizing black-box functions with an expensive evaluation cost. Such functions emerge in applications as diverse as hyperparameter tuning, drug discovery, and robotics. BO hinges on a Bayesian surrogate model to sequentially select query points so as to balance exploration with exploitation of the search space. Most existing works rely on a single Gaussian process (GP) based surrogate model, where the kernel function form is typically preselected using domain knowledge. To bypass such a design process, this paper leverages an ensemble (E) of GPs to adaptively select the surrogate model fit on-the-fly, yielding a GP mixture posterior with enhanced expressiveness for the sought function. Acquisition of the next evaluation input using this EGP-based function posterior is then enabled by Thompson sampling (TS) that requires no additional design parameters. To endow function sampling with scalability, random feature-based kernel approximation is leveraged per GP model. The novel EGP-TS readily accommodates parallel operation. To further establish convergence of the proposed EGP-TS to the global optimum, analysis is conducted based on the notion of Bayesian regret for both sequential and parallel settings. Tests on synthetic functions and real-world applications showcase the merits of the proposed method.
more » « less
Full Text Available
Scalable Bayesian Meta-Learning through Generalized Implicit Gradients

https://doi.org/10.1609/aaai.v37i9.26337

Zhang, Yilang; Li, Bingcong; Gao, Shijian; Giannakis, Georgios B. (June 2023, Proceedings of the AAAI Conference on Artificial Intelligence)

Meta-learning owns unique effectiveness and swiftness in tackling emerging tasks with limited data. Its broad applicability is revealed by viewing it as a bi-level optimization problem. The resultant algorithmic viewpoint however, faces scalability issues when the inner-level optimization relies on gradient-based iterations. Implicit differentiation has been considered to alleviate this challenge, but it is restricted to an isotropic Gaussian prior, and only favors deterministic meta-learning approaches. This work markedly mitigates the scalability bottleneck by cross-fertilizing the benefits of implicit differentiation to probabilistic Bayesian meta-learning. The novel implicit Bayesian meta-learning (iBaML) method not only broadens the scope of learnable priors, but also quantifies the associated uncertainty. Furthermore, the ultimate complexity is well controlled regardless of the inner-level optimization trajectory. Analytical error bounds are established to demonstrate the precision and efficiency of the generalized implicit gradient over the explicit one. Extensive numerical tests are also carried out to empirically validate the performance of the proposed method.
more » « less
Full Text Available

« Prev Next »

Search for: All records