Training Discrete Deep Generative Models via Gapped Straight-Through Estimator

Fan, Ting-Han and

Citation Details

While deep generative models have succeeded in image processing, natural language processing, and reinforcement learning, training that involves discrete random variables remains challenging due to the high variance of its gradient estimation process. Monte Carlo is a common solution used in most variance reduction approaches. However, this involves time-consuming resampling and multiple function evaluations. We propose a Gapped Straight-Through (GST) estimator to reduce the variance without incurring resampling overhead. This estimator is inspired by the essential properties of Straight-Through Gumbel-Softmax. We determine these properties and show via an ablation study that they are essential. Experiments demonstrate that the proposed GST estimator enjoys better performance compared to strong baselines on two discrete deep generative modeling tasks, MNIST-VAE and ListOps. more »

Award ID(s):: 1919452

PAR ID:: 10380235

Author(s) / Creator(s):: Fan, Ting-Han and

Editor(s):: Chaudhuri, Kamalika and

Date Published:: 2022-07-23

Journal Name:: Proceedings of Machine Learning Research

Volume:: 162

ISSN:: 2640-3498

Page Range / eLocation ID:: 6059-6073

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Journal Article:
The DOI is not currently available.

More Like this