Skip to main navigation Skip to search Skip to main content

A batch normalized inference network keeps the KL vanishing away

  • Qile Zhu
  • , Wei Bi
  • , Xiaojiang Liu
  • , Xiyao Ma
  • , Xiaolin Li
  • , Dapeng Wu

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Variational Autoencoder (VAE) is widely used as a generative model to approximate a model's posterior on latent variables by combining the amortized variational inference and deep neural networks. However, when paired with strong autoregressive decoders, VAE often converges to a degenerated local optimum known as “posterior collapse”. Previous approaches consider the Kullback-Leibler divergence (KL) individual for each datapoint. We propose to let the KL follow a distribution across the whole dataset, and analyze that it is sufficient to prevent posterior collapse by keeping the expectation of the KL's distribution positive. Then we propose Batch Normalized-VAE (BN-VAE), a simple but effective approach to set a lower bound of the expectation by regularizing the distribution of the approximate posterior's parameters. Without introducing any new model component or modifying the objective, our approach can avoid the posterior collapse effectively and efficiently. We further show that the proposed BN-VAE can be extended to conditional VAE (CVAE). Empirically, our approach surpasses strong autoregressive baselines on language modeling, text classification and dialogue generation, and rivals more complex approaches while keeping almost the same training time as VAE. © 2020 Association for Computational Linguistics
Original languageEnglish
Title of host publicationProceedings of the 58th Annual Meeting of the Association for Computational Linguistics
EditorsDan Jurafsky, Joyce Chai, Natalie Schluter, Joel Tetreault
PublisherAssociation for Computational Linguistics
Pages2636-2649
DOIs
Publication statusPublished - 2020
Externally publishedYes
Event58th Annual Meeting of the Association for Computational Linguistics, ACL 2020 - Virtual, Online, United States
Duration: 5 Jul 202010 Jul 2020

Publication series

NameProceedings of the Annual Meeting of the Association for Computational Linguistics
ISSN (Print)0736-587X

Conference

Conference58th Annual Meeting of the Association for Computational Linguistics, ACL 2020
PlaceUnited States
CityVirtual, Online
Period5/07/2010/07/20

Fingerprint

Dive into the research topics of 'A batch normalized inference network keeps the KL vanishing away'. Together they form a unique fingerprint.

Cite this