TY - GEN
T1 - A batch normalized inference network keeps the KL vanishing away
AU - Zhu, Qile
AU - Bi, Wei
AU - Liu, Xiaojiang
AU - Ma, Xiyao
AU - Li, Xiaolin
AU - Wu, Dapeng
PY - 2020
Y1 - 2020
N2 - Variational Autoencoder (VAE) is widely used as a generative model to approximate a model's posterior on latent variables by combining the amortized variational inference and deep neural networks. However, when paired with strong autoregressive decoders, VAE often converges to a degenerated local optimum known as “posterior collapse”. Previous approaches consider the Kullback-Leibler divergence (KL) individual for each datapoint. We propose to let the KL follow a distribution across the whole dataset, and analyze that it is sufficient to prevent posterior collapse by keeping the expectation of the KL's distribution positive. Then we propose Batch Normalized-VAE (BN-VAE), a simple but effective approach to set a lower bound of the expectation by regularizing the distribution of the approximate posterior's parameters. Without introducing any new model component or modifying the objective, our approach can avoid the posterior collapse effectively and efficiently. We further show that the proposed BN-VAE can be extended to conditional VAE (CVAE). Empirically, our approach surpasses strong autoregressive baselines on language modeling, text classification and dialogue generation, and rivals more complex approaches while keeping almost the same training time as VAE. © 2020 Association for Computational Linguistics
AB - Variational Autoencoder (VAE) is widely used as a generative model to approximate a model's posterior on latent variables by combining the amortized variational inference and deep neural networks. However, when paired with strong autoregressive decoders, VAE often converges to a degenerated local optimum known as “posterior collapse”. Previous approaches consider the Kullback-Leibler divergence (KL) individual for each datapoint. We propose to let the KL follow a distribution across the whole dataset, and analyze that it is sufficient to prevent posterior collapse by keeping the expectation of the KL's distribution positive. Then we propose Batch Normalized-VAE (BN-VAE), a simple but effective approach to set a lower bound of the expectation by regularizing the distribution of the approximate posterior's parameters. Without introducing any new model component or modifying the objective, our approach can avoid the posterior collapse effectively and efficiently. We further show that the proposed BN-VAE can be extended to conditional VAE (CVAE). Empirically, our approach surpasses strong autoregressive baselines on language modeling, text classification and dialogue generation, and rivals more complex approaches while keeping almost the same training time as VAE. © 2020 Association for Computational Linguistics
UR - https://www.scopus.com/pages/publications/85105994674
UR - https://www.scopus.com/record/pubmetrics.uri?eid=2-s2.0-85105994674&origin=recordpage
U2 - 10.18653/v1/2020.acl-main.235
DO - 10.18653/v1/2020.acl-main.235
M3 - RGC 32 - Refereed conference paper (with host publication)
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 2636
EP - 2649
BT - Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics
A2 - Jurafsky, Dan
A2 - Chai, Joyce
A2 - Schluter, Natalie
A2 - Tetreault, Joel
PB - Association for Computational Linguistics
T2 - 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020
Y2 - 5 July 2020 through 10 July 2020
ER -