Mixture modeling is a long established machine learning technique for learning large sets of multi-modal data.
While it is known that sequence-to-sequence models for dialog response generation suffer from the problem of low diversity, we hypothesize that it is because sequence-to-sequence models tend to learn a degenerate uni-modal distribution of responses.
We then propose to incorporate a mixture of decoders into sequence-to-sequence models and try to make each decoder learn specialized topics in order to improve the diversity of generated responses.
Our model is developed under the framework of conditional variational autoencoder (CVAE).
We evaluate our approach on an open domain chat corpus and show improvement over strong baselines in quantitative measures and human evaluation.