Key points are not available for this paper at this time.
We empirically characterize the performance of discriminative and generative models for text classification. We find that although RNN-based generative are more powerful than their bag-of-words ancestors (e. g. , they account conditional dependencies across words in a document), they have higher error rates than discriminatively trained RNN models. However we find that generative models approach their asymptotic error rate more than their discriminative counterparts---the same pattern that Ng & (2001) proved holds for linear classification models that make more conditional independence assumptions. Building on this finding, we that RNN-based generative classification models will be more robust shifts in the data distribution. This hypothesis is confirmed in a series of in zero-shot and continual learning settings that show that models substantially outperform discriminative models.
Yogatama et al. (Mon,) studied this question.