Although the Transformer translation model In this work, we extend the Transformer model with a new context encoder to represent document-level context, which is then incorporated into the original encoder and decoder. As large-scale document-level parallel corpora are usually not available, we introduce a two-step training method to take full advantage of abundant sentence-level parallel corpora and limited document-level parallel corpora. Experiments on the NIST Chinese-English datasets and the IWSLT French-English datasets show that our approach improves over Transformer significantly.
No takes yet. Share an insight, caveat, or question.
Zhang et al. (2018) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: