Controlled comparative analysis demonstrates superior language generation in Transformer encoder-decoders over recurrent architectures, highlighting the critical role of self-attention mechanisms.
Sequence-to-sequence systems have evolved from recurrent neural networks to gated recurrent architectures with attention and, ultimately, Transformer encoder-decoder models. This work analyzes these architectural transitions through controlled comparisons of RNN, LSTM, attention-augmented recurrent models, encoder-only Transformers, and full encoder-decoder Transformers. The results show that recurrent memory and attention independently address major limitations of vanilla RNNs, while the complete Transformer encoder-decoder architecture produces the largest improvement. The tuned Transformer achieves validation perplexity of 5.37, compared with 20.08 for the tuned LSTM-with-attention model, corresponding to a 73.3% relative reduction. Qualitative generation results similarly show substantially improved semantic coherence and reduced repetitive output. The study provides a component-level analysis of how memory, attention, autoregressive decoding, and hyperparameter tuning contribute to modern sequence-to-sequence generation.
No takes yet. Share an insight, caveat, or question.
Chandana Dayapule (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: