What question did this study set out to answer?

This research aims to enhance the efficiency of attention mechanisms in large language models by reducing redundancy between layers.

May 8, 2026Open Access

Cross-layer Attention Sharing for Pre-trained Large Language Models

Key Points

This research aims to enhance the efficiency of attention mechanisms in large language models by reducing redundancy between layers.
Introduced LiSA, a substitute for self-attention using tiny feed-forward networks and low-rank matrices.
Conducted comprehensive analyses across various large language models to identify redundancy in attention patterns.
Evaluated LiSA's performance on 13 benchmarks measuring accuracy and perplexity.
LiSA reduces redundant attention calculations within 53% − 84% of the layers.
Achieved a 6 × compression of Q and K matrices within the attention mechanism.
Maximum throughput improvements of 19.5%, 32.3%, and 40.1% for LLaMA3-8B, LLaMA2-7B, and LLaMA2-13B, respectively.

Abstract

Abstract To enhance the efficiency of the attention mechanism within large language models (LLMs), previous works primarily compress the Key-Value cache or group attention heads, while largely overlooking redundancy between layers. Our comprehensive analyses across various LLMs show that highly similar attention patterns persist within most layers. It’s intuitive to reduce the redundancy by sharing attention weights across layers. However, further analysis reveals two challenges: (1) Directly sharing the weight matrix without carefully rearranging the attention heads proves to be ineffective; (2) Shallow layers are vulnerable to small deviations in attention weights. Driven by these insights, we introduce LiSA, a lightweight substitute for self-attention in well-trained LLMs. LiSA employs tiny feed-forward networks to align attention heads between adjacent layers and low-rank matrices to approximate differences in layer-wise attention weights. Evaluations encompassing 13 typical benchmarks demonstrate that LiSA maintains high response quality in terms of accuracy and perplexity while reducing redundant attention calculations within 53% −84% of the total layers. Our implementations of LiSA achieve a 6 × compression of Q and K matrices within the attention mechanism, with maximum throughput improvements 19.5%, 32.3%, and 40.1% for LLaMA3-8B, LLaMA2-7B, and LLaMA2-13B, respectively. Our code is available at https://github.com/takagi97/lisa.

Read Full Paperexternally

Perguntar à IA

Bookmark

View Full Paper