PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 9, 20250 citationsOpen Access

Improved Algorithms for Kernel Matrix-Vector Multiplication Under Sparsity Assumptions

View Full Paper
PIPiotr IndykMKMichael KapralovKSKshiteej Sheth

Key Points

  • The algorithm achieves matrix-vector multiplication in time subquadratic in n while remaining linear in d, surpassing traditional methods.
  • Work shows that the sum of entries in the Gaussian kernel matrix scales linearly with n, facilitating faster computations.
  • Experimental validation confirms the effectiveness of the algorithm under sparsity assumptions in Gaussian kernel matrices.
  • This first subquadratic-time algorithm offers a breakthrough for processing attention matrices in large language models.

Abstract

Motivated by the problem of fast processing of attention matrices, we study fast algorithms for computing matrix-vector products for asymmetric Gaussian Kernel matrices K R^n n. K's columns are indexed by a set of n keys k₁, k₂, kₙ Rᵈ, rows by a set of n queries q₁, q₂, , qₙ Rᵈ, and its i, j entry is K₈₉ = e^-\|qᵢ-kⱼ\|₂²/2σ² for some bandwidth parameter σ>0. Given a vector x Rⁿ and error parameter ε>0, our task is to output a y Rⁿ such that \|Kx-y\|₂ ε\|x\|₂ in time subquadratic in n and linear in d. Our algorithms rely on the following modelling assumption about the matrices K: the sum of the entries of K scales linearly in n, as opposed to worst case quadratic growth. We validate this assumption experimentally, for Gaussian kernel matrices encountered in various settings such as fast attention computation in LLMs. We obtain the first subquadratic-time algorithm that works under this assumption, for unrestricted vectors.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Indyk et al. (2025) studied this question.

synapsesocial.com/papers/68e7103b90569dd607ee68e5https://doi.org/10.48550/arxiv.2507.23539
Ask AI
Helpful
Bookmark
Share
View Full Paper