PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 13, 20244 citationsOpen Access

An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

View Full Paper
ZMZiyang MaGYGuanrou YangYYYifan Yang

Key Points

Key points are not available for this paper at this time.

Abstract

In this paper, we focus on solving one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Recent works have complex designs such as compressing the output temporally for the speech encoder, tackling modal alignment for the projector, and utilizing parameter-efficient fine-tuning for the LLM. We found that delicate designs are not necessary, while an embarrassingly simple composition of off-the-shelf speech encoder, LLM, and the only trainable linear projector is competent for the ASR task. To be more specific, we benchmark and explore various combinations of LLMs and speech encoders, leading to the optimal LLM-based ASR system, which we call SLAM-ASR. The proposed SLAM-ASR provides a clean setup and little task-specific design, where only the linear projector is trained. To the best of our knowledge, SLAM-ASR achieves the best performance on the Librispeech benchmark among LLM-based ASR models and even outperforms the latest LLM-based audio-universal model trained on massive pair data. Finally, we explore the capability emergence of LLM-based ASR in the process of modal alignment. We hope that our study can facilitate the research on extending LLM with cross-modality capacity and shed light on the LLM-based ASR community.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ma et al. (2024) studied this question.

synapsesocial.com/papers/68e79585b6db6435877064c7https://doi.org/10.48550/arxiv.2402.08846
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Prompting Large Language Models with Speech Recognition Abilities2024 · 106 citations
  2. 2SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings2025
  3. 3Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets2024
  4. 4A Comprehensive Solution to Connect Speech Encoder and Large Language Model for ASR2024
  5. 5Leveraging Large Language Models for Exploiting ASR Uncertainty2024 · 11 citations