PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 3, 20240 citationsOpen Access

Investigating Decoder-only Large Language Models for Speech-to-text Translation

View Full Paper
CHChao-Wei HuangHLHui LuHGHongyu Gong

Key Points

  • The proposed model achieves state-of-the-art performance on two benchmarks, CoVoST 2 and FLEURS, demonstrating its effectiveness.
  • Performance metrics indicate significant improvement in translation accuracy compared to existing models.
  • The analysis utilized a decoder-only architecture that efficiently generates text from encoded speech representations after fine-tuning techniques were applied efficiently for better results in diverse tasks. May enable the development of more accessible and effective communication tools.

Abstract

Large language models (LLMs), known for their exceptional reasoning capabilities, generalizability, and fluency across diverse domains, present a promising avenue for enhancing speech-related tasks. In this paper, we focus on integrating decoder-only LLMs to the task of speech-to-text translation (S2TT). We propose a decoder-only architecture that enables the LLM to directly consume the encoded speech representation and generate the text translation. Additionally, we investigate the effects of different parameter-efficient fine-tuning techniques and task formulation. Our model achieves state-of-the-art performance on CoVoST 2 and FLEURS among models trained without proprietary data. We also conduct analyses to validate the design choices of our proposed model and bring insights to the integration of LLMs to S2TT.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Huang et al. (2024) studied this question.

synapsesocial.com/papers/68e6191db6db6435875abe65https://doi.org/10.48550/arxiv.2407.03169
Ask AI
Helpful
Bookmark
Share
View Full Paper