Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 15, 2025Open Access

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing

View Full Paper
Ask AI
Bookmark
Share

Authors

ZLZehua LiuXLXiaolou LiLGLi Guo

Discussion

Loading...

Member takes

Overview

Analysis reveals how large language models improve recognition accuracy in visual speech recognition, suggesting innovative approaches.

Key Points

  • Integrating large language models significantly enhances visual speech recognition performance, applying scaling laws.
  • Key findings include improved recognition accuracy through context-aware decoding and iterative polishing techniques.
  • Experimental results confirm the effectiveness of LLM scaling, adding context to the decoding process for improved outcomes.
  • These insights highlight innovative methods to maximize the potential of large language models in visual speech recognition.

Cite This Study

Liu et al. (2025) studied this question.

synapsesocial.com/papers/68f02c7d616531447b5f9415https://doi.org/10.48550/arxiv.2506.02012
View Full Paper
Ask AI
Bookmark
Share