PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 18, 20240 citationsOpen Access

TODM: Train Once Deploy Many Efficient Supernet-Based RNN-T Compression For On-Device ASR Models

View Full Paper
YSYuan ShangguanHYHaichuan YangDLDanni Li

Key Points

Key points are not available for this paper at this time.

Abstract

Automatic Speech Recognition (ASR) models need to be optimized for specific hardware before they can be deployed on devices. This can be done by tuning the model's hyperparameters or exploring variations in its architecture. Re-training and re-validating models after making these changes can be a resource-intensive task. This paper presents TODM (Train Once Deploy Many), a new approach to efficiently train many sizes of hardware-friendly on-device ASR models with comparable GPU-hours to that of a single training job. TODM leverages insights from prior work on Supernet, where Recurrent Neural Network Transducer (RNN-T) models share weights within a Supernet. It reduces layer sizes and widths of the Supernet to obtain subnetworks, making them smaller models suitable for all hardware types. We introduce a novel combination of three techniques to improve the outcomes of the TODM Supernet: adaptive dropout, an in-place Alpha-divergence knowledge distillation, and the use of ScaledAdam optimizer. We validate our approach by comparing Supernet-trained versus individually tuned Multi-Head State Space Model (MH-SSM) RNN-T using LibriSpeech. Results demonstrate that our TODM Supernet either matches or surpasses the performance of manually tuned models by up to a relative of 3% better in word error rate (WER), while efficiently keeping the cost of training many models at a small constant.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Shangguan et al. (2024) studied this question.

synapsesocial.com/papers/68e7397eb6db6435876b2ad1https://doi.org/10.1109/icassp48485.2024.10448025
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition2024
  2. 2Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition2024
  3. 3Dynamic Acoustic Model Architecture Optimization in Training for ASR2025
  4. 4T-SOT FNT: Streaming Multi-Talker ASR with Text-Only Domain Adaptation Capability2024
  5. 5HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation2025