Key points are not available for this paper at this time.
Speech segmentation is an essential part of speech translation (ST) systems in real-world scenarios. Since most ST models are designed to process speech segments, long-form audio must be partitioned into shorter segments before translation. Recently, data-driven approaches for the speech segmentation task have been developed. Although these approaches improve overall translation quality, a performance gap exists due to a mismatch between the models and ST systems.In addition, the prior works require large self-supervised speech models, which consume significant computational resources.In this work, we propose a segmentation model that achieves better speech translation quality with a small model size. We propose an ASR-with-punctuation task as an effective pre-training strategy for the segmentation model. We also show that proper integration of the speech segmentation model into the underlying ST system is critical to improve overall translation quality at inference time.
Building similarity graph...
Analyzing shared references across papers
Loading...
Lee et al. (Sun,) studied this question.
www.synapsesocial.com/papers/68e59e92b6db643587538a80 — DOI: https://doi.org/10.21437/interspeech.2024-790
Jaesong Lee
So Yoon Kim
Hanbyul Kim
Building similarity graph...
Analyzing shared references across papers
Loading...
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: