Benchmarking evaluation demonstrates robust translation and speech synthesis performance across eight languages on CPU hardware, highlighting the feasibility of offline voice translation.
Offline speech-to-speech translation can support privacy-sensitive and connectivity-constrained applications, but multilingual deployment on local hardware requires balancing recognition accuracy, language coverage, and latency. We present a modular, fully offline, turn-based system that combines language identification, confidence-aware automatic speech recognition routing, multilingual machine translation, and language-aware speech synthesis. Evaluation covers Hindi, Telugu, Tamil, Bengali, French, German, Arabic, and Japanese, with English as the common source/target language, using FLEURS, FLORES-200, and CVSS. The selected routing threshold (τ = 0.70) achieved 97.65% routing accuracy and substantially reduced recognition error across the four Indic languages compared with Whisper-only recognition. Clean-versus-cascaded translation experiments confirmed measurable propagation of recognition errors into machine translation. Across the five CVSS-supported source-to-English directions, end-to-end ASR-BLEU ranged from 31.4 to 39.2, while eight-language speech-synthesis evaluation produced mean opinion scores of 3.55–4.12 and back-ASR WER of 6.8–10.8%. The mean model-only CPU latency was 25.41 s, with translation and speech synthesis accounting for most of the runtime. These results characterize the accuracy–quality–latency trade-offs of modular offline multilingual speech translation under local CPU execution.
No takes yet. Share an insight, caveat, or question.
Gowroju et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: