Speech and language therapy (SLT) for individuals with speech sound disorders (SSDs) requires tools that can make articulatory movements, particularly tongue motion, more visible and interpretable. Ultrasound tongue imaging (UTI) is a promising modality for this purpose, as it is non-invasive, safe, and does not interfere with natural speech. In this work, we investigate the use of recurrent neural networks for acoustic-to-articulatory inversion (AAI), directly mapping speech audio features to corresponding UTI frames. We implement a deep LSTM architecture and evaluate its performance under different loss functions. Our experiments compare pixel-based losses (MSE, MAE), perceptual losses (SSIM), and hybrid combinations. While the models failed to generalize effectively to the unseen data, we observed that the choice of loss function significantly influenced perceptual frame quality. MSE loss emphasized sharper pixel contrast but missed structural accuracy, whereas SSIM improved structure at the cost of sharpness. Hybrid losses, particularly SSIM+MAE with weighting, produced the most visually convincing reconstructions, despite modest quantitative metrics. These findings demonstrate both the limitations of LSTM-based AAI for UTI generation and the importance of carefully designed loss functions. They suggest that future work should integrate hybrid loss formulations with alternative architectures to achieve more reliable mappings from speech to tongue imaging.
Dadgar et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: