We introduce a new unsupervised task, spoken language modeling: the learning linguistic representations from raw audio signals without any labels, along the Zero Resource Speech Benchmark 2021: a suite of 4 black-box, zero-shot probing for the quality of the learned models at 4 linguistic levels:, lexicon, syntax and semantics. We present the results and analyses a composite baseline made of the concatenation of three unsupervised: self-supervised contrastive representation learning (CPC), clustering(k-means) and language modeling (LSTM or BERT). The language models learn on basis of the pseudo-text derived from clustering the learned. This simple pipeline shows better than chance performance on four metrics, demonstrating the feasibility of spoken language modeling raw speech. It also yields worse performance compared to text-based'topline' systems trained on the same data, delineating the space to be by more sophisticated end-to-end models.
No takes yet. Share an insight, caveat, or question.
Nguyen et al. (2020) studied this question.