PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 30, 201885 citationsOpen Access

The CMU Statistical Language Modeling Toolkit and its use in the 1994 ARPA CSR Evaluation

RRRoni Rosenfeld

Key Points

Key points are not available for this paper at this time.

Abstract

The Carnegie Mellon Statistical Language Modeling (CMU SLM) Toolkit is a set of Unix software tools designed to facilitate language modeling work in the research community. The package, including source code, is freely available for research purposes. As of December 1994, the toolkit is in active use by 23 research groups in 8 countries. It was recently used to process the 2.5 GB NAB corpus for the ARPA CSR community. In this paper, I first discuss the design principles and features of the toolkit. Then, I describe the composition of the NAB corpus, and report on the ngram statistics, standard vocabulary and language models created using the SLM tools.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Roni Rosenfeld (2018) studied this question.

synapsesocial.com/papers/6a21048bff76e8667816dbd4https://doi.org/10.1184/r1/6610448.v1
Ask AI
Helpful
Bookmark
Share
View Full Paper