Statistical language modeling techniques have successfully been applied to large source code corpora, yielding a variety of new software development tools, such as tools for code suggestion, improving readability, and API migration. A major issue with these techniques is that code introduces new vocabulary at a far higher rate than natural language, as new identifier names proliferate. Both large vocabularies and out-of-vocabulary issues severely affect Neural Language Models (NLMs) of source code, degrading their performance and rendering them unable to scale.
No takes yet. Share an insight, caveat, or question.
Karampatsis et al. (2020) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: