PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 1, 2017Corpora201 citationsOpen Access

‘What is this corpus about?’: using topic modelling to explore a specialised corpus

View Full Paper
AMAkira MurakamiPTPaul ThompsonSHSusan Hunston

Key Points

  • The aim is to illustrate topic modelling as a machine learning technique for exploring a corpus and extracting meaningful insights.
  • Introduced topic modelling and its mechanism.
  • Described the model-building procedure and decisions involved.
  • Conducted analysis on topic emergence in academic papers and journal evolution.
  • Identified prominent topics in various parts of academic papers, showcasing diversity.
  • Investigated chronological changes in journal content over time.
  • Compared topic modelling's effectiveness against semantic annotation and keywords analysis.

Abstract

This paper introduces topic modelling, a machine learning technique that automatically identifies ‘topics’ in a given corpus. The paper illustrates its use in the exploration of a corpus of academic English. It first offers the intuitive explanation of the underlying mechanism of topic modelling and describes the procedure for building a model, including the decisions involved in the model-building process. The paper then explores the model. A topic in topic models is characterised by a set of co-occurring words, and we will demonstrate that such topics bring us rich insights into the nature of a corpus. As exemplary tasks, this paper identifies the prominent topics in different parts of papers, investigates the chronological change of a journal, and reveals different types of papers in the journal. The paper further compares topic modelling to two more traditional techniques in corpus linguistics, semantic annotation and keywords analysis, and highlights the strengths of topic modelling. We believe that topic modelling is particularly useful in the initial exploration of a corpus.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Murakami et al. (2017) studied this question.

synapsesocial.com/papers/69d96cb700ab073a27836840https://doi.org/10.3366/cor.2017.0118
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1An algorithm for suffix stripping1980 · 8,159 citations
  2. 2Proceedings of the 23rd international conference on Machine learning2006 · 2,593 citations
  3. 3Journal of Digital Humanities2013 · 154 citations
  4. 4Problems in investigating keyness, or clearing the undergrowth and marking out trails…2010 · 83 citations
  5. 5Dimensions of Register Variation1995 · 1,267 citations