PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 1, 19993,957 citationsOpen Access

Probabilistic latent semantic indexing

THThomas Hofmann

Key Points

  • To develop a probabilistic method for automated document indexing that improves semantic understanding of text.
  • Utilized a statistical latent class model for factor analysis of count data.
  • Applied the Expectation Maximization algorithm to train on a corpus of text documents.
  • Conducted retrieval experiments on multiple test collections.
  • Achieved significant performance gains in document retrieval compared to direct term matching methods.
  • Demonstrated superior results over standard Latent Semantic Indexing (LSI).
  • Combination of models with different dimensionalities enhanced indexing effectiveness.

Abstract

Probabilistic Latent Semantic Indexing is a novel approach to automated document indexing which is based on a statistical latent class model for factor analysis of count data. Fitted from a training corpus of text documents by a generalization of the Expectation Maximization algorithm, the utilized model is able to deal with domain speci c synonymy as well as with polysemous words. In contrast to standard Latent Semantic Indexing LSI by Singular Value Decomposition, the probabilistic variant has a solid statistical foundation and de nes a proper generative data model. Retrieval experiments on a number of test collections indicate substantial performance gains over direct term matching metho d s a s w ell as over LSI. In particular, the combination of models with di erent dimensionalities has proven to be advantageous.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Thomas Hofmann (1999) studied this question.

synapsesocial.com/papers/69df087bb46aaead81614075https://doi.org/10.1145/312624.312649
Ask AI
Helpful
Bookmark
Share
View Full Paper