The sequential information bottleneck (sIB) algorithm clusters co-occurrence data such as text documents vs. words. We introduce a variant that models sparse co-occurrence data by a generative process. This turns the objective function of sIB, mutual information, into a Bayes factor, while keeping it intact asymptotically, for non-sparse data. Experimental performance of the new algorithm is comparable to the original sIB for large data sets, and better for smaller, sparse sets.
No takes yet. Share an insight, caveat, or question.
Peltonen et al. (2004) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: