PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 16, 2026Statistics and Computing0 citationsOpen Access

Approximate learning of parsimonious Bayesian context trees

DGDaniyar GhaniNHNicholas A. HeardFPFrancesco Sanna Passino

Key Points

  • The aim is to develop a Bayesian modeling framework that effectively captures complex dependence structures in categorical sequences while maintaining memory efficiency.
  • Introduced parsimonious Bayesian context trees as variable-order Markov models.
  • Utilized conjugate prior distributions to reduce parameter counts.
  • Employed a model-based agglomerative clustering procedure for approximate inference and processing.
  • Tested the framework on synthetic and real-world datasets, including protein sequences and computer malware traces.
  • The proposed framework shows improved predictive performance over traditional fixed-order Markov models.
  • Fewer parameters were required to achieve effective modeling of dependencies.
  • Outperformed existing sequence models when applied to real protein sequences and honeypot terminal sessions.

Abstract

Abstract Models for categorical sequences typically assume exchangeable or first-order dependent sequence elements. These are common assumptions, for example, in models of computer malware traces and protein sequences. Although such simplifying assumptions lead to computational tractability, these models fail to capture long-range, complex dependence structures that may be harnessed for greater predictive power. To this end, a Bayesian modelling framework is proposed to parsimoniously capture rich dependence structures in categorical sequences, with memory efficiency suitable for real-time processing of data streams. Parsimonious Bayesian context trees are introduced as a form of variable-order Markov model with conjugate prior distributions. The novel framework requires fewer parameters than fixed-order Markov models by dropping redundant dependencies and clustering sequential contexts. Approximate inference on the context tree structure is performed via a computationally efficient model-based agglomerative clustering procedure. The proposed framework is tested on synthetic and real-world data examples, and it outperforms existing sequence models when fitted to real protein sequences and honeypot computer terminal sessions.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ghani et al. (2026) studied this question.

synapsesocial.com/papers/69b79e638166e15b153ab9c9https://doi.org/10.1007/s11222-026-10835-7
Ask AI
Helpful
Bookmark
Share
View Full Paper