PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 31, 20260 citationsOpen Access

When Does BiMamba Beat Transformers in JEPA-style Masked Latent Prediction? Evidence from Image and Video Benchmarks

View Full Paper
BKBrian Kim

Key Points

  • This research aims to compare the performance of BiMamba and Transformer architectures in masked latent prediction tasks within the JEPA framework.
  • Compared 7 architectures: Transformer, Vanilla Mamba, BiMamba, and 4 Sequential Attention variants.
  • Evaluated across 5 datasets from simple images to complex videos.
  • Measured performance using mean squared error (MSE) for different tasks.
  • BiMamba achieves half the MSE of Transformer for fine-grained temporal tasks (0.55× ± 0.02).
  • Transformers show competitive performance on coarse temporal structure tasks (ImageNet: BiMamba/TF = 1.10×; UCF-101: 1.31× ± 0.10).
  • Sequential Attention methods were found to structurally fail for Mamba.

Abstract

Wepresent, to our knowledge, the first empirical comparison of Transformer attention and Mamba (Structured State Space Model) in Joint-Embedding Predictive Architecture (JEPA). While Mamba has shown competitive results in classification and generation tasks, its applicability to JEPA’s masked latent prediction objective remains unexplored. Wecompare 7 architectures—Transformer, Vanilla Mamba, Bidirectional Mamba (BiMamba), and 4 Sequential Attention variants—across 5 datasets ranging from simple images (Moving MNIST) to complex videos (HMDB-51). Our key finding is that fine grained temporal ambiguity in the task correlates with architecture suitability: on tasks with coarse temporal structure, Transformer remains competitive or better (ImageNet: BiMamba/TF = 1.10×; UCF-101: 1.31× ± 0.10), while on tasks requiring fine-grained temporal discrimination (HMDB-51), BiMamba consistently achieves roughly half the MSE of Transformer (0.55× ± 0.02, reproducible across 3 seeds).Wealso demonstrate why Sequential Attention approaches structurally fail for Mamba and confirm that modality-specific FFN separation remains beneficial even when allmodalities share the same loss function. This is a toy-scale empirical study. We study architectural trends rather than claim state of-the-art capability. Reported ratios should be interpreted as directional evidence, not production-ready benchmarks. All code, checkpoints, and results are publicly available.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Brian Kim (2026) studied this question.

synapsesocial.com/papers/69cb64f0e6a8c024954b90bdhttps://doi.org/10.5281/zenodo.19323215
Ask AI
Helpful
Bookmark
Share
View Full Paper