Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
January 18, 2026PLoS ONEOpen Access

MViT: A vision transformer with fractal path reordering and dynamic positional encoding

View Full Paper
Ask AI
Bookmark
Share

Authors

BLBomin LiuLHLinjun HeYZYan Zhu

Discussion

Loading...

Member takes

Overview

MViT improves classification in computer vision, highlighting better adaptability to complex structures.

Key Points

  • The aim is to improve Vision Transformers' ability to represent complex images by addressing limitations in spatial continuity and positional encoding.
  • Implemented a multi-order fractal mapping to optimize patch reordering.
  • Designed a dynamic partitioning template with a boundary compensation algorithm.
  • Integrated a period-aware positional encoding module with convolutional features.
  • Achieved a 0.52% increase in classification accuracy on CIFAR-100 compared to ViT-B/16.
  • Attained a 0.31% increase in accuracy on ImageNet-21k.
  • Showed improved metrics in PSNR and SSIM, demonstrating robust performance against rotation and variations.

Cite This Study

Liu et al. (2026) studied this question.

synapsesocial.com/papers/696c77f1eb60fb80d13961f5https://doi.org/10.1371/journal.pone.0340788
View Full Paper
Ask AI
Bookmark
Share