PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 12, 2026Sensors0 citationsOpen Access

TrackRefine: A Plug-and-Play Decoupled Enhancement Framework for Online Multi-Object Tracking and Segmentation

View Full Paper
LQLongfei QieCCChunlei ChaiRWRuixue Wang

Key Points

  • The aim is to enhance multi-object tracking and segmentation performance in video sequences by addressing existing challenges.
  • Proposed TrackRefine framework for online multi-object tracking and segmentation.
  • Introduced Fast GrabCut-based mask refinement for better mask boundaries.
  • Utilized a multimodal long-short-term memory bank for identity modeling.
  • Achieved 69.4 sMOTSA and 82.7 MOTSA on MOTS20.
  • Attained 62.4/73.7 sMOTSA and 78.0/85.4 MOTSA on KITTI for pedestrians and cars, respectively.
  • Confirmed the effectiveness of modules through ablation studies.

Abstract

Multi-object tracking and segmentation (MOTS) aims to jointly perform pixel-level instance segmentation and temporal identity association for multiple objects in video sequences. Existing online decoupled MOTS methods face several challenges in complex scenarios, including limited front-end mask quality, corruption of memory representations under prolonged occlusion, and unstable data association and trajectory recovery. To address these limitations, we propose TrackRefine, a plug-and-play decoupled enhancement framework. TrackRefine enhances overall performance through back-end refinement without modifying the architecture of the front-end instance segmenter or relying on additional end-to-end joint training. Specifically, we introduce a lightweight Fast GrabCut-based mask refinement module to optimize mask boundaries, a multimodal long-short-term memory bank that integrates appearance, semantic, and shape cues for identity modeling, and a progressive three-stage association strategy for stable matching and long-term trajectory recovery. Experimental results on MOTS20 show that TrackRefine achieves 69.4 sMOTSA, 82.7 MOTSA, and 478 Frag. Experimental results on KITTI MOTS show that it achieves 62.4/73.7 sMOTSA and 78.0/85.4 MOTSA for pedestrians and cars, respectively. Extensive experiments with different front-end instance segmenters verify its plug-and-play flexibility and decoupled design, while ablation studies confirm the effectiveness of each core module. These results show that TrackRefine provides an efficient and practical solution for online MOTS in complex scenarios.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Qie et al. (2026) studied this question.

synapsesocial.com/papers/6a2ba2448101cf8926f013f2https://doi.org/10.3390/s26123696
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Seg2Track-SAM2: SAM2-based Multi-object Tracking and Segmentation2025
  2. 2DETrack: Multi-Object Tracking Algorithm Based on Feature Decomposition and Feature Enhancement2024 · 1 citations
  3. 3AuxTrack: Auxiliary Detection and Spatio-Temporal Attention Matching for Robust Multi-Object Tracking2026 · 1 citations
  4. 4Robust online multi-object tracking with conditional diffusion motion hypotheses and time-aware contrastive prototypes2026
  5. 5Sampling-Resilient Multi-Object Tracking2024 · 5 citations