PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 11, 2026IEEE Transactions on Image Processing0 citations

Language Supervised Multi-Camera Multi-Object Tracking

View Full Paper
KMKaige MaoXHXiaopeng HongXFXiaopeng Fan

Key Points

  • This research aims to improve multi-camera multi-object tracking by utilizing language descriptions instead of complex identity annotations.
  • Developed a novel approach called LaVST for language-to-vision weakly-supervised learning.
  • Implemented tracklet-level cross-modality matching to generate reliable pseudo-labels.
  • Designed an ID-aware projection self-correction mechanism for more accurate projections.
  • Models trained with LaVST show competitive performance compared to state-of-the-art identity-supervised methods, demonstrating a 20.0% average gain in IDF1 during cross-dataset evaluations.
  • The language annotations improve tracking accuracy and simplify the labeling process.

Abstract

Recent multi-camera multi-object tracking (MC-MOT) algorithms are primarily trained using per-detection identity annotations, which are complicated to obtain. In contrast, labeling a language description per-object is a more natural and human-friendly way. In this paper, we explore MCMOT in a language-supervised manner (LS-MCMOT) and propose a novel approach LaVST, which performs language-to-vision weakly-supervised learning based on reliable pseudo-labels generated via tracklet-level cross-modality matching. In addition, we design an ID-aware projection self-correction mechanism to correct inaccurate image-to-ground projection in a self-supervised manner. The models trained with our approach exhibit promising performance in LS-MCMOT. Surprisingly, they perform favorably against state-of-the-art identity-supervised methods, especially in cross-dataset evaluation (with an average gain by 20.0% in IDF1), underscoring the potential of language annotations in MCMOT. Codes and language annotations will be available here.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mao et al. (2026) studied this question.

synapsesocial.com/papers/6a2a503380c8f91e7f39cbd9https://doi.org/10.1109/tip.2026.3699088
Ask AI
Helpful
Bookmark
Share
View Full Paper