Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
March 24, 2024Open Access

ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided Optimization

View Full Paper
Ask AI
Bookmark
Share

Authors

HWHao WangJiangnan UniversityFLFang LiuXidian UniversityLJLicheng JiaoInstitut polytechnique de Grenoble

Discussion

Loading...

Member takes

Implication

Key Points

Key points are not available for this paper at this time.

Cite This Study

Wang et al. (2024) studied this question.

synapsesocial.com/papers/68e72954b6db6435876a2dfahttps://doi.org/10.1609/aaai.v38i6.28347
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision2021 · 1,195 citations
  2. 2Is Space-Time Attention All You Need for Video Understanding?2021 · 1,468 citations
  3. 3An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020 · 21,793 citations
  4. 4Fine-tuned CLIP Models are Efficient Video Learners2022 · 5 citations
  5. 5Expanding Language-Image Pretrained Models for General Video Recognition2022 · 16 citations