Editorial highlights innovations in deep multimodal learning for information generation and retrieval, suggesting new directions.
This editorial introduces the Special Issue on Deep Multimodal Generation and Retrieval , hosted by the ACM Transactions on Multimedia Computing, Communications, and Applications in 2024. Information generation (IG) and information retrieval (IR) are two key representative approaches of information acquisition, i.e. , producing content either via generation or via retrieval. While traditional IG and IR have achieved great success within the scope of languages, the under-utilization of varied data sources in different modalities ( i.e. , text, images, audio, and video) would hinder IG and IR techniques from giving the full advances and thus limit the applications in the real world. Knowing the fact that our world is replete with multimedia information, this special issue encourages the development of deep multimodal learning for the research of IG and IR. Benefiting from a variety of data types and modalities, some of the latest prevailing techniques are extensively invented to show great facilitation in multimodal IG and IR learning. With this special issue, we encourage explorations in Deep Multimodal Generation and Retrieval, providing a platform for researchers to share insights and advancements in this rapidly evolving domain. Each article provides novel insights into areas and challenges such as Multimodal Semantics Understanding, Generative Models for Vision Synthesis, Multimodal Information Retrieval, Explainable and Reliable Multimodal Learning . We summarize the main contributions of the included works and emphasize their role in advancing the field of multimodal generation or via retrieval. Finally, we discuss ongoing challenges and future opportunities in this rapidly evolving domain, particularly in the context of large foundational models.
No takes yet. Share an insight, caveat, or question.
Fei et al. (2025) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: