With recent advancements in sensory technologies and computational power, multimodel analysis and learning have gained significant attention in the machine learning (ML) community and have been applied across a wide array of domains. Nevertheless, the effective integration of multimodal data sources poses significant challenges to the extraction of discriminative representations, as well as the design of robust fusion strategies capable of capturing complex intermodal relationships. To address these challenges, in this work, a multimodal based representation learning solution, the deep discriminative multiple canonical correlation analysis (DDMCCA), is proposed. To verify the generic naturalness and effectiveness of DDMCCA, we conduct experiments on three databases with different types of input data sources, including face recognition, handwritten digital recognition, and object recognition. Experimental results validate the superiority of the presented solution over state‐of‐the‐art (SOTA).
Liang et al. (Thu,) studied this question.