Zero-shot learning (ZSL) typically leverages semantic knowledge and textual descriptions of classes to forge connections between seen and unseen classes. ZSL can classify new categories of data unseen in the training set. Prior research has focused on aligning image features with their corresponding auxiliary information, overlooking the limitation whereby individual features may not capture the full spectrum of information inherent in the original image. Additionally, there are concerns regarding the bias in the predicted results towards seen classes during the testing phase. In this work, we introduce a novel approach termed Feature Fusion Transformer Network (FFusion), which decomposes visual and semantic features into a multitude of functionally distinct features, fusing them to form new representations that better encapsulate the information content of the original images during training. Furthermore, we implement novel loss functions to balance the model’s focus on unseen and seen classes within the ZSL framework. Experiments show that our model achieves accuracy of 76.8% in the CUB dataset, while reducing the bias between seen and unseen classes to a mere 0.2%.
No takes yet. Share an insight, caveat, or question.
Tao et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: