This framework enhances attribute disentanglement and reduces overconfidence in models for zero-shot learning, emphasizing multimodal embedding benefits.
Key Points
State-of-the-art performance was achieved on three challenging datasets through a new framework.
The method utilizes MLLM embeddings, which provide superior representation for unseen attributes and objects.
Disentanglement challenges are addressed using feature adaptive aggregation and learnable condition masks, enhancing model robustness.
Attribute smoothing is employed to mitigate overconfidence in seen compositions, improving generalization capabilities.