Medical image segmentation is challenging due to subtle pathological patterns and the inherent ambiguity of clinical descriptions. Although vision–language models have shown promise, they frequently lack fine-grained perception of structural variability. To address these limitations, we propose the Symmetry- and Variability-Perceiving Conditional Variational Autoencoder (SVP-CVAE). The proposed method integrates a clinical attribute encoder with a morphology-aware enhancement module that incorporates a cross-bilateral symmetry mechanism to explicitly capture symmetry-related variations. By reformulating the segmentation task as a probabilistic prior-to-posterior inference process, SVP-CVAE models the one-to-many mapping between textual attributes and visual realizations. Furthermore, we introduce an attribute-latent contrastive objective to ensure that the latent space encodes discriminative morphological information. Extensive experiments demonstrate that the proposed framework achieves superior segmentation accuracy compared to state-of-the-art methods. Results indicate that SVP-CVAE effectively captures diverse yet anatomically plausible structural variations while maintaining high sensitivity to bilateral symmetry. Comprehensive ablation studies confirm that the performance gains are synergistically driven by the proposed symmetry-perceiving module and the contrastive semantic alignment objective, rather than relying solely on the probabilistic formulation. In conclusion, integrating explicit symmetry perception with probabilistic modeling significantly enhances the reliability and interpretability of multimodal medical image segmentation in complex clinical scenarios.
Jiang et al. (2026) studied this question.