Key points are not available for this paper at this time.
For speaker recognition, pseudo-labeling has shown advantages in alleviating the scarcity of labeled data. Inspired by image classification tasks, existing methods typically adopt threshold-based strategies to identify reliable pseudo-labels. However, compared to image classification, speaker recognition requires finer-grained class discrimination for open-set identity verification, and thus often adopts margin-based losses that amplify gradients near decision boundaries. While effective under full supervision, this design increases sensitivity to noisy or sparse pseudo-labels, limiting the effectiveness of threshold-based selection. In this work, we propose SpeakerMatch , a novel distribution-based framework for semi-supervised and self-supervised speaker recognition. SpeakerMatch models the confidence distribution to distinguish reliable pseudo-labels from noisy ones globally and selects those whose confidence values and confidence prediction behaviors closely align with high-quality signals. Systematic evaluation across five settings shows that our method outperforms existing approaches, achieving a 13.7% relative improvement over the best semi-supervised speaker recognition baseline, while also delivering a lower equal error rate (EER) and reduced training costs compared to self-supervised methods. • A unified framework for semi-supervised and self-supervised speaker recognition. • Confidence distribution modeling enables global selection of reliable pseudo-labels. • Confidence matching aligns pseudo-labels with high-quality signals. • Consistency matching expands pseudo-labels by mining stable prediction patterns.
Liu et al. (Fri,) studied this question.