Speaker diarization is the task of partitioning an audio stream into segments with the objective of determining “who spoke when.” Traditionally, speaker diarization systems have relied on datasets that typically feature a limited number of speakers, ranging from 2–6 per session often consist of clean or simulated speech. These datasets do not reflect the speaker variability, overlapping speech, or channel diversity observed in real-world operations. In our study, we use the Fearless Steps Apollo-11 (FS-A11) corpus, which presents a wide range of speakers, diverse acoustic environments, varying speaker utterance durations, and an unknown total number of speakers. To address these challenges, we leverage LLMs to generate initial speaker segmentation labels. LLMs are capable of analyzing conversational structure, initiation patterns, and lexical cues specific to Apollo communications. These labels are refined using variable word-based segmentation. Finally, we propose a semi-supervised graph transformer architecture for speaker diarization to enhance speaker embeddings by modeling relationships between segments, taking into account both temporal and contextual information. Final clustering is performed using agglomerative hierarchical clustering (AHC). The training set includes 80 h of labeled audio from five channels, with 282 unique speakers. Evaluation and test sets each consist of 10 h.
Shekar et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: