Replication case series reveals intra-speaker vocal variation degrades speaker embeddings across three neural encoders, highlighting fundamental pitch measurement instability.
Status: frozen. Data collection is closed. The analyses in Section 6 were pre-registered before execution; the pre-registration is included. A prior-art audit was completed after analysis and before release; it is included in full, including the claim it withdrew. What this is. That controlled within-speaker vocal variation degrades automatic speaker recognition is established (Hughes et al., Interspeech 2023; Prieto et al., DSP 2022; Gonzalez Hautamaki et al., JASA 2019). This study does not claim that effect. It reproduces it on three modern encoders with untrained lay speakers, and extends it with per-utterance voice-quality measurement that those studies did not perform. Four speakers produced identical sentences under controlled changes in pitch, phonation and articulation, at fixed microphone position and gain within each speaker. All 30 speaker x condition x encoder cells are negative, 28 unanimous across utterances (ECAPA-TDNN, ResNet, WavLM-base-plus-sv). Articulation timing (staccato versus legato) does not displace the embedding in any encoder, giving a within-corpus null an order of magnitude smaller than the effects. Extension. With speaker and corpus fixed effects, pitch deviation and jitter are independently associated with displacement in all three encoders; harmonics-to-noise ratio and cepstral peak prominence add nothing once jitter is included. HNR correlates 0.55 with jitter and loses all independent power beside it. Principal methodological result: differential measurement error in F0. Rough phonation breaks the pitch tracker, injecting values up to 32.1 semitones (a 6.4x frequency ratio). All 10 impossible observations fall in rough phonation and none in modal phonation (Fisher exact p = 2.4e-11); they carry 2.7x the leverage of modal observations and up to 13x the Cook's distance. Because octave errors are one-directional, the contamination loads asymmetrically onto the upward-pitch coefficient. Two defensible corrections then give opposite answers about the direction of the pitch effect (p = 0.947 contaminated; p = 0.099 dropping only impossible values, favouring upward; p = 0.013 restricting to the reliably tracked corpus, favouring downward). The pre-registered direction prediction is therefore withdrawn rather than reported: the restriction that yields p = 0.013 also removes an entire corpus and so cannot be attributed to F0 cleaning alone. The instability is reported as the result. Also reported as limits. One speaker shows large unanimous displacement not explained by any of the six acoustic variables measured. Encoder differences are descriptive only: one model per family, differing simultaneously in architecture, training data, objective and head. Four speakers, single session, corpus confounded with manipulation type. This is a case series. Contents. Manuscript, pre-registration, full prior-art audit, per-utterance measurement tables for all three encoders (233 rows each), a provenance manifest mapping every row to its source recording and time offsets, all analysis code, and the 137 source recordings. Every number in the manuscript is derivable from the CSVs without audio. Audio licensing - read before downloading. The recordings are included under restricted terms (LICENSE_AUDIO.md): research, benchmarking, evaluation and teaching are permitted; use as training, fine-tuning or distillation data for any machine-learning model, and use for voice cloning or synthetic-voice generation, are prohibited. The speakers are four identifiable adults who consented to public release on these terms; three are under separate commercial voice contract to the author. The manuscript, tables, manifest and code are CC BY 4.0; the record-level licence field is set to the more restrictive of the two because a single field cannot express both. No priority claim is made for any finding in this record.
No takes yet. Share an insight, caveat, or question.
Panagiotis Gkilis (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: