Randomized trial evaluates music models on timbre, technique, key, and mode in various music genres, suggesting performance disparities.
Music foundation models are increasingly used as feature extractors in music information retrieval; yet, their treatment of Chinese tonal organization has received little controlled evaluation. We probe frozen MERT, MuQ, CultureMERT, and an 88-dimensional librosa baseline on instrument timbre, guzheng playing technique, Western key, and Chinese pentatonic mode. Probes are evaluated with task-specific splits that guard against leakage and with chance-normalized accuracy. Timbre is near ceiling (0.994–0.998), whereas source-aware guzheng-technique accuracy ranges from 0.382 to 0.787. In the unmatched external comparison, GiantSteps key scores 0.441–0.604 and CNPM mode pattern scores 0.294–0.355; this difference is descriptive because the corpora differ in genre, production, provenance, and label distribution. Same-audio CNPM controls identify an important boundary: the TongGong system reaches 0.649–0.734, and absolute tonic reaches 0.433–0.480, whereas their matched five-way mode-pattern labels reach 0.256–0.350 and 0.243–0.366, respectively. Chroma ablations show that pitch-class energy is central to the external key task but does not by itself account for CNPM mode-pattern performance. These results map which musical dimensions are readily decodable from the evaluated frozen representations while avoiding a cultural-capability interpretation of the unmatched cross-corpus comparison.
No takes yet. Share an insight, caveat, or question.
Bai et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: