A recent study~ reported an inverted-U relationship between listener preference on the Free Music Archive (FMA) and of pooled CLAP audio embeddings. That result is clean and large (z = 11. 6 for the negative quadratic at 2-bit quantization on the 23, 395-track joined subset used here), but it leaves open what the ``natural basis'' actually sees, and whether the signal can be reproduced with recognizable music-theoretic variables. We report three findings on the same 23, 395 tracks. , the top ten CLAP principal components load dominantly on timbral (MFCC) and broadband spectral features, with weak harmonic (chroma, tonnetz) loadings. , a compressibility analogue of the main paper's pipeline, run on the 518-feature per-track librosa bank at any retained dimension K \9, 16, 32, 64, 128\ and any bit width \2, 3, 4\, fails to recover the Wundt quadratic (max |z_ | = 2. 6 across 30 cells, versus z_ = 5. 43 uncorrected and z_ = 6. 04 Bonferroni thresholds). Per-track acoustic summary statistics do not carry the signal. , the signal does reappear on summaries of the time-resolved encoder token sequences. Under a pre-registered discovery/confirmation split and a permutation-calibrated null, MERT~ Mahalanobis atypicality---how far a track's token distribution sits from the corpus prior Gaussian---shows a held-out within-corpus Wundt quadratic at z_ = -6. 07 (bootstrap 95\ -7. 84, -4. 38; sign preserved under leave-one-genre-out across all nine FMA genres). CLAP Mahalanobis replicates the direction at weaker magnitude. , a pure- librosa follow-up that extracts time-resolved 43-dimensional feature sequences from raw audio---no neural encoder---and computes Gaussian-distance summaries in exactly the same way recovers a weaker but directionally-consistent effect: symmetric-KL divergence between each track's within-track Gaussian and the corpus prior gives held-out z_ = -2. 52 (bootstrap 95\ -0. 51], CI excludes zero; full-data z = -3. 18). The magnitude hierarchy across four feature classes is monotone in structural-ness: per-track librosa marginals (null, |z| \! \! 2. 6) librosa within-track covariance (|z| = 2. 52 held-out, partially bridging to pure acoustics) neural-encoder structural features (|z| = 6. 07) CLAP-cos in the main paper (|z| = 11. 57 full-data). The aesthetic axis is geometric and partially accessible from classical MIR covariance, but the majority of the signal lives in representations that neural encoders expose and classical per-track summary statistics do not. Author preprint deposited for archival and citation. Draft — pending author review.
Andrew Bond (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: