Musical DNA is a two-component computational system connecting music theory, machine learning, and copyright law. Component A is a style-fingerprinting engine trained on 353 MIDI files from six classical composers (Bach, Beethoven, Chopin, Debussy, Mozart, Rachmaninoff), extracting 21 musical features across melodic, harmonic, rhythmic, and structural dimensions. A Random Forest classifier reaches 88.7% test accuracy on six-composer classification (an SVM marginally higher at 91.5%); its most discriminative features are pitch range, key stability, and average pitch. A generalization test produced one unexpected result: an AI-generated Carnatic classical vocal piece was confidently classified as Bach, exposing a genre-bias limitation relevant to evaluating AI-generated music. Component B applies pairwise similarity metrics to a hand-curated dataset of 43 real music copyright cases spanning 1946 to 2023, 33 of them fully scored. A logistic regression model finds that non-musical factors (plaintiff fame, defendant commercial success, expert testimony, litigation forum) predict case outcomes substantially better than musical similarity alone (AUC 0.86 vs. 0.39 under cross-validation), with plaintiff fame the single strongest predictor. Five outlier cases where computed similarity and the court's ruling diverge most are examined as detailed studies; notably, Williams v. Gaye ("Blurred Lines") is among the least stylistically similar pairs despite its infringement verdict. The paper contributes an open-source audio-to-MIDI similarity pipeline and evidence that the legal standard of "substantial similarity" tracks non-musical case factors more closely than any computational similarity measure tested here.
Nihal Gundluru (Sat,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: