Algorithmic study reveals improved style similarity and rhythmic smoothness in synthetic Chinese folk melodies, indicating viable deep-learning frameworks for cultural music preservation.
This study addresses challenges in Chinese folk music generation, including insufficient style feature extraction, weak melodic coherence, and limited automation. An automatic melody generation model for Chinese folk music style is proposed. First, a comprehensive database is constructed using pitch, rhythm, and timbre features from multi-ethnic musical samples, including Han, Tibetan, Yi, Mongolian, Korean, and Dai music. Mel-frequency cepstral coefficients are used to extract timbre features, and long short-term memory networks are used to model temporal dependence in melodies. An attention mechanism is then introduced to enhance the capture of stylistic features, while a variational autoencoder controls the latent melodic style space. A custom loss function incorporates pitch deviation, rhythmic smoothness, and style consistency to improve generation quality. In 1,000 test samples, the generated melodies show a 32.7% improvement in style similarity and a 24.9% improvement in rhythmic smoothness. The average subjective human rating reaches 4.32 out of 5, approximately 20% higher than the baseline model. The study provides a modeling framework for intelligent generation of culturally distinctive melodies.
No takes yet. Share an insight, caveat, or question.
S. K. Wang (2026) studied this question.