This paper presents a deep learning–based investigation of automatic music transcription (AMT) for the Thai xylophone (ranat ek), a central instrument in Thai classical music that exhibits distinctive acoustic and performance characteristics. Unlike Western instruments, the Thai xylophone employs a non–equal-tempered tuning system and is commonly performed with both hard and soft mallets, producing markedly different timbral properties that pose significant challenges for transcription. Owing to the lack of publicly available datasets for this instrument, a dedicated dataset was constructed, comprising 63 solo performances (33 with hard mallets and 30 with soft mallets) with a total duration of 113.37 min, each precisely aligned with corresponding MIDI annotations. The proposed AMT framework evaluates the effectiveness of three widely used feature extraction techniques—Mel-spectrogram, Mel-frequency cepstral coefficients (MFCC), and Constant-Q Transform (CQT)—in combination with convolutional neural networks (CNNs). In addition, the CNN-based approach is systematically compared with Long Short-Term Memory (LSTM) and Deep Neural Network (DNN) architectures. Experimental results demonstrate that the CNN model combined with Mel-spectrogram features yields the best performance, achieving F1-scores of 90.97% for frame-level detection and 95.03% for onset detection, and consistently outperforming LSTM and DNN models. In addition to objective evaluation metrics, this study assesses the perceptual validity of AMT outcomes through a large-scale listening experiment involving 400 participants. Participants compared transcriptions generated by different models and assessed their similarity to original performances. Statistical analyses, including chi-square tests and Eta coefficients, reveal strong agreement between human judgments and quantitative evaluation outcomes, confirming the superiority of the CNN-based model. Moreover, musical ability and perceived difficulty of the listening task are found to be significantly correlated with participants’ agreement with model evaluations.
Huaysrijan et al. (Mon,) studied this question.