Rap is a rhythmic vocal style characterized by a distinct flow and rhyming patterns synchronized to the beat. While previous studies have primarily focused on the linguistic features of rap lyrics, limited attention has been paid to vocal characteristics. Such approaches often fail to capture performer-specific vocal traits and expressive techniques. In our prior work, we developed a method to quantify rhyme similarity based on formant trajectories. Building on this, the present study integrates formant-based similarity with RMS amplitude-based similarity to better represent both vowel quality and loudness dynamics. Formant trajectories (F1 and F2) are extracted via linear predictive coding (LPC), while RMS amplitudes are computed through frame-based processing. Similarity scores for each feature are independently calculated using dynamic time warping (DTW) and subsequently standardized and integrated through Z-score normalization to produce a unified similarity metric. We applied this method to 30 Japanese rap songs, extracting rhymed segments and computing similarity scores. Experimental results demonstrate that the integrated metric outperforms individual features in rhyme detection, yielding improvements in both area under the ROC curve (AUC) and average precision (AP). Furthermore, we propose a novel methodology for representing the structural organization of rap songs based on similarity-based analysis.
Torii et al. (Wed,) studied this question.