Music tag words that describe music audio by text have different levels of. Taking this issue into account, we propose a music classification that aggregates multi-level and multi-scale features using pre-trained extractors. In particular, the feature extractors are trained in-level deep convolutional neural networks using raw waveforms. We show this approach achieves state-of-the-art results on several music datasets.
No takes yet. Share an insight, caveat, or question.
Lee et al. (2017) studied this question.