This paper deals with the extraction of an instrument from music by using a deep neural network. As prior information, we only assume to know the instrument types that are present in the mixture and, using this information, we generate the training data from a database with solo instrument performances. The neural network is built up from rectified linear units where each hidden layer has the same number of nodes as the output layer. This allows a least squares initialization of the layer weights and speeds up the training of the network considerably compared to a traditional random initialization. We give results for two mixtures, each consisting of three instruments, and evaluate the extraction performance using BSS Eval for a varying number of hidden layers.
No takes yet. Share an insight, caveat, or question.
Uhlich et al. (2015) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: