Fast Music Source Separation

Rescuing my Master's thesis website — isolating bass, drums, vocals and other instruments from a stereo mix with convolutional neural networks

Introduction

This post revives the website I built for my Master’s thesis, which unfortunately stopped working over the years. The original site lived in the docs/ folder of the Fast-Music-Source-Separation repository, so everything you see here — the samples, the figures, and the links — comes straight from it.

A music source separation system capable of isolating bass, drums, vocals and other instruments from a stereophonic audio mix is presented. The system was developed for the fulfilment of my degree thesis “Separación de fuentes musicales mediante redes neuronales convolucionales”. The thesis report and Python scripts to perform separations and train the neural network, together with DemixME, a standalone Windows x64 application I created which separates instruments in realtime and allows remixing songs on the fly, are provided in the following links:

And here is a demo video of DemixME remixing a song on the fly:

Listen

The samples below come from the DSD100 dataset test set. For each song you can listen to the original mixture and the four estimated sources.

Patrick Talbot - A reason to leave

Mixture
Bass
Drums
Others
Vocals

Secretariat - Over the Top

Mixture
Bass
Drums
Others
Vocals

Hollow Ground - Ill Fate

Mixture
Bass
Drums
Others
Vocals

Bobby Nobody - Stitch Up

Mixture
Bass
Drums
Others
Vocals

System overview

In the first stage, the audio mixture is turned into an image-like representation. This can be achieved applying a Short Time Fourier Transform and calculating its magnitude. The result is a spectrogram, where the vertical axis corresponds to frequency and the horizontal axis to time.

The second stage consists in turning the spectrogram of the audio mixture into the spectrogram of the desired source. This is achieved using a convolutional neural network. The DSD100 dataset is used to train it. It provides mixtures and sources (bass, drums, vocals and others) from 100 different songs (50 for training and 50 for testing).

The proposed neural network is inspired in DeepConvSep (paper). Some of the differences with DeepConvSep are:

The convolutional neural network takes as input a stereophonic spectrogram and reduces its dimensionality using two parallel encoders. The latent space consists of a fully connected layer and its output is a compact representation of the mixture. Multiple decoders take as input this compact representation. The first layer of the decoders is fully connected and its purpose is to transform the representation of the mixture into a representation of the source to isolate. The next layers, which are convolutional, reconstruct spectrograms from the transformed latent space.

Instead of using a neural network for each source to separate, a single neural network is used, which outputs the estimation of the four sources (bass, drums, vocals and others). This way, processing time is reduced and information of all sources is shared among decoders.

Once the spectrograms are estimated, in the final stage a time-varying filter is performed through soft masks. This way, the processing reduces to filtering the input audio mixture, instead of generating from scratch new audios with the neural network. The resulting spectrograms are taken back to time domain applying the Griffin-Lim algorithm.

Citation

If this was useful for your research, please reference it as:

@thesis{pepino2019,
  title  = {Separación de fuentes musicales mediante redes neuronales convolucionales},
  author = {Pepino, Leonardo},
  school = {Universidad Nacional de Tres de Febrero},
  year   = {2019}
}