Harmonic-aware tri-path convolution recurrent network for singing voice separation

Yih-Liang Shen; Ya-Ching Lai; Tai-Shih Chi

Overview

Temporal coherence and spectral regularity are critical cues for human auditory streaming processes and are considered in many sound separation models. Some examples include the Conv-tasnet model, which focuses on temporal coherence using short length kernels to analyze sound, and the dual-path convolution recurrent network (DPCRN) model, which uses two recurring neural networks to analyze general patterns along the temporal and spectral dimensions on a spectrogram. By expanding DPCRN, a harmonic-aware tri-path convolution recurrent network model via the addition of an inter-band RNN is proposed. Evaluation results on public datasets show that this addition can further boost the separation performances of DPCRN.

Schematic diagram of the HA-TPCRN singing voice separation architecture
Schematic diagram

Listening Examples

Demo 1

DSD100 excerpt 43
MixtureOriginal mixed audio
Model Separated vocal Separated music
HA-TPCRN (Proposed)
DPCRN
DPRNN
SpleeterDeepUNet

Demo 2

DSD100 excerpt 38
MixtureOriginal mixed audio
Model Separated vocal Separated music
HA-TPCRN (Proposed)
DPCRN
DPRNN
SpleeterDeepUNet

Demo 3

DSD100 excerpt 14
MixtureOriginal mixed audio
Model Separated vocal Separated music
HA-TPCRN (Proposed)
DPCRN
DPRNN
SpleeterDeepUNet