Audio Style Transfer Using Shallow Convolutional Networks and Random Filters.

Jiyou Chen,Gaobo Yang,Huihuang Zhao,Manimaran Ramasamy
DOI: https://doi.org/10.1007/s11042-020-08798-6
IF: 2.577
2020-01-01
Multimedia Tools and Applications
Abstract:Recently, with the advent of Convolutional Neural Network (CNN) era, Neural style transfer on images has become a very active research topic and the style of an image can be transferred to another image through a CNN so that the image retains both its own content and another style of image. In this work, we propose an algorithm for audio style transfer that uses the force of CNN to generate a new audio from a style audio. We use Continuous Wavelet Transfer(CWT) to convert the audio into a spectrogram and then use the spectrogram as the representation of the audio image through image style transfer method to obtain a new image, and finally, generate an audio using iterative phase reconstruction with Griffin-Lim. We succeed in transferring audio such as light music but had difficulty in transferring audio that has lyrics and high-level metrics such as emotion or tone. We propose several measures to improve the quality of audio and a lot of experimental results shows that our method is better than other methods in terms of sound quality.
What problem does this paper attempt to address?