Abstract:Lyric-to-melody generation is a highly challenging task in the field of AI music generation. Due to the difficulty of learning strict yet weak correlations between lyrics and melodies, previous methods have suffered from weak controllability, low-quality and poorly structured generation. To address these challenges, we propose CSL-L2M, a controllable song-level lyric-to-melody generation method based on an in-attention Transformer decoder with fine-grained lyric and musical controls, which is able to generate full-song melodies matched with the given lyrics and user-specified musical attributes. Specifically, we first introduce REMI-Aligned, a novel music representation that incorporates strict syllable- and sentence-level alignments between lyrics and melodies, facilitating precise alignment modeling. Subsequently, sentence-level semantic lyric embeddings independently extracted from a sentence-wise Transformer encoder are combined with word-level part-of-speech embeddings and syllable-level tone embeddings as fine-grained controls to enhance the controllability of lyrics over melody generation. Then we introduce human-labeled musical tags, sentence-level statistical musical attributes, and learned musical features extracted from a pre-trained VQ-VAE as coarse-grained, fine-grained and high-fidelity controls, respectively, to the generation process, thereby enabling user control over melody generation. Finally, an in-attention Transformer decoder technique is leveraged to exert fine-grained control over the full-song melody generation with the aforementioned lyric and musical conditions. Experimental results demonstrate that our proposed CSL-L2M outperforms the state-of-the-art models, generating melodies with higher quality, better controllability and enhanced structure. Demos and source code are available at <a class="link-external link-https" href="https://lichaiustc.github.io/CSL-L2M/" rel="external noopener nofollow">this https URL</a>.

SongGLM: Lyric-to-Melody Generation with 2D Alignment Encoding and Multi-Task Pre-Training

MelodyGLM: Multi-task Pre-training for Symbolic Melody Generation

Unsupervised Melody-to-Lyric Generation

Unsupervised Melody-Guided Lyrics Generation

LOAF-M2L: Joint Learning of Wording and Formatting for Singable Melody-to-Lyric Generation

SongMASS: Automatic Song Writing with Pre-training and Alignment Constraint

Conditional LSTM-GAN for Melody Generation from Lyrics

SongComposer: A Large Language Model for Lyric and Melody Composition in Song Generation

Telemelody: Lyric-to-melody generation with a template-based two-stage method

Interpretable Melody Generation from Lyrics with Discrete-Valued Adversarial Training

Deep Attention-Based Alignment Network for Melody Generation from Incomplete Lyrics

Lyrics-Conditioned Neural Melody Generation

ReLyMe: Improving Lyric-to-Melody Generation by Incorporating Lyric-Melody Relationships

CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical Controls

Melody Generation from Lyrics with Local Interpretability

Syllable-level lyrics generation from melody exploiting character-level language model

Translate the Beauty in Songs: Jointly Learning to Align Melody and Translate Lyrics

Neural Melody Composition from Lyrics

Agent-Driven Large Language Models for Mandarin Lyric Generation

Melody-Guided Music Generation