C3LLM: Conditional Multimodal Content Generation Using Large Language Models
2024 Β· Zixuan Wang, Qinkai Duan, Yu-Wing Tai, et al.
Abstract
We introduce C3LLM (Conditioned-on-Three-Modalities Large Language Models), a novel framework combining three tasks of video-to-audio, audio-to-text, and text-to-audio together. C3LLM adapts the Large Language Model (LLM) structure as a bridge for aligning different modalities, synthesizing the given conditional information, and making multimodal generation in a discrete manner. Our contributions are as follows. First, we adapt a hierarchical structure for audio generation tasks with pre-trained audio codebooks. Specifically, we train the LLM to generate audio semantic tokens from the given conditions, and further use a non-autoregressive transformer to generate different levels of acoustic tokens in layers to better enhance the fidelity of the generated audio. Second, based on the intuition that LLMs were originally designed for discrete tasks with the next-word prediction method, we use the discrete representation for audio generation and compress their semantic meanings into acousti
Authors
(none)
Tags
Stats
Related papers
- Uniaudio 1.5: Large Language Model-driven Audio Codec Is A Few-shot Audio Task Learner (2024)0.00
- Llms Meet Multimodal Generation And Editing: A Survey (2024)5.48
- Audiolm: A Language Modeling Approach To Audio Generation (2022)18.91
- Large Language Models Are Strong Audio-visual Speech Recognition Learners (2024)9.59
- A Review Of Multi-modal Large Language And Vision Models (2024)0.00
- M\(^{2}\)ugen: Multi-modal Music Understanding And Generation With The Power Of Large Language Models (2023)0.00
- Semantically Consistent Video-to-audio Generation Using Multimodal Language Large Model (2024)0.00
- X-LLM: Bootstrapping Advanced Large Language Models By Treating Multi-modalities As Foreign Languages (2023)0.00