← all papers · overview

C3LLM: Conditional Multimodal Content Generation Using Large Language Models

Abstract

We introduce C3LLM (Conditioned-on-Three-Modalities Large Language Models), a novel framework combining three tasks of video-to-audio, audio-to-text, and text-to-audio together. C3LLM adapts the Large Language Model (LLM) structure as a bridge for aligning different modalities, synthesizing the given conditional information, and making multimodal generation in a discrete manner. Our contributions

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).