← all papers · overview

Modaverse: Efficiently Transforming Modalities With Llms

Abstract

Humans possess the capability to comprehend diverse modalities and seamlessly transfer information between them. In this work, we introduce ModaVerse, a Multi-modal Large Language Model (MLLM) capable of comprehending and transforming content across various modalities including images, videos, and audio. Predominant MLLM frameworks have largely relied on the alignment of latent spaces of textual a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).