← all papers · overview

Seed-bench-2: Benchmarking Multimodal Large Language Models

Abstract

Multimodal large language models (MLLMs), building upon the foundation of powerful large language models (LLMs), have recently demonstrated exceptional capabilities in generating not only texts but also images given interleaved multimodal inputs (acting like a combination of GPT-4V and DALL-E 3). However, existing MLLM benchmarks remain limited to assessing only models' comprehension ability of si

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).