MMMU
Canonical34papers using it
2023first seen
A massive multi-discipline multimodal benchmark of college-level questions requiring image understanding plus expert reasoning.
Papers using MMMU (34)
- STEP3-VL-10B Technical ReportQwen3-VL Technical ReportMorphoQuant: Modality-Aware Quantization for Omni-modal Large Language ModelsInternvl3.5: Advancing Open-source Multimodal Models In Versatility, Reasoning, And EfficiencyChatVLA: Unified Multimodal Understanding and Robot Control with
Vision-Language-Action ModelData-Driven Calibration of Prediction Sets in Large Vision-Language Models Based on Inductive Conformal PredictionTVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal UnderstandingVision Verification Enhanced Fusion of VLMs for Efficient Visual ReasoningSame Answer, Different Representations: Hidden instability in VLMsLinMU: Multimodal Understanding Made LinearThinking with Video: Video Generation as a Promising Multimodal Reasoning ParadigmVisPlay: Self-Evolving Vision-Language Models from ImagesSimple Vision-language Math Reasoning Via Rendered TextDiagnosing Visual Reasoning: Challenges, Insights, and a Path ForwardFusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language ModelLLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy ModelSAIL-VL2 Technical ReportTraining Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons LearnedCoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual GroundingMultimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language ModelsReinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-TrainingSkywork-r1v3 Technical ReportASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLMConstructive Distortion: Improving Mllms With Attention-guided Image WarpingHow Far Have Medical Vision-language Models Come? A Comprehensive Benchmarking StudyMmjee-eval: A Bilingual Multimodal Benchmark For Evaluating Scientific Reasoning In Vision-language ModelsTest-time Warmup For Multimodal Large Language ModelsGam-agent: Game-theoretic And Uncertainty-aware Collaboration For Complex Visual ReasoningLaViDa: A Large Diffusion Language Model for Multimodal UnderstandingSkywork R1V: Pioneering Multimodal Reasoning with Chain-of-ThoughtMOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused
Vision-Language ProcessingAre We on the Right Way for Evaluating Large Vision-Language Models?MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkText as Images: Can Multimodal Large Language Models Follow Printed
Instructions in Pixels?