GenEval
Emerging21papers using it
2025first seen
GenEval is a benchmark used to evaluate the performance of models in unified multimodal understanding and generation tasks.
Papers using GenEval (21)
- IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image GenerationRethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-trainingCan AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal ModelsMural: Transferring LLM knowledge to image generation via Mixture-of-TransformersPrompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion TransformersDynin-Omni: Omnimodal Unified Large Diffusion Language ModelEvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and GenerationCheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and GenerationHYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized TokenizationEnhancing Alignment for Unified Multimodal Models via Semantically-Grounded SupervisionM3: High-fidelity Text-to-Image Generation via Multi-Modal, Multi-Agent and Multi-Round Visual ReasoningGenAgent: Scaling Text-to-Image Generation via Agentic Multimodal ReasoningMammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and GenerationUniGame: Turning a Unified Multimodal Model Into Its Own AdversaryReconstruction Alignment Improves Unified Multimodal ModelsLavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and GenerationThe Telephone Game: Evaluating Semantic Drift in Unified ModelsOvis-U1 Technical ReportUnigen: Enhanced Training & Test-time Strategies For Unified Multimodal Understanding And GenerationOpenUni: A Simple Baseline for Unified Multimodal Understanding and GenerationHarmonizing Visual Representations for Unified Multimodal Understanding
and Generation