Qwen-2.5-VL-7B
Emerging6papers using it
2025first seen
The 'Qwen-2.5-VL-7B' is a benchmark used to evaluate the performance of Vision-Language Models (VLMs) in terms of their alignment and ability to reduce hallucinations during inference.
Papers using Qwen-2.5-VL-7B (6)
- Stepwise Token Selection for Efficient Multimodal Large Language ModelsGRIP: Feedback-Guided Prompt Retrieval for Large Multimodal ModelsMuCRASP: Multimodal Chain-of-thought Reasoning aware Structured PruningSelf-Correction Inside the Model: Leveraging Layer Attention to Mitigate Hallucinations in Large Vision Language ModelsHALP: Detecting Hallucinations in Vision-Language Models without Generating a Single TokenToken-Level Inference-Time Alignment for Vision-Language Models