LLaVA-1.5-7B
Emerging6papers using it
2025first seen
'LLaVA-1.5-7B' is a benchmark used to evaluate the performance of Vision-Language Models (VLMs) in terms of their general understanding and reduction of hallucinations across various tasks.
Papers using LLaVA-1.5-7B (6)
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and InferenceSelf-Correction Inside the Model: Leveraging Layer Attention to Mitigate Hallucinations in Large Vision Language ModelsCost-Efficient Multimodal LLM Inference via Cross-Tier GPU HeterogeneityToken-Level Inference-Time Alignment for Vision-Language ModelsVISOR: Visual Input-based Steering for Output Redirection in Vision-Language ModelsMitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models