LVIS
Emerging5papers using it
2026first seen
Progress on object detection is enabled by datasets that focus the research community's attention on open challenges. This process led us from simple images to complex scenes and from bounding boxes to segmentation masks. In this work, we introduce LVIS (pronounced `el-vis'): a new dataset for Large Vocabulary Instance
Papers using LVIS (5)
- VL-DINO: Leveraging CLIP Vision-Language Knowledge for Open-Vocabulary Object DetectioMoondream Segmentation: From Words to MasksQATMA: Quantization-Aware Training with Multimodal Alignment for Open-Vocabulary Object DetectionDetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object DetectionEnhancing Open-Vocabulary Object Detection through Multi-Level Fine-Grained Visual-Language Alignment