COCO-2017
Emerging8papers using it
20HF downloads
0HF likes
2024first seen
COCO 2017 image captions in Vietnamese The dataset is firstly introduced in dinhanhx/VisualRoBERTa. I use VinAI tools to translate COCO 2027 image caption (2017 Train/Val annotations) from English to Vietnamese. Then we merge UIT-ViIC dataset into it. To load the dataset, one can take a look at this code in VisualRoBER
π€ Hugging Faceβ unknown
Papers using COCO-2017 (8)
- Can Multimodal Large Language Models Understand Spatial Relations?Context-Dependent Affordance Computation in Vision-Language ModelsEnhancing Open-Vocabulary Object Detection through Multi-Level Fine-Grained Visual-Language AlignmentCoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language TasksExtreme Model Compression For Edge Vision-language Models: Sparse Temporal Token Fusion And Adaptive Neural CompressionCaprecover: A Cross-modality Feature Inversion Attack Framework On Vision Language ModelsSJTU:Spatial judgments in multimodal models towards unified segmentation through coordinate detectionAn Enhanced Large Language Model For Cross Modal Query Understanding
System Using DL-KeyBERT Based CAZSSCL-MPGPT