ADE20K
Emerging8papers using it
2022first seen
The ADE-20K dataset is a comprehensive benchmark that contains a diverse set of images annotated with pixel-level semantic segmentation, used to evaluate the performance of models in understanding and interpreting complex scenes.
Papers using ADE20K (8)
- A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIPSpatialBoost: Enhancing Visual Representation through Language-Guided ReasoningLangHOPS: Language Grounded Hierarchical Open-Vocabulary Part SegmentationScene-aware Urban Design: A Human-ai Recommendation Framework Using Co-occurrence Embeddings And Vision-language ModelsImage as a Foreign Language: BEiT Pretraining for All Vision and
Vision-Language TasksONE-PEACE: Exploring One General Representation Model Toward Unlimited
ModalitiesMTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text AlignmentEmergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models