COCO
Canonical111papers using it
28,119HF downloads
82HF likes
2016first seen
Common Objects in Context โ 330k images with object-detection, segmentation, keypoint, and captioning annotations.
Papers using COCO (111)
- Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language RetrievalDite-HRNet: Dynamic Lightweight High-Resolution Network for Human Pose EstimationHSA: Hierarchical Slot Attention for Multi-granularity Scene-DecompositionRethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware PerspectiveFRFDet: Efficient UAV Small Object Detection with Symmetric Sampling and Scalable FusionRepurposing CLIP to Localize at Pixel LevelConfidence Scores in Open-Vocabulary Detection Are a Biased Mixture of Scale and SemanticsSlot-RAE: Streamlining Object-Centric Learning via Direct Representation Auto-EncodersInhibited Self-Attention: Sharpening Focus in Vision TransformersPractical Insights into Semi-Supervised Object Detection ApproachesCompletely Weakly Supervised Class-Incremental Learning for Semantic SegmentationLINEA: Fast and Accurate Line Detection Using Scalable TransformersTeach YOLO to Remember: A Self-Distillation Approach for Continual Object DetectionYOLOv4: A Breakthrough in Real-Time Object DetectionUnveiling the Unknown: Open Vocabulary Object Detection with Scene GraphsTraining-Free Metrics for Synthetic Object Detection Data: A Proxy for Detector PerformanceA Turbo-Inference Strategy for Object Detection and Instance SegmentationWhat Helps---and What Hurts: Bidirectional Explanations for Vision TransformersExploring Open-Vocabulary Object Recognition in Images using CLIPYOLO Object Detectors for Robotics -- a Comparative StudyEnhancing Open-Vocabulary Object Detection through Multi-Level Fine-Grained Visual-Language AlignmentLe-DETR: Revisiting Real-Time Detection Transformer with Efficient Encoder DesignZENITH: Automated Gradient Norm Informed Stochastic OptimizationSSR: Semantic and Spatial Rectification for CLIP-based Weakly Supervised SegmentationDiffusion Is Your Friend in Show, Suggest and TellLAPX: Lightweight Hourglass Network with Global ContextMulti-label Classification with Panoptic Context Aggregation NetworksDSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive ArchitectureUtilizing dynamic sparsity on pretrained DETRA Training-Free Framework for Open-Vocabulary Image Segmentation and Recognition with EfficientNet and CLIPPyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose EstimationLLM-Guided Agentic Object Detection for Open-World UnderstandingTest-time Vocabulary Adaptation for Language-driven Object DetectionunMORE: Unsupervised Multi-Object Segmentation via Center-Boundary ReasoningMultiple Object Stitching for Unsupervised Representation LearningPose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO SubsetsLeveraging Vision-Language Pre-training for Human Activity Recognition in Still ImagesMask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal RetrievalVisual Textualization for Image Prompted Object DetectionDecoupling Classifier for Boosting Few-shot Object Detection and Instance SegmentationThe Missing Point in Vision Transformers for Universal Image SegmentationDIVE: Inverting Conditional Diffusion Models for Discriminative TasksFractional Correspondence Framework in Detection TransformerExploring Token-Level Augmentation in Vision Transformer for
Semi-Supervised Semantic SegmentationToward Lightweight and Fast Decoders for Diffusion Models in Image and
Video GenerationApproximate Size Targets Are Sufficient for Accurate Semantic
SegmentationDynamic Relation Inference via Verb EmbeddingsLooking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient SegmentationLP-DETR: Layer-wise Progressive Relations for Object DetectionSelf-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image GenerationProgressive Token Length Scaling in Transformer Encoders for Efficient
Universal SegmentationPosition Focused Attention Network For Image-text MatchingData Augmentation To Improve Robustness Of Image Captioning SolutionsCornerNet: Detecting Objects as Paired KeypointsDeformable DETR: Deformable Transformers for End-to-End Object DetectionAssociative Embedding: End-to-End Learning for Joint Detection and
GroupingFew-Shot Object Detection with Fully Cross-TransformerCPTR: Full Transformer Network for Image CaptioningImage Captioning: Transforming Objects into WordsConv2Former: A Simple Transformer-Style ConvNet for Visual RecognitionCausal Intervention for Weakly-Supervised Semantic SegmentationOne-Shot Instance SegmentationSelf-EMD: Self-Supervised Object Detection without ImageNetComprehensive Attention Self-Distillation for Weakly-Supervised Object
DetectionFace Detection Using Improved Faster RCNNTFPose: Direct Human Pose Estimation with TransformersISTR: End-to-End Instance Segmentation with TransformersFully Convolutional Instance-aware Semantic SegmentationEfficient Visual Pretraining with Contrastive DetectionUnsupervised Discovery of the Long-Tail in Instance Segmentation Using
Hierarchical Self-SupervisionMViTv2: Improved Multiscale Vision Transformers for Classification and
DetectionCRCNet: Few-shot Segmentation with Cross-Reference and Region-Global
Conditional NetworksDiffusionInst: Diffusion Model for Instance SegmentationFeature-Driven Super-Resolution for Object DetectionImplicit Feature Pyramid Network for Object DetectionMulti-class Token Transformer for Weakly Supervised Semantic
SegmentationEnd-to-End Object Detection with Fully Convolutional NetworkDeep Occlusion-Aware Instance Segmentation with Overlapping BiLayersCAT: Cross-Attention Transformer for One-Shot Object DetectionHierarchical Attention Network for Few-Shot Object Detection via
Meta-Contrastive LearningUnderstanding Gaussian Attention Bias of Vision Transformers Using
Effective Receptive FieldsTowards Few-Annotation Learning for Object Detection: Are
Transformer-based Models More Efficient ?T-VSE: Transformer-Based Visual Semantic EmbeddingSpatial Reasoning for Few-Shot Object DetectionAnalysis of Visual Reasoning on One-Stage Object DetectionHCFormer: Unified Image Segmentation with Hierarchical ClusteringA Simple Latent Diffusion Approach for Panoptic Segmentation and Mask
InpaintingPPT: token-Pruned Pose Transformer for monocular and multi-view human
pose estimationContextual Relabelling of Detected ObjectsDETReg: Unsupervised Pretraining with Region Priors for Object DetectionTask Specific Attention is one more thing you need for object detectionTime-rEversed diffusioN tEnsor Transformer: A new TENET of Few-Shot
Object DetectionCan the Query-based Object Detector Be Designed with Fewer Stages?IvaNet: Learning to jointly detect and segment objets with the help of
Local Top-Down ModulesLearning to Inpaint by Progressively Growing the Mask RegionsImage Captioning using Multiple Transformers for Self-Attention
MechanismModulating Localization and Classification for Harmonized Object
DetectionPoseur: Direct Human Pose Regression with TransformersACORT: A Compact Object Relation Transformer for Parameter Efficient
Image CaptioningSeqCo-DETR: Sequence Consistency Training for Self-Supervised Object
Detection with TransformersCLIP-DIY: CLIP Dense Inference Yields Open-Vocabulary Semantic
Segmentation For-FreeCOMNet: Co-Occurrent Matching for Weakly Supervised Semantic
SegmentationDECO: Unleashing the Potential of ConvNets for Query-based Detection and
SegmentationA Simple and Generalist Approach for Panoptic SegmentationCOCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image
UnderstandingWaterfall Transformer for Multi-person Pose EstimationOPCap:Object-aware Prompting CaptioningEOV-Seg: Efficient Open-Vocabulary Panoptic SegmentationSingle-Shot Panoptic SegmentationPE-former: Pose Estimation TransformerDeep Multi-Task Networks For Occluded Pedestrian Pose Estimation