CLIP
Emerging15papers using it
2023first seen
CLIP (Contrastive Language-Image Pretraining) is a dataset and model that contains paired images and text, used to evaluate multimodal alignment and preferences in vision-language tasks.
Papers using CLIP (15)
- Clip-handid: Vision-language Model For Hand-based Person IdentificationExperimental Evaluation Of Static Image Sub-region-based Search Models Using CLIPCan Argus Judge Them All? Comparing VLMs Across DomainsCAPT: Confusion-Aware Prompt Tuning for Reducing Vision-Language MisalignmentDeepSight: Bridging Depth Maps and Language with a Depth-Driven Multimodal ModelRAZOR: Ratio-Aware Layer Editing for Targeted Unlearning in Vision Transformers and Diffusion ModelsAligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal AlignmentImportance Sampling for Multi-Negative Multimodal Direct Preference OptimizationDisentangling 3D From Large Vision-language Models For Controlled Portrait GenerationDRIP: Dynamic Patch Reduction Via Interpretable PoolingCompositional Semantics for Open Vocabulary Spatio-semantic RepresentationsMultimodality Helps Unimodality: Cross-Modal Few-Shot Learning with
Multimodal ModelsExtending Multi-modal Contrastive RepresentationsLinear Spaces of Meanings: Compositional Structures in Vision-Language
ModelsFM-OV3D: Foundation Model-based Cross-modal Knowledge Blending for
Open-Vocabulary 3D Detection