Words Are All You Need? Language As An Approximation For Human Similarity Judgments
2022 Β· Raja Marjieh, Pol van Rijn, Ilia Sucholutsky, et al.
Abstract
Human similarity judgments are a powerful supervision signal for machine learning applications based on techniques such as contrastive learning, information retrieval, and model alignment, but classical methods for collecting human similarity judgments are too expensive to be used at scale. Recent methods propose using pre-trained deep neural networks (DNNs) to approximate human similarity, but pre-trained DNNs may not be available for certain domains (e.g., medical images, low-resource languages) and their performance in approximating human similarity has not been extensively tested. We conducted an evaluation of 611 pre-trained models across three domains -- images, audio, video -- and found that there is a large gap in performance between human similarity judgments and pre-trained DNNs. To address this gap, we propose a new class of similarity approximation methods based on language. To collect the language data required by these new methods, we also developed and validated a novel
Authors
(none)
Tags
Stats
Related papers
- Learning From One And Only One Shot (2022)5.24
- Dreamsim: Learning New Dimensions Of Human Visual Similarity Using Synthetic Data (2023)5.84
- Learning Similarity Measures From Data (2020)12.68
- It's The Best Only When It Fits You Most: Finding Related Models For Serving Based On Dynamic Locality Sensitive Hashing (2020)0.00
- Description-based Text Similarity (2023)0.00
- Learning Non-metric Visual Similarity For Image Retrieval (2017)11.58
- Why Do Nearest Neighbor Language Models Work? (2023)3.56
- The Curious Layperson: Fine-grained Image Recognition Without Expert Labels (2021)9.99