Clotho
Emerging7papers using it
2023first seen
The 'Clotho' dataset is a benchmark that contains audio clips paired with descriptive text, used to evaluate the performance of multi-modal audio-language models in understanding and relating audio content to natural language.
Papers using Clotho (7)
- FORTE: FOL-guided Optimal Refinement for Text-audio rEtrievalOmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher DistillationAC/DC: LLM-based Audio Comprehension via Dialogue ContinuationONE-PEACE: Exploring One General Representation Model Toward Unlimited
ModalitiesZero-Shot Audio Captioning Using Soft and Hard PromptsMINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley
Audio Content Planning and GenerationEnhancing Audio-Language Models through Self-Supervised Post-Training
with Text-Audio Pairs