Musictm-dataset For Joint Representation Learning Among Sheet Music, Lyrics, And Musical Audio
2020 Β· Donghuo Zeng, Yi Yu, Keizo Oyama
Abstract
This work present a music dataset named MusicTM-Dataset, which is utilized in improving the representation learning ability of different types of cross-modal retrieval (CMR). Little large music dataset including three modalities is available for learning representations for CMR. To collect a music dataset, we expand the original musical notation to synthesize audio and generated sheet-music image, and build musical notation based sheet-music image, audio clip and syllable-denotation text as fine-grained alignment, such that the MusicTM-Dataset can be exploited to receive shared representation for multimodal data points. The MusicTM-Dataset presents 3 kinds of modalities, which consists of the image of sheet-music, the text of lyrics and synthesized audio, their representations are extracted by some advanced models. In this paper, we introduce the background of music dataset and express the process of our data collection. Based on our dataset, we achieve some basic methods for CMR tasks
Authors
(none)
Tags
Stats
Related papers
- Unified Cross-modal Translation Of Score Images, Symbolic Music, And Performance Audio (2025)0.00
- Mumu-llama: Multi-modal Music Understanding And Generation Via Large Language Models (2024)6.34
- MERGE -- A Bimodal Audio-lyrics Dataset For Static Music Emotion Recognition (2024)0.00
- Fakemusiccaps: A Dataset For Detection And Attribution Of Synthetic Music Generated Via Text-to-music Models (2024)0.00
- Musilingo: Bridging Music And Text With Pre-trained Language Models For Music Captioning And Query Response (2023)9.03
- Exploiting Synchronized Lyrics And Vocal Features For Music Emotion Detection (2019)0.00
- Deep Cross-modal Correlation Learning For Audio And Lyrics In Music Retrieval (2017)14.06
- Yourmt3+: Multi-instrument Music Transcription With Enhanced Transformer Architectures And Cross-dataset Stem Augmentation (2024)11.84