← all papers · overview

Tagengo: A Multilingual Chat Dataset

Abstract

Open source large language models (LLMs) have shown great improvements in recent times. However, many of these models are focused solely on popular spoken languages. We present a high quality dataset of more than 70k prompt-response pairs in 74 languages which consist of human generated prompts and synthetic responses. We use this dataset to train a state-of-the-art open source English LLM to chat

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).