← all papers · overview

Tuning Large Language Model For End-to-end Speech Translation

Abstract

With the emergence of large language models (LLMs), multimodal models based on LLMs have demonstrated significant potential. Models such as LLaSM, X-LLM, and SpeechGPT exhibit an impressive ability to comprehend and generate human instructions. However, their performance often falters when faced with complex tasks like end-to-end speech translation (E2E-ST), a cross-language and cross-modal transl

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).