← all papers · overview

Speak While You Think: Streaming Speech Synthesis During Text Generation

Abstract

Large Language Models (LLMs) demonstrate impressive capabilities, yet interaction with these models is mostly facilitated through text. Using Text-To-Speech to synthesize LLM outputs typically results in notable latency, which is impractical for fluent voice conversations. We propose LLM2Speech, an architecture to synthesize speech while text is being generated by an LLM which yields significant l

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).