← all papers · overview

A Family Of Llms Liberated From Static Vocabularies

Abstract

Tokenization is a central component of natural language processing in current large language models (LLMs), enabling models to convert raw text into processable units. Although learned tokenizers are widely adopted, they exhibit notable limitations, including their large, fixed vocabulary sizes and poor adaptability to new domains or languages. We present a family of models with up to 70 billion p

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).