← all papers · overview

Tokenization Counts: The Impact Of Tokenization On Arithmetic In Frontier Llms

Abstract

Tokenization, the division of input text into input tokens, is an often overlooked aspect of the large language model (LLM) pipeline and could be the source of useful or harmful inductive biases. Historically, LLMs have relied on byte pair encoding, without care to specific input domains. With the increased use of LLMs for reasoning, various number-specific tokenization schemes have been adopted,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).