← all papers · overview

Qtok: A Comprehensive Framework For Evaluating Multilingual Tokenizer Quality In Large Language Models

Abstract

In the development of Large Language Models (LLMs), considerable attention has been given to the quality of training datasets. However, the role of tokenizers in the LLM training pipeline, particularly for multilingual models, has received less focus. The quality of tokenization can significantly impact a model's ability to handle diverse languages effectively. We introduce Qtok, a tool designed t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).