← all papers · overview

Tokenization Matters! Degrading Large Language Models Through Challenging Their Tokenization

Abstract

Large Language Models (LLMs) have shown remarkable capabilities in language understanding and generation. Nonetheless, it was also witnessed that LLMs tend to produce inaccurate responses to specific queries. This deficiency can be traced to the tokenization step LLMs must undergo, which is an inevitable limitation inherent to all LLMs. In fact, incorrect tokenization is the critical point that hi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).