Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
tokenization
loadingβ¦
π€
Ask AI
Awesome tokenization β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
tokenization
15 papers tagged tokenization β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
15 papers Β· trending (default)
numbers = π₯ heat
What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?
(2026)
Gagan Bhatia et al.
1.94
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
(2026)
Xu Ouyang et al.
1.94
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
(2026)
Xiang An et al.
1.94
(1D) Ordered Tokens Enable Efficient Test-Time Search
(2026)
Zhitong Gao et al.
1.83
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
(2026)
Aditya Kumar Singh et al.
1.72
GeoMotionGPT: Geometry-Aligned Motion Understanding with Large Language Models
(2026)
Zhankai Ye et al.
1.67
SmolVLM: Redefining small and efficient multimodal models
(2025)
AndrΓ©s Marafioti et al.
1.28
Imperceptible Jailbreaking against Large Language Models
(2025)
Kuofeng Gao et al.
1.28
Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
(2025)
Ziyuan Huang et al.
1.28
PASTA: Pretrained Action-State Transformer Agents
(2023)
Raphael Boige et al.
β
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
(2023)
Dongchao Yang et al.
β
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
(2023)
Bin Lin et al.
β
BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
(2024)
Mateusz Εajszczak et al.
β
OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents
(2024)
Zihao Wang et al.
β
Emu3: Next-Token Prediction is All You Need
(2024)
Xinlong Wang et al.
β