Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
softmax
loadingβ¦
π€
Ask AI
Awesome softmax β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
softmax
11 papers tagged softmax β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
11 papers Β· trending (default)
numbers = π₯ heat
The Entropic Bound for Transformers: Why Static Rank Fails and Attention-Native Rank Recovers
(2026)
Byeong Hoon Yoon
4.39
A Coulomb Particle Model for Learning Kernel Attention in Transformers
(2026)
Masoud Badiei Khuzani et al.
4.39
MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention
(2026)
Jianlin Yu et al.
4.39
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing
(2026)
Tommaso Cerruti et al.
2.00
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
(2026)
Yingfa Chen et al.
1.94
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
(2026)
Haoyi Zhu et al.
1.94
ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention
(2026)
Joe Sharratt
1.94
Native Hybrid Attention for Efficient Sequence Modeling
(2025)
Jusen Du et al.
1.28
Error-Free Linear Attention is a Free Lunch: Exact Solution from Continuous-Time Dynamics
(2025)
Jingdi Lei et al.
1.28
Cottention: Linear Transformers With Cosine Attention
(2024)
Gabriel Mongaras et al.
β
Differential Transformer
(2024)
Tianzhu Ye et al.
β