Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
linear
loadingβ¦
π€
Ask AI
Awesome linear β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
linear
21 papers tagged linear β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
21 papers Β· trending (default)
numbers = π₯ heat
The Entropic Bound for Transformers: Why Static Rank Fails and Attention-Native Rank Recovers
(2026)
Byeong Hoon Yoon
4.39
A Coulomb Particle Model for Learning Kernel Attention in Transformers
(2026)
Masoud Badiei Khuzani et al.
4.39
Convergent Evolution: How Different Language Models Learn Similar Number Representations
(2026)
Deqing Fu et al.
1.94
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
(2026)
Haoyi Zhu et al.
1.94
Linear representations in language models can change dramatically over a conversation
(2026)
Andrew Kyle Lampinen et al.
1.67
Linear Correlation in LM's Compositional Generalization and Hallucination
(2025)
Letian Peng et al.
1.28
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
(2025)
Hung-Yueh Chiang et al.
1.28
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
(2025)
Ali Behrouz et al.
1.28
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
(2025)
Danil Sivtsov et al.
1.28
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
(2025)
Jikai Jin et al.
1.28
Ovis2.5 Technical Report
(2025)
Shiyin Lu et al.
1.28
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
(2025)
Junsong Chen et al.
1.28
Why Can't Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls
(2025)
Xiaoyan Bai et al.
1.28
Native Hybrid Attention for Efficient Sequence Modeling
(2025)
Jusen Du et al.
1.28
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
(2025)
Hongyuan Tao et al.
1.28
Error-Free Linear Attention is a Free Lunch: Exact Solution from Continuous-Time Dynamics
(2025)
Jingdi Lei et al.
1.28
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
(2023)
Albert Gu et al.
β
Knowledge Composition using Task Vectors with Learned Anisotropic Scaling
(2024)
Frederic Z. Zhang et al.
β
LinFusion: 1 GPU, 1 Minute, 16K Image
(2024)
Songhua Liu et al.
β
Cottention: Linear Transformers With Cosine Attention
(2024)
Gabriel Mongaras et al.
β
If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs
(2024)
Muhammad Khalifa et al.
β