Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
self-attention
loadingβ¦
π€
Ask AI
Awesome self-attention β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
self-attention
22 papers tagged self-attention β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
22 papers Β· trending (default)
numbers = π₯ heat
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing
(2026)
Tommaso Cerruti et al.
2.00
UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation
(2026)
Ruiheng Zhang et al.
1.94
MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers
(2026)
Ajay Jaiswal et al.
1.94
Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention
(2026)
Vishesh Tripathi et al.
1.94
Residual Stream Duality in Modern Transformer Architectures
(2026)
Yifan Zhang
1.78
LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
(2026)
Jiazheng Xing et al.
1.78
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
(2025)
Md Mohaiminul Islam et al.
1.28
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
(2025)
Weiming Ren et al.
1.28
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
(2025)
Xuan Ju et al.
1.28
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
(2025)
Zikang Liu et al.
1.28
Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling
(2025)
Rishiraj Acharya
1.28
Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
(2025)
Sangmin Bae et al.
1.28
How to Teach Large Multimodal Models New Skills
(2025)
Zhen Zhu et al.
1.28
Brainformers: Trading Simplicity for Efficiency
(2023)
Yanqi Zhou et al.
β
Localizing and Editing Knowledge in Text-to-Image Generative Models
(2023)
Samyadeep Basu et al.
β
Hiformer: Heterogeneous Feature Interactions Learning with Transformers for Recommender Systems
(2023)
Huan Gui et al.
β
DreamTuner: Single Image is Enough for Subject-Driven Generation
(2023)
Miao Hua et al.
β
MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation
(2024)
Kunpeng Song et al.
β
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations
(2024)
Can Qin et al.
β
LinFusion: 1 GPU, 1 Minute, 16K Image
(2024)
Songhua Liu et al.
β
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
(2024)
Hongyan Zhi et al.
β
MV-Adapter: Multi-view Consistent Image Generation Made Easy
(2024)
Zehuan Huang et al.
β