Awesome Large Language Models
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Mayank Mishra — most-cited papers & profile · Large Language Models
← authors
·
overview
Mayank Mishra
11
papers ·
177
citations
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Reducing Transformer Key-value Cache Size With Cross-layer Attention
2024 · 109 citations
Taylor-Calibrate: Principled Initialization for Hybrid Linear Attention Distillation
2026
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
2025
StarCoder: may the source be with you!
2023
StarCoder 2 and The Stack v2: The Next Generation
2024
Aurora-M: The First Open Source Multilingual Language Model Red-teamed according to the U.S. Executive Order
2024
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
2024
Top co-authors
Niklas Muennighoff
· 3
Terry Yue Zhuo
· 3
Tri Dao
· 3
Alex Gu
· 2
Arjun Guha
· 2
Armel Zebaze
· 2
Carlos Muñoz Ferrandis
· 2
Carolyn Jane Anderson
· 2
Chenghao Mou
· 2
Christopher Akiki
· 2
Denis Kocetkov
· 2
Dmitry Abulkhanov
· 2
Topics
Efficiency
Model Architecture
Training Techniques
Code
Large
Language
Models
LLMs
The
Stack