Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
automatic
loadingβ¦
π€
Ask AI
Awesome automatic β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
automatic
17 papers tagged automatic β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
17 papers Β· trending (default)
numbers = π₯ heat
Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon
(2026)
VΓctor Gallego
1.94
MetaAgent-X : Breaking the Ceiling of Automatic Multi-Agent Systems via End-to-End Reinforcement Learning
(2026)
Yaolun Zhang et al.
1.94
SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking
(2026)
Guohong Liu et al.
1.94
VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
(2025)
Zuwei Long et al.
1.28
Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models
(2025)
Zhiyuan Hu et al.
1.28
RExBench: Can coding agents autonomously implement AI research extensions?
(2025)
Nicholas Edwards et al.
1.28
AHELM: A Holistic Evaluation of Audio-Language Models
(2025)
Tony Lee et al.
1.28
arXiVeri: Automatic table verification with GPT
(2023)
Gyungin Shin et al.
β
VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use
(2023)
Yonatan Bitton et al.
β
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
(2024)
Rohan Wadhawan et al.
β
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
(2024)
Xiaoyi Dong et al.
β
ArzEn-LLM: Code-Switched Egyptian Arabic-English Translation and Speech Recognition Using LLMs
(2024)
Ahmed Heakl et al.
β
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
(2024)
Philippe Laban et al.
β
Very Large-Scale Multi-Agent Simulation in AgentScope
(2024)
Xuchen Pan et al.
β
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
(2024)
Yang Liu et al.
β
OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision
(2024)
Cong Wei et al.
β
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
(2024)
Renqiu Xia et al.
β