Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
questions
loadingβ¦
π€
Ask AI
Awesome questions β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
questions
22 papers tagged questions β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
22 papers Β· trending (default)
numbers = π₯ heat
CoDiQ: Test-Time Scaling for Controllable Difficult Question Generation
(2026)
Zhongyuan Peng et al.
1.94
Hierarchical Abstract Tree for Cross-Document Retrieval-Augmented Generation
(2026)
Ziwen Zhao et al.
1.89
When AI Navigates the Fog of War
(2026)
Ming Li et al.
1.78
V2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multi-Modal Large Language Models
(2025)
Hsu-kuang Chiu et al.
1.28
ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering
(2025)
Kaisi Guan et al.
1.28
FREESON: Retriever-Free Retrieval-Augmented Reasoning via Corpus-Traversing MCTS
(2025)
Chaeeun Kim et al.
1.28
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
(2025)
Run-Ze Fan et al.
1.28
What Is Your AI Agent Buying? Evaluation, Implications and Emerging Questions for Agentic E-Commerce
(2025)
Amine Allouah et al.
1.28
ViExam: Are Vision Language Models Better than Humans on Vietnamese Multimodal Exam Questions?
(2025)
Vy Tuong Dang et al.
1.28
ReviewScore: Misinformed Peer Review Detection with Large Language Models
(2025)
Hyun Ryu et al.
1.28
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
(2025)
Jingru Lin et al.
1.28
Multi-hop Reasoning via Early Knowledge Alignment
(2025)
Yuxin Wang et al.
1.28
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
(2024)
Haoji Zhang et al.
β
Self-Recognition in Language Models
(2024)
Tim R. Davidson et al.
β
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
(2024)
Fanqing Meng et al.
β
Text2SQL is Not Enough: Unifying AI and Databases with TAG
(2024)
Asim Biswal et al.
β
PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
(2024)
Yijia Shao et al.
β
Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
(2024)
Satyapriya Krishna et al.
β
Revealing the Barriers of Language Agents in Planning
(2024)
Jian Xie et al.
β
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
(2024)
Chengke Zou et al.
β
Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models
(2024)
YiFan Zhang et al.
β
MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models
(2024)
Mahir Labib Dihan et al.
β