Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
evaluations
loadingβ¦
π€
Ask AI
Awesome evaluations β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
evaluations
12 papers tagged evaluations β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
12 papers Β· trending (default)
numbers = π₯ heat
Analyzing Toxic Behavior and Its Impact on the Mastodon Community
(2026)
Pasan Kamburugamuwa et al.
5.01
NodeRAG: Structuring Graph-based RAG with Heterogeneous Nodes
(2025)
Tianyang Xu et al.
1.28
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
(2025)
Mingjie Liu et al.
1.28
An Agentic System for Rare Disease Diagnosis with Traceable Reasoning
(2025)
Weike Zhao et al.
1.28
ModelCitizens: Representing Community Voices in Online Safety
(2025)
Ashima Suvarna et al.
1.28
ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
(2024)
Siming Yan et al.
β
Trajectory Consistency Distillation
(2024)
Jianbin Zheng et al.
β
SafeInfer: Context Adaptive Decoding Time Safety Alignment for Large Language Models
(2024)
Somnath Banerjee et al.
β
Dolphin: Long Context as a New Modality for Energy-Efficient On-Device Language Models
(2024)
Wei Chen et al.
β
OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision
(2024)
Cong Wei et al.
β
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
(2024)
Shivalika Singh et al.
β
Contextualized Counterspeech: Strategies for Adaptation, Personalization, and Evaluation
(2024)
Lorenzo Cima et al.
β