Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
workflows
loadingβ¦
π€
Ask AI
Awesome workflows β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
workflows
19 papers tagged workflows β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
19 papers Β· trending (default)
numbers = π₯ heat
AgenticDataBench: A Comprehensive Benchmark for Data Agents
(2026)
Zhaoyan Sun et al.
2.00
Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents
(2026)
Bowen Ye et al.
1.94
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
(2026)
Tingjia Miao et al.
1.94
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
(2026)
Sharareh Younesian et al.
1.94
Code as Agent Harness
(2026)
Xuying Ning et al.
1.94
EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory
(2026)
Chang Nie et al.
1.94
Boundary-Aware Context Grounding for A Low-Channel EEG Agent
(2026)
Zhiyuan Xu et al.
1.94
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents
(2026)
Shoufa Chen et al.
1.94
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions
(2026)
Mohit Raghavendra et al.
1.94
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
(2026)
Qianchu Liu et al.
1.94
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
(2026)
Ruofeng Yang et al.
1.89
ClawGym: A Scalable Framework for Building Effective Claw Agents
(2026)
Fei Bai et al.
1.83
LawFlow : Collecting and Simulating Lawyers' Thought Processes
(2025)
Debarati Das et al.
1.28
UniBiomed: A Universal Foundation Model for Grounded Biomedical Image Interpretation
(2025)
Linshan Wu et al.
1.28
LLM Context Conditioning and PWP Prompting for Multimodal Validation of Chemical Formulas
(2025)
Evgeny Markhasin
1.28
ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback
(2025)
Litao Guo et al.
1.28
Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
(2025)
Yueqi Song et al.
1.28
GenAgent: Build Collaborative AI Systems with Automated Workflow Generation -- Case Studies on ComfyUI
(2024)
Xiangyuan Xue et al.
β
Agent Workflow Memory
(2024)
Zora Zhiruo Wang et al.
β