Key papers π₯ Trending (default) π Most cited π Newest first π€ A β Z by title 35 papers Β· trending (default) numbers = π₯ heat
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities (2026) Zichen Ding et al.
12.26 Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation (2026) Jasmine Brazilek et al.
5.87 SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents (2026) Tianming Sha et al.
4.39 The Patchwork Problem in LLM-Generated Code (2026) Viraaji Mothukuri et al.
4.39 SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation (2026) Sijia Gu et al.
4.39 Git-Assistant: Planning-Based Support for Updating Git Repositories (2026) Alfredo Garrach\'on Ruiz et al.
4.39 Diversifying to Verify: When Task-Equivalent Programs Differ in Verifiability (2026) Shirley Yu et al.
4.39 Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry (2026) Oriana Presacan et al.
4.39 The Hitchhiker's Guide to Monoculture (2026) Gordon Burtch
4.39 Analyzing Curricular Pattern Complexity Using AI to Improve On-Time Graduation Rates (2026) Lynn Vonderhaar et al.
4.39 SemaDiff: Identifying Semantic-Changing Commits with Generated Code and Tests (2026) Maha Ayub et al.
4.39 HEDGEHOG: Hierarchical Evaluation of Drug Generators Through Rigorous Filtration (2026) Daria A. Ryabchenko (Ligand Pro et al.
4.39 RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar (2026) Marek \v{S}uppa et al.
4.39 The Caf\'e in Amsterdam: When the Incumbent Becomes the Oracle (2026) Augusto Camargo
4.39 DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments (2026) Huatao Li et al.
4.39 Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning (2026) Chih-Hsuan Yang et al.
4.39 Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3? (2026) Sergey Rodionov
4.39 Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation (2026) Himel Dev et al.
4.39 Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data (2026) Kiwan Kwon et al.
4.39 AEGIS: Assay-Aware Protocol Validation and Runtime Monitoring for Open-Source Liquid Handling Robots (2026) Priyanka V. Setty et al.
4.39 Presentation, Not Mechanism: A Render Confound in Deprecation-Aware Memory Evaluation (2026) Zhaoyang Jiang et al.
4.39 Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance (2026) Ali Toygar Abak
4.39 Binding Drift in Multi-Step Tool-Augmented Agents (2026) Rahul Suresh Babu et al.
4.39 CODENS: Transforming Code Changes into Living, Accessible, and Queryable Documentation (2026) Abdelhak Kelious et al.
4.39 Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts (2026) Haodi Fan et al.
4.39 SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation (2026) Weifeng Sun et al.
4.39 Test Case Prioritization for DNNs via Neural Collapse Instability (2026) Chunyu Liu et al.
4.39 Multi-stage Dynamic Selection for Cross-Project Defect Prediction (2026) Juscimara G. Avelino et al.
4.39 TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language Models (2026) Amirreza Esmaeili et al.
3.51 Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety (2026) Lifei Liu et al.
3.51 STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle (2026) Sagar Deb et al.
3.51 CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data (2026) Vipul Gupta et al.
3.51 Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity (2026) Brett Reynolds
2.00 Agent Step Value: Auditing Evaluator-Channel Reversals in Black-Box Agent Traces (2026) Andrew Zhang et al.
2.00 AEVAL: From Anecdotal to Deterministic Testing for Agentic Skill Workflows (2026) Tejas Singh Anand et al.
2.00