Key papers π₯ Trending (default) π Most cited π Newest first π€ A β Z by title 60 papers Β· trending (default) numbers = π₯ heat
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents (2026) Xiangchen Cheng et al.
12.42 AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents (2026) Kunlun Zhu et al.
11.84 AgenticDataBench: A Comprehensive Benchmark for Data Agents (2026) Zhaoyan Sun et al.
11.22 SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use (2026) Jiayin Zhu et al.
10.21 PACE: A Proxy for Agentic Capability Evaluation (2026) Yueqi Song et al.
8.24 NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs (2026) Jiarong Zhao et al.
7.96 Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges (2026) Tuo Liang et al.
5.87 Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism (2026) Minjong Cheon
5.49 Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks (2026) Hiroki Tamba
5.49 ExPerT: Personalizing LLM Responses to Users' Domain Expertise via Query-Wise Semantic and Keystroke Behavioral Cues (2026) Yeji Park et al.
5.01 Office Comprehension Benchmark (2026) Firoz Shaik et al.
5.01 TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue (2026) Hao Zhang et al.
5.01 MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering (2026) Dang Quang Thien Tran et al.
5.01 IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs (2026) Samir Abdaljalil et al.
5.01 Grounded Optimization: A Layered Engineering Framework for Reducing LLM Hallucination in Automated Personal Document Rewriting (2026) Shashank Indukuri et al.
5.01 DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents (2026) Tianyi Zhang et al.
5.01 Evaluating Chunking Strategies for Retrieval-Augmented Generation on Academic Texts (2026) Valentin J. J. Kreileder et al.
5.01 TUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27B (2026) Baran Bingol et al.
5.01 AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations (2026) Javier Irigoyen et al.
5.01 PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation (2026) Peng Yun et al.
5.01 Object Aligner: A Configurable JSON Schema Similarity Score for Graphs, Applied to LLM Prompt Optimization (2026) Jan Drchal
5.01 OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets (2026) Rheeya Uppaal et al.
5.01 SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses (2026) Anna Chorna
5.01 AutoIndex: Learning Representation Programs for Retrieval (2026) Sam O'Nuallain et al.
5.01 Bounding Boxes to Improve Small Language Model Performance on Vision-Based Grading Tasks (2026) Lachlan McGinness
5.01 Measuring Reward-Seeking via Contrastive Belief Updates (2026) Axel H{\o}jmark et al.
5.01 GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus (2026) Daekeun Kim
5.01 Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception (2026) Ali Asad et al.
5.01 The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs (2026) Adam Rigby et al.
5.01 ShriNep@EEUCA 2026: RAKSHAK - Multi-Task DeBERTa with Rationale Distillation and Jigsaw-Augmented Training for Toxic Intent Classification (2026) Binayak Karki et al.
5.01 Semantic Field Theory: Historical Origin, Higher-Order Interaction, and Stabilized Semantic Inference (2026) Dimitris Vartziotis
5.01 A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction (2026) Jessica Sena et al.
5.01 RE-AD: Real-Time Requirement Adherence for Data Labeling (2026) Siddarth Malreddy et al.
5.01 Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions? (2026) Yuzhi Tang et al.
5.01 Can Valence Reflect Morality in Natural Language? A Preliminary Annotation Study (2026) Jonny O'Dwyer et al.
5.01 Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1 (2026) Sarah Y. Li et al.
5.01 From Agent Failures to Text Policies: What Works and What Breaks (2026) Jaideep Ray et al.
5.01 Rushes: A Human Preference Dataset for Pluralistic Alignment (2026) Michael Xu et al.
5.01 Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles (2026) Donghwan Kim
5.01 The Geometry of Personality: Activation Steering with Jungian Cognitive Functions (2026) Liu Zai (University of Glasgow) et al.
5.01 CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation (2026) Hehao Zhang et al.
5.01 QuantiBias: Benchmarking Quantization-Induced Bias in LLMs (2026) Emilio Ferrara
5.01 One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies (2026) Minh Ngoc Ta et al.
5.01 slang.gr as a Large-Scale Crowdsourced Resource for Non-Standard Greek (2026) Panagiotis Papadakos et al.
5.01 Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models (2026) Netanel Eliav
5.00 Learning User-Aware Recall: Personalized Retrieval in Long-Term Conversational Memory (2026) ZhiShu Jiang et al.
4.39 Structuring the Space of Sociotechnical Alignment (2026) Esra D\"onmez et al.
4.39 Black-Box Inference of LLM Architectural Properties with Restrictive API Access (2026) Christopher Ellis et al.
4.39 Risk Architecture for AI-Native Engineering Teams: An Organizational Framework for Agentic System Governance (2026) Laxmipriya Ganesh Iyer
4.39 On the Utility and Factual Reliability of Pruned Mixture-of-Experts Models in the Biomedical Domain (2026) Atsuki Yamaguchi et al.
4.39 Multi-Head Recurrent Memory Agents (2026) Jiatong Li et al.
4.39 Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation (2026) Junyi Wen et al.
4.39 Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing (2026) Joshua Penman
4.39 Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge (2026) Alex Brooker et al.
4.39 Safety Targeted Embedding Exploit via Refinement (2026) Joshua Adrian Cahyono
4.39 Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing (2026) Siyuan Li et al.
4.39 Steerability via constraints: a substrate for scalable oversight of coding agents (2026) Thomas Winninger
4.39 Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach (2026) Manuel Alonso-Carracedo et al.
4.39 What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates (2026) Arman Ghaffarizadeh et al.
4.39 AEVAL: From Anecdotal to Deterministic Testing for Agentic Skill Workflows (2026) Tejas Singh Anand et al.
4.39