Key papers π₯ Trending (default) π Most cited π Newest first π€ A β Z by title 60 papers Β· trending (default) numbers = π₯ heat
EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration (2026) Jia-Kai Dong et al.
7.37 Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation (2026) Jasmine Brazilek et al.
5.87 Validating the Single Item Kawaii Measure (2026) Katie Seaborn et al.
5.01 The Geometry of Personality: Activation Steering with Jungian Cognitive Functions (2026) Liu Zai (University of Glasgow) et al.
5.01 SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents (2026) Tianming Sha et al.
4.39 HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration (2026) Chang Liu et al.
4.39 Analyzing Curricular Pattern Complexity Using AI to Improve On-Time Graduation Rates (2026) Lynn Vonderhaar et al.
4.39 AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized (2026) Chiara Marcoccia et al.
4.39 CoWeaver: A Bi-directional, Learnable and Explainable Matching Engine for Mixed Human-Agent Science Collaboration (2026) Jiayao Gu et al.
4.39 When Not to Automate: A Formal Protocol for Human Preservation in AI-Optimized Organizations (2026) Jose Manuel de la Chica Rodriguez et al.
4.39 LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization (2026) Mazene Ameur et al.
4.39 Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent (2026) Shiva Pochampally et al.
4.39 Public perceptions of AI-driven decision-making in healthcare: A structural equation modeling approach (2026) Leonie Westerbeek et al.
4.39 The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems (2026) Gjergji Kasneci et al.
4.39 DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making (2026) Raffi Khatchadourian
4.39 Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning (2026) Ajay Patel et al.
3.51 Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries (2026) Mohammad Arvan et al.
3.51 Towards an Automated Test of LLM Security Knowledge (2026) Shufan Chai et al.
3.51 Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists (2026) Arnavi Chheda-Kothary et al.
3.51 AutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalism (2026) Himel Ghosh et al.
3.51 Generative and multimodal AI for materials prediction and design: Progress, challenges, and perspectives (2026) Xianyuan Liu et al.
3.51 Behavior-Adaptive Conversational Agents: Toward a Fluid Personality Framework (2026) Hasibur Rahman et al.
2.00 BrainPilot: Automating Brain Discovery with Agentic Research (2026) Haoxuan Li et al.
2.00 AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets (2026) Ming Chen et al.
2.00 When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering (2026) Ziteng Hu et al.
2.00 Control panels to clarify user intent with Large Language Models (2026) Ben Shneiderman
2.00 Analyzing Middle School Students' Dialogue and Behaviors during Collaborative AI Chatbot Development Using Ordered Network Analysis (2026) Shan Zhang et al.
2.00 From Obligation to Specification: A Survey on Validating EU AI Act Requirements in RE (2026) T. Y. Emmy Lai et al.
2.00 A Systematic Survey on Image Description Techniques for STEM Domains (2026) Marco Cardia et al.
2.00 Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions (2026) Pin Qian et al.
2.00 Defining AI-Native Systems: Autonomy as Revision Authority (2026) Cheng Tan
2.00 What AI Red-Team Evaluations Can and Cannot Prove (2026) Bandana Kaur
2.00 Co-design of LLM-based preference agents: participation may drive overtrust (2026) Michael J. Fell
2.00 AI-Integrated Scientific Inquiry: A Practice-Centered Vision for Science Education (2026) Arne Bewersdorff et al.
2.00 How Do AI Coding Agents Contribute to Software Development? an Empirical Study of Agentic Pull Requests (2026) Iren Mazloomzadeh et al.
2.00 ToolGuardian: Declarative Security for AI Agent-Tool Interactions (2026) Arun Ravindran et al.
2.00 Multi-Agent System-driven Digital Twins for predictive maintenance: architectures, technologies and open research challenges (2026) Korota Ars\`ene Coulibaly et al.
2.00 Benchmarking Text-to-SQL under Role-Based Access Control (2026) Yang Fei et al.
2.00 Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity (2026) Pengzhao Lyu et al.
2.00 AI4PLE: A Methodology for Integrating AI into Product Line Engineering (2026) Bedir Tekinerdogan
2.00 A Roadmap to Impactful Pluralistic Alignment Research (2026) Elinor Poole-Dayan et al.
2.00 Towards Trustworthy and Cost-Efficient Data Integration: From Na\"ive RAG to Agentic RAG (2026) Chuangtao Ma et al.
2.00 Teachy Mini: Development and Preliminary Evaluation of a Knowledge-Based Generative Social Robot for Higher Education (2026) Stephan Vonschallen et al.
2.00 Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI (2026) Jiaqi Shao et al.
2.00 Kutti AI: A Voice-First, Offline-Capable Learning Companion with Real-Time Struggle Detection for Visually-Impaired Children (2026) Kadharmoideen Fadurudeen
2.00 A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource Allocation (2026) Fin Gentzen et al.
2.00 Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture (2026) Halil Burak Noyan
2.00 Beyond Perspectives: A Trio-Ethnography of Interpretation Evolution in LLM-Supported Programming Education (2026) Jennie Ren et al.
2.00 CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference (2026) Jiyuan Tan et al.
2.00 ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents (2026) Jianan Ma et al.
1.89 WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics (2026) Sneha Maurya et al.
1.83 Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control (2026) Mahiro Nakao et al.
1.83 Math Education Digital Shadows for Investigating Learning with GenAI: Mathematics Performance, Anxiety, and Confidence in LLMs (2026) Naomi Esposito et al.
1.83 Administrative Law's Fourth Settlement: AI and the Scrutable State (2026) Nicholas Caputo
1.72 LLAMA LIMA: A Living Meta-Analysis on the Effects of Generative AI on Learning Mathematics (2026) Anselm Strohmaier et al.
1.67 InteractComp: Evaluating Search Agents With Ambiguous Queries (2025) Mingyi Deng et al.
1.50 From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem (2025) James Jewitt et al.
1.44 From Mind to Machine: The Rise of Manus AI as a Fully Autonomous Digital Agent (2025) Minjie Shen et al.
1.22 When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas (2025) Steffen Backmann et al.
1.22 Analyzing the Ethical Logic of Eight Large Language Models (2025) W. Russell Neuman et al.
1.00