← all papers · overview

Agentnoisebench: Benchmarking Robustness Of Tool-using LLM Agents Under Noisy Condition

Abstract

Recent advances in large language models have enabled LLM-based agents to achieve strong performance on a variety of benchmarks. However, their performance in real-world deployments often that observed on benchmark settings, especially in complex and imperfect environments. This discrepancy largely arises because prevailing training and evaluation paradigms are typically built on idealized assumpt

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).