WebArena
Canonical21papers using it
2024first seen
WebArena is a benchmark dataset used to evaluate the performance of large language model web agents by measuring their ability to execute structured tool actions based on web interactions.
Papers using WebArena (20)
- Mobile-agent-v3.5: Multi-platform Fundamental GUI AgentsMulti-Agent Transactive MemoryDevil's Advocate: Anticipatory Reflection For LLM AgentsThe Compliance Trap: Diagnosing How AI Agents Consume Conflicting MemoryOnline Skill Learning for Web Agents via State-Grounded Dynamic RetrievalBeyond Domains: Reusing Web Skills via Transferable Interaction PatternsWeasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data SelectionAPEX: Autonomous Policy Exploration for Self-Evolving LLM AgentsAdarubric: Task-adaptive Rubrics For LLM Agent EvaluationDoes The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web AgentsBranch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action MemoryWebUncertainty: Dual-Level Uncertainty Driven Planning and Reasoning For Autonomous Web AgentAgenther: Hindsight Experience Replay For LLM Agent Trajectory RelabelingEnvironment Maps: Structured Environmental Representations For Long-horizon AgentsOpAgent: Operator Agent for Web NavigationJust-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient UpdatesSkyRL-Agent: Efficient RL Training for Multi-turn LLM AgentWEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at ScaleSurfer 2: The Next Generation Of Cross-platform Computer Use AgentsCoAct: A Global-Local Hierarchy for Autonomous Agent Collaboration