OSWorld
Canonical30papers using it
2024first seen
The 'OSWorld' dataset is a benchmark used to evaluate GUI automation tasks and tool-calling tasks for open-source models.
Papers using OSWorld (23)
- Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use AgentsMobile-agent-v3.5: Multi-platform Fundamental GUI AgentsOS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using AgentWindows Agent Arena: Evaluating Multi-modal OS Agents At ScaleInfantagent-next: A Multimodal Generalist Agent For Automated Computer InteractionLearning from Failure: Inference-Time Self-Improvement for Computer-Use AgentsRecovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI AgentsIntentScore: Intent-Conditioned Action Evaluation for Computer-Use AgentsAgent Alpha: Tree Search Unifying Generation, Exploration and Evaluation for Computer-Use AgentsAgent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path ForwardInfiniteWeb: Scalable Web Environment Synthesis for GUI Agent TrainingBEAP-Agent: Backtrackable Execution and Adaptive Planning for GUI AgentsSurfer 2: The Next Generation Of Cross-platform Computer Use AgentsScaling Agents For Computer UseAgentic Lybic: Multi-Agent Execution System with Tiered Reasoning and OrchestrationMano Technical ReportInstruction Agent: Enhancing Agent With Expert DemonstrationUI-TARS-2 Technical Report: Advancing GUI Agent With Multi-turn Reinforcement LearningCoAct-1: Computer-using Multi-Agent System with Coding ActionsOSWorld-Human: Benchmarking the Efficiency of Computer-Use AgentsUi-evol: Automatic Knowledge Evolving For Computer Use AgentsOsworld: Benchmarking Multimodal Agents For Open-ended Tasks In Real Computer EnvironmentsAgent S: An Open Agentic Framework that Uses Computers Like a Human