OSWorld
Emerging10papers using it
2025first seen
The 'OSWorld' dataset/benchmark contains a variety of tasks used to evaluate the performance of reinforcement learning systems, specifically in the context of large language models and agentic scenarios.
Papers using OSWorld (10)
- Reinforcement Learning for Computer-Use Agents with Autonomous EvaluationIntentScore: Intent-Conditioned Action Evaluation for Computer-Use AgentsRLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL SystemOS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic ModelsSTEP: Success-Rate-Aware Trajectory-Efficient Policy OptimizationUI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement LearningEfficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data CurationComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use AgentsMobile-Agent-v3: Fundamental Agents for GUI AutomationZeroGUI: Automating Online GUI Learning at Zero Human Cost