AppWorld
Emerging11papers using it
2025first seen
The 'AppWorld' dataset is a benchmark used to evaluate the effectiveness of context engineering approaches for large language model agents in decision-making tasks.
Papers using AppWorld (11)
- Hera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM AgentsReinforcement Learning for Long-Horizon Interactive LLM AgentsKeep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use AgentsLearning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM AgentsHINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon AgentsCLEAR: Context Augmentation from Contrastive Learning of Experience via Agentic ReflectionSkill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM AgentsSeeUPO: Sequence-Level Agentic-RL with Convergence GuaranteesCuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RLReinforcement Learning for Self-Improving Agent with Skill LibrarySALT: Step-level Advantage Assignment for Long-horizon Agents via Trajectory Graph