AppWorld
Emerging15papers using it
2025first seen
The 'AppWorld' dataset/benchmark contains a collection of applications and their associated contexts, used to evaluate the ability of language agents to ground user instructions in the relevant environmental information.
Papers using AppWorld (15)
- From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness FlawsACCORD: Action-Conditioned Contextual Grounding for Language AgentsHera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM AgentsExpGraph: Model-Agnostic Experience Learning with Graph-Structured Memory for LLM AgentsPushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and ActivationMetis: Bridging Text and Code Memory for Self-Evolving AgentsKeep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use AgentsLearning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM AgentsHINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon AgentsCoEvolve: Training LLM Agents via Agent-Data Mutual EvolutionThree Roles, One Model: Role Orchestration At Inference Time To Close The Performance Gap Between Small And Large AgentsToward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM AgentsReinforcement Learning for Self-Improving Agent with Skill LibraryACON: Optimizing Context Compression for Long-horizon LLM AgentsProst: Progressive Sub-task Training For Pareto-optimal Multi-agent Systems Using Small Language Models