ALFWorld
Canonical99papers using it
2023first seen
ALFWorld is a dataset and benchmark designed to evaluate the ability of language agents to ground their actions in contextual information from their environment.
Papers using ALFWorld (89)
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement LearningHarnessX: A Composable, Adaptive, and Evolvable Agent Harness FoundrySkill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement LearningTurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent TrainingRLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon AgentsSKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit AssignmentBlueprint First, Model Second: A Framework for Deterministic LLM WorkflowWhat and When to Distill: Selective Hindsight Distillation for Multi-Turn AgentsSelf-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making TasksMetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill EvolutionBiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM AgentsGraph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent SkillsKnowAgent: Knowledge-Augmented Planning for LLM-Based AgentsWorld Model Implanting for Test-time Adaptation of Embodied AgentsLatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM AgentsMulti-Agent Transactive MemoryBeyond Policy Optimization: A Data Curation Flywheel for Sparse-Reward Long-Horizon PlanningNo Time Like the Present: Agentic Test-Time Training for LLM AgentsRSPO: Reward-Swap Policy Optimization for Multi-Turn LLM AgentsSTAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent TrainingTask Decomposition-Guided Reranking for Adaptive Agent Skill RetrievalSkill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference LearningUnified Context Evolution for LLM AgentsSIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent TrainingSkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at ScaleAdaMEM: Test-Time Adaptive Memory for Language AgentsSelf-evolving LLM agents with in-distribution Optimization3SPO: State-Score-Supervised Policy Optimization for LLM AgentsOrganize then Retrieve: Hierarchical Memory Navigation for Efficient AgentsOn-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn AgentsACCORD: Action-Conditioned Contextual Grounding for Language AgentsEnvRL: Learn from Environment Dynamics in Agentic Reinforcement LearningUncertainty Decomposition for Clarification Seeking in LLM AgentsThe Interplay of Harness Design and Post-Training in LLM AgentsSemantic Consistency Policy Optimization for Reinforcement Learning of LLM AgentsJoint Learning of Experiential Rules and Policies for Large Language Model AgentsATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic TasksDuoMem: Towards Capable On-Device Memory Agents via Dual-Space DistillationTraining Language Agents to Learn from ExperienceHera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM AgentsStepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement LearningSkill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM AgentsHonest Lying: Understanding Memory Confabulation in Reflexive AgentsSkillsInjector: Dynamic Skill Context Construction for LLM AgentsExpGraph: Model-Agnostic Experience Learning with Graph-Structured Memory for LLM AgentsSkill Reuse as Compression in Agentic RLWhere LLM Agents Fail And How They Can Learn From FailuresSpinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful DemonstrationsProgress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement LearningRetrospective Progress-Aware Self-Refinement for LLM Agent TrainingPatchBoard: Schema-Grounded State Mutation for Reliable and Auditable LLM Multi-Agent CollaborationSKILL0: In-Context Agentic Reinforcement Learning for Skill InternalizationRewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon AgentsPolicy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language AgentsWhen Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent TrainingSelective Memory Retention for Long-Horizon LLM AgentsUCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-DistillationSelf-Evolving World Models for LLM Agent PlanningComplete Cyclic Subtask Graphs For Tool-using LLM Agents: Flexibility, Cost, And Bottlenecks In Multi-agent WorkflowsTAPE: Tool-guided Adaptive Planning And Constrained Execution In Language Model AgentsReAct-Diffuse: An Integrated Agentic and Generative Diffusion Framework for Autonomous Multi-Step Task Reasoning and ExecutionGrasp: Graph-structured Skill Compositions For LLM AgentsDynamic Skill Lifecycle Management for Agentic Reinforcement LearningSkills on the Fly: Test-Time Adaptive Skill Synthesis for LLM AgentsHierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM AgentsFrom Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM AgentsAsk Only When Needed: Proactive Retrieval from Memory and Skills for Experience-Driven Lifelong AgentsDPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based AgentsHiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM AgentsDynamic Dual-Granularity Skill Bank for Agentic RLSkillnet: Create, Evaluate, And Connect AI SkillsMemSkill: Learning and Evolving Memory Skills for Self-Evolving AgentsDynamic Mixed-Precision Routing for Efficient Multi-step LLM InteractionEmbodied Task Planning via Graph-Informed Action Generation with Large Language ModelsReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process RoutingPADME: Procedure Aware DynaMic ExecutionNeSyPr: Neurosymbolic Proceduralization For Efficient Embodied ReasoningGraph-Enhanced Policy Optimization in LLM Agent TrainingAutocontext: Instance-level Context Learning For LLM AgentsIntrinsic Memory Agents: Heterogeneous Multi-agent LLM Systems Through Structured Contextual MemoryFact-Augmented Lookahead Planning for LLM AgentsUnleashing Embodied Task Planning Ability in LLMs via Reinforcement LearningDivide, Optimize, Merge: Fine-Grained LLM Agent Optimization at ScaleStructured Agent Distillation for Large Language ModelDivide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningDebFlow: Automating Agent Creation via Agent DebateBetter Than Your Teacher: LLM Agents That Learn From Privileged AI FeedbackAdaPlanner: Adaptive Planning from Feedback with Language ModelsADaPT: As-Needed Decomposition and Planning with Language Models