Sokoban
Emerging13papers using it
15HF downloads
0HF likes
2025first seen
The 'Sokoban' dataset/benchmark contains a series of puzzle scenarios used to evaluate the performance of reinforcement learning agents in solving multi-turn interactive tasks.
Papers using Sokoban (13)
- Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMsLearning to Search and Searching to Learn for Generalization in PlanningSkill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM AgentsFreshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement LearningHiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM AgentsProAct: Agentic Lookahead in Interactive EnvironmentsTSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM AgentsPaying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM AgentsMeta-RL Induces Exploration in Language AgentsDyna-Mind: Learning to Simulate from Experience for Better AI AgentsInternalizing World Models via Self-Play Finetuning for Agentic RLCogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and PlanningInterpreting Emergent Planning in Model-Free Reinforcement Learning