Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
environments
loadingβ¦
π€
Ask AI
Awesome environments β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
environments
22 papers tagged environments β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
22 papers Β· trending (default)
numbers = π₯ heat
ResearchGym: Evaluating Language Model Agents on Real-World AI Research
(2026)
Aniketh Garikaparthi et al.
1.94
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
(2026)
Dadi Guo et al.
1.94
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
(2026)
Hao Zhang et al.
1.94
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
(2026)
Han Wang et al.
1.94
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
(2026)
Zhekai Chen et al.
1.94
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
(2026)
Minrui Xu et al.
1.89
Synthetic Sandbox for Training Machine Learning Engineering Agents
(2026)
Yuhang Zhou et al.
1.83
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
(2026)
Mingyu Ouyang et al.
1.83
ClawGym: A Scalable Framework for Building Effective Claw Agents
(2026)
Fei Bai et al.
1.83
AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
(2026)
Weiyi Wang et al.
1.67
ASTRA: Automated Synthesis of agentic Trajectories and Reinforcement Arenas
(2026)
Xiaoyu Tian et al.
1.67
Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills
(2025)
Yuquan Xie et al.
1.28
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
(2025)
Guibin Zhang et al.
1.28
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
(2025)
Darshan Deshpande et al.
1.28
CWM: An Open-Weights LLM for Research on Code Generation with World Models
(2025)
FAIR CodeGen team et al.
1.28
Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations
(2025)
Jinkun Chen et al.
1.28
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
(2025)
Zhiyuan Zeng et al.
1.28
From Word to World: Can Large Language Models be Implicit Text-based World Models?
(2025)
Yixia Li et al.
1.28
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
(2023)
Wenlong Huang et al.
β
Holodeck: Language Guided Generation of 3D Embodied AI Environments
(2023)
Yue Yang et al.
β
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
(2024)
Baichuan Zhou et al.
β
Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
(2024)
Rogerio Bonatti et al.
β