Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
off-policy
loadingβ¦
π€
Ask AI
Awesome off-policy β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
off-policy
8 papers tagged off-policy β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
8 papers Β· trending (default)
numbers = π₯ heat
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
(2026)
Chenlu Ye et al.
1.94
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
(2026)
Zhong Guan et al.
1.94
Online Causal Kalman Filtering for Stable and Effective Policy Optimization
(2026)
Shuo He et al.
1.72
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
(2025)
Wenhao Zhang et al.
1.28
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
(2025)
Avishai Elmakies et al.
1.28
Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
(2025)
Siwei Wang et al.
1.28
The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
(2025)
Farid Bagirov et al.
1.28
SePPO: Semi-Policy Preference Optimization for Diffusion Alignment
(2024)
Daoan Zhang et al.
β