← all papers · overview

Isocompute Playbook: Optimally Scaling Sampling Compute For LLM RL

Abstract

While scaling laws guide compute allocation for LLM pre-training, analogous prescriptions for reinforcement learning (RL) post-training of large language models (LLMs) remain poorly understood. We study the compute-optimal allocation of sampling compute for on-policy RL methods in LLMs, framing scaling as a compute-constrained optimization over three resources: parallel rollouts per problem, numbe

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).