β all topics overview
loadingβ¦
on-policy is one of the most active areas in Awesome Large Language Models β 25 papers in this collection. A strong starting point is "TREK: Distill to Explore, Reinforce to Refine".