← all papers · overview

Automated Rewards Via Llm-generated Progress Functions

Abstract

Large Language Models (LLMs) have the potential to automate reward engineering by leveraging their broad domain knowledge across various tasks. However, they often need many iterations of trial-and-error to generate effective reward functions. This process is costly because evaluating every sampled reward function requires completing the full policy optimization process for each function. In this

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).