Reward-directed Score-based Diffusion Models Via Q-learning
2024 Β· Xuefeng Gao, Jiale Zha, Xun Yu Zhou
Abstract
We propose a new reinforcement learning (RL) formulation for training continuous-time score-based diffusion models for generative AI to generate samples that maximize reward functions while keeping the generated distributions close to the unknown target data distributions. Different from most existing studies, ours does not involve any pretrained model for the unknown score functions of the noise-perturbed data distributions, nor does it attempt to learn the score functions. Instead, we formulate the problem as entropy-regularized continuous-time RL and show that the optimal stochastic policy has a Gaussian distribution with a known covariance matrix. Based on this result, we parameterize the mean of Gaussian policies and develop an actor--critic type (little) q-learning algorithm to solve the RL problem. A key ingredient in our algorithm design is to obtain noisy observations from the unknown score function via a ratio estimator. Our formulation can also be adapted to solve pure score
Authors
(none)
Tags
Stats
Related papers
- Learning A Diffusion Model Policy From Rewards Via Q-score Matching (2023)0.00
- Scores As Actions: A Framework Of Fine-tuning Diffusion Models By Continuous-time Reinforcement Learning (2024)0.00
- Diffusion Policy Through Conditional Proximal Policy Optimization (2026)0.00
- Sampling From Energy-based Policies Using Diffusion (2024)0.00
- Q-learning In Continuous Time (2022)0.00
- Goal-driven Reward By Video Diffusion Models For Reinforcement Learning (2025)0.00
- Bellman Diffusion: Generative Modeling As Learning A Linear Operator In The Distribution Space (2024)0.00
- A Diffusion Model Framework For Maximum Entropy Reinforcement Learning (2025)0.00