← all papers · overview

Visualising Policy-reward Interplay To Inform Zeroth-order Preference Optimisation Of Large Language Models

Abstract

Fine-tuning Large Language Models (LLMs) with first-order methods like back-propagation is computationally intensive. Zeroth-Order (ZO) optimisation uses function evaluations instead of gradients, reducing memory usage, but suffers from slow convergence in high-dimensional models. As a result, ZO research in LLMs has mostly focused on classification, overlooking more complex generative tasks. In t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).