← all papers · overview

To Mix Or To Merge: Toward Multi-domain Reinforcement Learning For Large Language Models

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) plays a key role in stimulating the explicit reasoning capability of Large Language Models (LLMs). We can achieve expert-level performance in some specific domains via RLVR, such as coding or math. When a general multi-domain expert-level model is required, we need to carefully consider the collaboration of RLVR across different domains. The cu

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).