Qwen-2.5-math-7B
Emerging5papers using it
2025first seen
The 'Qwen-2.5-Math-7B' dataset/benchmark is used to evaluate reinforcement learning with verifiable rewards (RLVR) in training large language models on deterministic outcome reasoning tasks.
Papers using Qwen-2.5-math-7B (5)
- SFT-then-RL Outperforms Mixed-Policy Methods for LLM ReasoningBeyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional PairingOn the optimization dynamics of RLVR: Gradient gap and step size thresholdsDCPO: Dynamic Clipping Policy OptimizationIncentivizing LLMs to Self-Verify Their Answers