Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
Verifiable
loadingβ¦
π€
Ask AI
Awesome Verifiable β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
Verifiable
19 papers tagged Verifiable β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
19 papers Β· trending (default)
numbers = π₯ heat
RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation
(2026)
Sunzhu Li et al.
1.94
Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR
(2026)
Fanfan Liu et al.
1.94
Beyond Length Scaling: Synergizing Breadth and Depth for Generative Reward Models
(2026)
Qiyuan Zhang et al.
1.94
CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test
(2026)
Zhangyi Hu et al.
1.94
Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration
(2026)
Zili Wang et al.
1.94
Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs
(2025)
Yufa Zhou et al.
1.28
Perception-Aware Policy Optimization for Multimodal Reasoning
(2025)
Zhenhailong Wang et al.
1.28
The Invisible Leash: Why RLVR May Not Escape Its Origin
(2025)
Fang Wu et al.
1.28
VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
(2025)
Dongfu Jiang et al.
1.28
ΞL Normalization: Rethink Loss Aggregation in RLVR
(2025)
Zhiyuan He et al.
1.28
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
(2025)
Zhilin Wang et al.
1.28
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
(2025)
Long Xing et al.
1.28
Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
(2025)
Xin Xu et al.
1.28
ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
(2025)
Yuqi Liu et al.
1.28
The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
(2025)
Farid Bagirov et al.
1.28
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
(2025)
Yihe Deng et al.
1.28
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
(2025)
Zhiyuan Zeng et al.
1.28
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
(2025)
Songyang Gao et al.
1.28
OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification
(2025)
Zijian Wu et al.
1.28