14B
Emerging1papers using it
2025first seen
The '14B' dataset/benchmark is used to evaluate the performance of Large Language Models (LLMs) in mathematical reasoning tasks, specifically focusing on the effectiveness of a novel CoT compression framework based on step entropy.