AMC 2023
Emerging1papers using it
2025first seen
The 'AMC 2023' dataset/benchmark is used to evaluate the reasoning capabilities and conciseness of Large Language Models (LLMs) in generating long Chain-of-Thought (CoT) responses.
The 'AMC 2023' dataset/benchmark is used to evaluate the reasoning capabilities and conciseness of Large Language Models (LLMs) in generating long Chain-of-Thought (CoT) responses.