MATH
Emerging14papers using it
2024first seen
The 'MATH' dataset is a benchmark that contains mathematical problems used to evaluate the performance of large language models in solving complex reasoning tasks.
Papers using MATH (10)
- DynaGraph: Lightweight Multi-Model Interaction Framework via Dynamic Topological ReconfigurationKV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time ScalingGraph-of-Agents: A Graph-based Framework for Multi-Agent LLM CollaborationCARD: Towards Conditional Design of Multi-agent Topological StructuresPlan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM ReasoningHow to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMsPutting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With VerifiersDebFlow: Automating Agent Creation via Agent DebateGATE: Graph-based Adaptive Tool Evolution Across Diverse TasksQ*: Improving Multi-step Reasoning for LLMs with Deliberative Planning