MT-Bench
Emerging6papers using it
2022first seen
The 'MT-Bench' is a benchmark designed to evaluate the performance of large language models in controlled settings, focusing on their capabilities and limitations in single-session tasks.
Papers using MT-Bench (5)
- ASA: Backbone-Training-Free Representation Engineering for Tool-Calling AgentsEvaluating Agentic AI In The Wild: Failure Modes, Drift Patterns, And A Production Evaluation FrameworkMMoA: An AI-Agent framework with recurrence for Memoried Mixure-of-AgentCoEvol: Constructing Better Responses for Instruction Finetuning through Multi-Agent CooperationA System for Morphology-Task Generalization via Unified Representation
and Behavior Distillation