Ο-Bench
Emerging9papers using it
2025first seen
The 'Ο-Bench' dataset is used to evaluate the behavioral similarity of tool-use habits among different language model agents by modeling their actions as directed graphs.
Papers using Ο-Bench (9)
- Goal Alignment in LLM-Based User Simulators for Conversational AIRemember When It Matters: Proactive Memory Agent for Long-Horizon AgentsScaling Agentic Capabilities via Grounded Interaction SynthesisUncertainty-Aware Clarification in LLM Agents with Information GainWhen Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use BehaviorsLiteResearcher: A Scalable Agentic RL Training Framework for Deep Research AgentReThinker: Scientific Reasoning by Rethinking with Guided Reflection and Confidence ControlScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent TrainingSearch More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization