SkillsBench
Emerging15papers using it
2026first seen
Warning: The leaderboard above is generated by Hugging Face eval-results and may be incomplete until evaluation_framework: benchflow is accepted and deployed. The audited SkillsBench v1.1 result archive is https://huggingface.co/datasets/benchflow/skillsbench-leaderboard, with all retained submissions normalized under
Papers using SkillsBench (15)
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and EvaluationGraph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent SkillsSkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill RevisionSkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at ScaleAIP: A Graph Representation for Learning and Governing Agent SkillsWhat Should a Skill Remember? Quality--Cost Trade-offs in Cost-Aware Skill Rewriting for Language Model AgentsSkillAxe: Sharpening LLM-Authored Agent Skills Through Evaluation-Guided Self-RefinementSkillJuror: Measuring How Agent Skill Organization Changes Runtime BehaviorSkillsInjector: Dynamic Skill Context Construction for LLM AgentsSkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM AgentsSkillMOO: Multi-Objective Optimization of Agent Skills for Software EngineeringCoevoskills: Self-evolving Agent Skills Via Co-evolutionary VerificationSkCC: Portable and Secure Skill Compilation for Cross-Framework LLM AgentsSkillSmith: Compiling Agent Skills into Boundary-Guided Runtime InterfacesClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation