← all papers · overview

Assessing The Performance Of Human-capable Llms -- Are Llms Coming For Your Job?

Abstract

The current paper presents the development and validation of SelfScore, a novel benchmark designed to assess the performance of automated Large Language Model (LLM) agents on help desk and professional consultation tasks. Given the increasing integration of AI in industries, particularly within customer service, SelfScore fills a crucial gap by enabling the comparison of automated agents and human

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).