← all papers · overview

Benchmarking LLM Powered Chatbots: Methods And Metrics

Abstract

Autonomous conversational agents, i.e. chatbots, are becoming an increasingly common mechanism for enterprises to provide support to customers and partners. In order to rate chatbots, especially ones powered by Generative AI tools like Large Language Models (LLMs) we need to be able to accurately assess their performance. This is where chatbot benchmarking becomes important. In this paper, we prop

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).