← all papers · overview

Generating Unseen Code Tests In Infinitum

Abstract

Large Language Models (LLMs) are used for many tasks, including those related to coding. An important aspect of being able to utilize LLMs is the ability to assess their fitness for specific usages. The common practice is to evaluate LLMs against a set of benchmarks. While benchmarks provide a sound foundation for evaluation and comparison of alternatives, they suffer from the well-known weakness

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).