← all papers · overview

How Secure Is Ai-generated Code: A Large-scale Comparison Of Large Language Models

Abstract

This study compares state-of-the-art Large Language Models (LLMs) on their tendency to generate vulnerabilities when writing C programs using a neutral zero-shot prompt. Tihanyi et al. introduced the FormAI dataset at PROMISE'23, featuring 112,000 C programs generated by GPT-3.5-turbo, with over 51.24% identified as vulnerable. We extended that research with a large-scale study involving 9 state-o

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).