← all papers · overview

Chipbench: A Next-step Benchmark For Evaluating LLM Performance In Ai-aided Chip Design

Abstract

While Large Language Models (LLMs) show significant potential in hardware engineering, current benchmarks suffer from saturation and limited task diversity, failing to reflect LLMs' performance in real industrial workflows. To address this gap, we propose a comprehensive benchmark for AI-aided chip design that rigorously evaluates LLMs across three critical tasks: Verilog generation, debugging, an

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).