← all papers · overview

CIDR: A Large-Scale Industrial Source Code Dataset for Software Engineering Research

Abstract

We present the Curated Industrial Developer Repository (CIDR), a large-scale dataset of real-world software repositories collected from industrial partners. The dataset comprises 4,225 repositories spanning 75 programming languages, totaling 832 million raw lines of code (581 million logical lines), along with structured metadata at the repository level, full version control history, and engineering-practice attributes such as continuous integration usage and the presence of automated tests. All repositories were collected, filtered, and anonymized through a multi-stage pipeline developed specifically for this purpose. We additionally report an exploratory fine-tuning study that adapts a 3-billion-parameter code language model to CIDR and quantifies the effect on held-out enterprise code. CIDR is intended to support research in code intelligence, software quality analysis, developer tooling, and related software engineering tasks. Access to CIDR is provided under a restricted license; details on eligibility and terms are available at https://fermatix.ai/#Contact.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).