Awesome AI for Code
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
datasets
loadingβ¦
π€
Ask AI
Awesome datasets β curated papers, datasets & benchmarks Β· Awesome AI for Code
β all topics
overview
datasets
15 papers tagged datasets β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
15 papers Β· trending (default)
numbers = π₯ heat
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
(2026)
Hao Liang et al.
4.39
Pre-Hoc Predictions in AutoML: Leveraging LLMs to Enhance Model Selection and Benchmarking for Tabular datasets
(2025)
Yannis Belkhiter et al.
1.50
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
(2025)
Eunsu Kim et al.
1.28
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
(2025)
Guilherme Penedo et al.
1.28
Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
(2025)
Yueqi Song et al.
1.28
DB-BERT: a Database Tuning Tool that "Reads the Manual"
(2021)
Immanuel Trummer
β
Empower Large Language Model to Perform Better on Industrial Domain-Specific Question Answering
(2023)
Fangkai Yang et al.
β
FlexKBQA: A Flexible LLM-Powered Framework for Few-Shot Knowledge Base Question Answering
(2023)
Zhenyu Li et al.
β
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
(2024)
Guilherme Penedo et al.
β
Data-Prep-Kit: getting your data ready for LLM application development
(2024)
David Wood et al.
β
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
(2024)
Joongwon Kim et al.
β
LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
(2024)
Yushi Bai et al.
β
RoundTable: Leveraging Dynamic Schema and Contextual Autocomplete for Enhanced Query Precision in Tabular Question Answering
(2024)
Pratyush Kumar et al.
β
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
(2024)
Jun Shern Chan et al.
β
UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages
(2024)
Bethel Melesse Tessema et al.
β