Awesome AI for Code
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
training
loadingβ¦
π€
Ask AI
Awesome training β curated papers, datasets & benchmarks Β· Awesome AI for Code
β all topics
overview
training
24 papers tagged training β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
24 papers Β· trending (default)
numbers = π₯ heat
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
(2025)
Yujia Qin et al.
5.13
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
(2025)
GLM-4. 5 Team et al.
5.10
OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale
(2026)
Jingze Shi et al.
3.23
Union of Experts: Adapting Hierarchical Routing to Equivalently Decomposed Transformer
(2025)
Yujiao Yang et al.
2.87
Fast Data Aware Neural Architecture Search via Supernet Accelerated Evaluation
(2025)
Emil Njor et al.
2.65
Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning
(2026)
Shijie Li et al.
2.00
Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model
(2026)
Xinyin Ma et al.
2.00
RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation
(2026)
Haoyu Zhao et al.
2.00
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing
(2026)
Tommaso Cerruti et al.
2.00
Xmodel-2.5: 1.3B Data-Efficient Reasoning SLM
(2025)
Yang Liu et al.
1.56
Benchmarking Optimizers for Large Language Model Pretraining
(2025)
Andrei Semenov et al.
1.44
CMoE: Fast Carving of Mixture-of-Experts for Efficient LLM Inference
(2025)
Zehua Pei et al.
1.28
A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
(2025)
Jitai Hao et al.
1.28
Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
(2025)
Yueqi Song et al.
1.28
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
(2025)
Taishi Nakamura et al.
1.06
Joey NMT: A Minimalist NMT Toolkit for Novices
(2019)
Julia Kreutzer et al.
β
Knowledge Graph Embedding with Atrous Convolution and Residual Learning
(2020)
Feiliang Ren et al.
β
MetaMIML: Meta Multi-Instance Multi-Label Learning
(2021)
Yuanlin Yang et al.
β
Multi-Head Deep Metric Learning Using Global and Local Representations
(2021)
Mohammad K. Ebrahimpour et al.
β
Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
(2022)
Haokun Liu et al.
β
Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-Experts
(2024)
Tong Zhu et al.
β
Approximating Two-Layer Feedforward Networks for Efficient Transformers
(2023)
RΓ³bert CsordΓ‘s et al.
β
Horizon-Length Prediction: Advancing Fill-in-the-Middle Capabilities for Code Generation with Lookahead Planning
(2024)
Yifeng Ding et al.
β
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
(2024)
Jun Shern Chan et al.
β