Awesome AI Agents
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β authors
Β·
overview
Loading authorβ¦
π€
Ask AI
Xingyuan Bu β most-cited papers & profile Β· AI Agents
β authors
Β·
overview
Xingyuan Bu
6
papers Β·
14
citations
Google Scholar β
Semantic Scholar β
OpenAlex β
Most-cited papers
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
2024 Β· 12 citations
TEGEE: Task dEfinition Guided Expert Ensembling for Generalizable and Few-shot Learning
2024 Β· 2 citations
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
2025
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
2025
Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level
2024
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
2024
Topics
models
cs.CL
Evaluation
Code Models
Bug Detection
Testing
Software Engineering
Chain-of-Thought
Large
Language