Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
GPT-4
loadingβ¦
π€
Ask AI
Awesome GPT-4 β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
GPT-4
20 papers tagged GPT-4 β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
20 papers Β· trending (default)
numbers = π₯ heat
Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1
(2026)
Sarah Y. Li et al.
5.01
Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate
(2025)
Yubo Wang et al.
1.28
Understanding Gen Alpha Digital Language: Evaluation of LLM Safety Systems for Content Moderation
(2025)
Manisha Mehta et al.
1.28
German4All - A Dataset and Model for Readability-Controlled Paraphrasing in German
(2025)
Miriam AnschΓΌtz et al.
1.28
From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence
(2025)
Jian Yang et al.
1.28
From FLOPs to Footprints: The Resource Cost of Artificial Intelligence
(2025)
Sophia Falk et al.
1.28
arXiVeri: Automatic table verification with GPT
(2023)
Gyungin Shin et al.
β
Scaling Clinical Trial Matching Using Large Language Models: A Case Study in Oncology
(2023)
Cliff Wong et al.
β
Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification
(2023)
Aojun Zhou et al.
β
GPT Can Solve Mathematical Problems Without a Calculator
(2023)
Zhen Yang et al.
β
From Sparse to Dense: GPT-4 Summarization with Chain of Density Prompting
(2023)
Griffin Adams et al.
β
Holodeck: Language Guided Generation of 3D Embodied AI Environments
(2023)
Yue Yang et al.
β
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
(2024)
Raghav Kapoor et al.
β
LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report
(2024)
Justin Zhao et al.
β
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
(2024)
Xinyu Fang et al.
β
APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
(2024)
Zuxin Liu et al.
β
Integrating Large Language Models into a Tri-Modal Architecture for Automated Depression Classification
(2024)
Santosh V. Patapati
β
PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
(2024)
Yijia Shao et al.
β
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
(2024)
Sara Ghaboura et al.
β
MATATA: a weak-supervised MAthematical Tool-Assisted reasoning for Tabular Applications
(2024)
Vishnou Vinayagame et al.
β