Key papers π₯ Trending (default) π Most cited π Newest first π€ A β Z by title 60 papers Β· trending (default) numbers = π₯ heat
MentalThink: Shaping Thoughts in Mental SVG World (2026) Kangheng Lin et al.
2.00 DSAEval: Evaluating Data Science Agents on a Wide Range of Real-World Data Science Problems (2026) Maojun Sun et al.
1.94 Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR (2026) Fanfan Liu et al.
1.94 ContextBench: A Benchmark for Context Retrieval in Coding Agents (2026) Han Li et al.
1.94 SenTSR-Bench: Thinking with Injected Knowledge for Time-Series Reasoning (2026) Zelin He et al.
1.94 PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research (2026) Tingjia Miao et al.
1.94 PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination (2026) Qiyao Wang et al.
1.94 PreScam: A Benchmark for Predicting Scam Progression from Early Conversations (2026) Weixiang Sun et al.
1.94 MINTEval: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems (2026) Hyunji Lee et al.
1.94 K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts (2026) Nahyun Lee et al.
1.94 Benchmark Everything Everywhere All at Once (2026) Shiyun Xiong et al.
1.94 EmpiriGraph-Psy: A Dataset and LLM Pipeline for Extracting Empirical Relation Graphs from Psychology Abstracts (2026) Danqin Zhao et al.
1.94 Understanding the Behaviors of Environment-aware Information Retrieval (2026) Ruifeng Yuan et al.
1.94 Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs (2026) Fahd Seddik et al.
1.94 ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration (2026) Ruofeng Yang et al.
1.89 Self-Execution Simulation Improves Coding Models (2026) Gallil Maimon et al.
1.83 Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers (2026) Bozhou Li et al.
1.72 Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs (2026) Yining Hong et al.
1.72 Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks (2026) Qihua Dong et al.
1.72 X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests (2026) Jie Wu et al.
1.67 Dispider: Enabling Video LLMs with Active Real-Time Interaction via
Disentangled Perception, Decision, and Reaction (2025) Rui Qian et al.
1.28 WebWalker: Benchmarking LLMs in Web Traversal (2025) Jialong Wu et al.
1.28 Control LLM: Controlled Evolution for Intelligence Retention in LLM (2025) Haichao Wei et al.
1.28 O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning (2025) Haotian Luo et al.
1.28 PlotGen: Multi-Agent LLM-based Scientific Data Visualization via
Multimodal Feedback (2025) Kanika Goswami et al.
1.28 Competitive Programming with Large Reasoning Models (2025) OpenAI et al.
1.28 V2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with
Multi-Modal Large Language Models (2025) Hsu-kuang Chiu et al.
1.28 The Mirage of Model Editing: Revisiting Evaluation in the Wild (2025) Wanli Yang et al.
1.28 FLAG-Trader: Fusion LLM-Agent with Gradient-based Reinforcement Learning
for Financial Trading (2025) Guojun Xiong et al.
1.28 How to Steer LLM Latents for Hallucination Detection? (2025) Seongheon Park et al.
1.28 IterPref: Focal Preference Learning for Code Generation via Iterative
Debugging (2025) Jie Wu et al.
1.28 Using Mechanistic Interpretability to Craft Adversarial Attacks against
Large Language Models (2025) Thomas Winninger et al.
1.28 Should VLMs be Pre-trained with Image Data? (2025) Sedrick Keh et al.
1.28 Communication-Efficient Language Model Training Scales Reliably and
Robustly: Scaling Laws for DiLoCo (2025) Zachary Charles et al.
1.28 DAPO: An Open-Source LLM Reinforcement Learning System at Scale (2025) Qiying Yu et al.
1.28 Judge Anything: MLLM as a Judge Across Any Modality (2025) Shu Pu et al.
1.28 Efficient Model Selection for Time Series Forecasting via LLMs (2025) Wang Wei et al.
1.28 Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections
of Images (2025) Boyang Deng et al.
1.28 SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction
Fine-Tuning (2025) Prabhat Pandey et al.
1.28 NodeRAG: Structuring Graph-based RAG with Heterogeneous Nodes (2025) Tianyang Xu et al.
1.28 BookWorld: From Novels to Interactive Agent Societies for Creative Story
Generation (2025) Yiting Ran et al.
1.28 PHYBench: Holistic Evaluation of Physical Perception and Reasoning in
Large Language Models (2025) Shi Qiu et al.
1.28 Decoding Open-Ended Information Seeking Goals from Eye Movements in
Reading (2025) Cfir Avraham Hadar et al.
1.28 The Aloe Family Recipe for Open and Specialized Healthcare LLMs (2025) Dario Garcia-Gasulla et al.
1.28 Perception, Reason, Think, and Plan: A Survey on Large Multimodal
Reasoning Models (2025) Yunxin Li et al.
1.28 EfficientLLM: Efficiency in Large Language Models (2025) Zhengqing Yuan et al.
1.28 Code Graph Model (CGM): A Graph-Integrated Large Language Model for
Repository-Level Software Engineering Tasks (2025) Hongyuan Tao et al.
1.28 Fostering Video Reasoning via Next-Event Prediction (2025) Haonan Wang et al.
1.28 The Entropy Mechanism of Reinforcement Learning for Reasoning Language
Models (2025) Ganqu Cui et al.
1.28 DINGO: Constrained Inference for Diffusion LLMs (2025) Tarun Suresh et al.
1.28 SiLVR: A Simple Language-based Video Reasoning Framework (2025) Ce Zhang et al.
1.28 LIFT the Veil for the Truth: Principal Weights Emerge after Rank
Reduction for Reasoning-Focused Supervised Fine-Tuning (2025) Zihang Liu et al.
1.28 Unleashing the Reasoning Potential of Pre-trained LLMs by Critique
Fine-Tuning on One Problem (2025) Yubo Wang et al.
1.28 TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management
in LLM-based Agentic Multi-Agent Systems (2025) Shaina Raza et al.
1.28 MIRIAD: Augmenting LLMs with millions of medical query-response pairs (2025) Qinyue Zheng et al.
1.28 VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos (2025) Jiashuo Yu et al.
1.28 Profiling News Media for Factuality and Bias Using LLMs and the
Fact-Checking Methodology of Human Experts (2025) Zain Muhammad Mujahid et al.
1.28 RExBench: Can coding agents autonomously implement AI research
extensions? (2025) Nicholas Edwards et al.
1.28 Evaluating, Synthesizing, and Enhancing for Customer Support
Conversation (2025) Jie Zhu et al.
1.28 Select to Know: An Internal-External Knowledge Self-Selection Framework
for Domain-Specific Question Answering (2025) Bolei He et al.
1.28