Key papers π₯ Trending (default) π Most cited π Newest first π€ A β Z by title 60 papers Β· trending (default) numbers = π₯ heat
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models (2026) Yueyi Sun et al.
12.02 AutoMem: Automated Learning of Memory as a Cognitive Skill (2026) Shengguang Wu et al.
11.71 Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs (2026) Gabrielle Kaili-May Liu et al.
10.98 Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel (2026) Sietse Schelpe
9.96 Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE (2026) Haozhan Tang et al.
9.60 Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe (2026) Qian Zhao et al.
8.80 Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents (2026) Changdae Oh et al.
8.17 Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification (2026) Yunhao Feng et al.
7.98 GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity (2026) Yong Yi Bay et al.
7.37 HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment (2026) Shei Pern Chua et al.
7.37 CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression (2026) Morayo Danielle Adeyemi et al.
6.91 ReFreeKV: Towards Threshold-Free KV Cache Compression (2025) Xuanfan Ni et al.
6.60 AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation (2026) Bao Long Nguyen Huu et al.
5.88 Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction (2026) Mingzhe Du et al.
5.88 LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning (2026) Matteo Boglioni et al.
5.49 Bounded Morality: Defining the Space of Moral Computation (2026) Max Kanwal et al.
5.01 NeuroCogMap Reveals Cognitive Organization of Large Language Models (2026) Zhongxiang Sun et al.
5.01 Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval (2026) Ivan Ji (Zihao) et al.
5.01 Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization (2026) Xuefeng Liu et al.
5.01 Evaluating Chunking Strategies for Retrieval-Augmented Generation on Academic Texts (2026) Valentin J. J. Kreileder et al.
5.01 AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations (2026) Javier Irigoyen et al.
5.01 Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring (2026) William Hackett et al.
5.01 The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry (2026) Yuan Yuan
5.01 How Do I Know What to Say Next? Barenholtz's Autogenerative Theory as an Enrichment of Harrisean Integrationism (2026) J. Mark Bishop et al.
5.01 Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework (2026) Riccardo Revalor et al.
5.01 Two Axes of LLM Abstention: Answer Correctness and Question Answerability (2026) Benedikt J. Wagner
5.01 When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability (2026) Zongyou Yang et al.
5.01 Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs (2026) Shayda Moezzi et al.
5.01 Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility (2026) Nyx Iskandar
5.01 Towards Reliable AI-Assisted Analog Design: Template-Constrained LLM Agents for SAR ADC Generation (2026) Dimple Vijay Kochar et al.
5.01 OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios (2026) Chengyu Shen et al.
5.01 Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents (2026) Dylan Van Mulders et al.
5.01 Symbal: Detecting Systematic Misalignments in Model-Generated Captions (2026) Maya Varma et al.
5.01 Surprisal Theory is Tautological (without Rational Grounding) (2026) Ryan Cotterell
5.01 Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023 (2026) Sherzod Turaev et al.
4.95 IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources (2026) Arash Ghafouri et al.
4.95 Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software (2026) Arastoo Zibaeirad et al.
4.95 TokenMinds: Pretrained User Tokens and Embeddings for User Understanding in Large Recommender Systems (2026) Qingyun Liu et al.
4.95 Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training (2026) Jasmine Brazilek et al.
4.95 Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare (2026) Jasmine Brazilek et al.
4.95 DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models (2026) Yuxuan Yang et al.
4.95 Synthetic Feature Augmentation Improves Generalization Performance of
Language Models (2025) Ashok Choudhary et al.
4.71 BaRA: Budget-constrained and Reliable Web Data Collection Agent (2026) Soojeong Lee et al.
4.39 SchemaRAG: Dynamic Large Schema Reduction for LLM-driven Structured Information Extraction (2026) Sin Yu Bonnie Ho et al.
4.39 Libra: Training the Environment for Agentic Information Retrieval (2026) Xuan Zhao et al.
4.39 Learning User-Aware Recall: Personalized Retrieval in Long-Term Conversational Memory (2026) ZhiShu Jiang et al.
4.39 Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection (2026) Bo Chen
4.39 Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows (2026) Edward Y. Chang et al.
4.39 PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents (2026) Ke Zhang et al.
4.39 Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences (2026) Mark Russinovich et al.
4.39 Exploring the Semantic Gap in Agentic Data Systems: A Formative Study of Operationalization Failures in Analytical Workflows (2026) Jalal Mahmud et al.
4.39 Skills Are Not Islands: Measuring Dependency and Risk in Agent Skill Supply Chains (2026) Changguo Jia et al.
4.39 Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows (2026) Karthikeya Aditya Vissa et al.
4.39 Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL (2026) Juliette Decugis et al.
4.39 Revisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought Learning (2026) Hongyang He et al.
4.39 Safe and Adaptive Cloud Healing: Verifying LLM-Generated Recovery Plans with a Neural-Symbolic World Model (2026) Junyan Tan et al.
4.39 Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing (2026) Joshua Penman
4.39 Meta-Benchmarks for Financial-Services LLM Evaluation (2026) Blair Hudson
4.39 Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge (2026) Alex Brooker et al.
4.39 Atomic Task Graph: A Unified Framework for Agentic Planning and Execution (2026) Yue Zhang et al.
4.39