Key papers π₯ Trending (default) π Most cited π Newest first π€ A β Z by title 60 papers Β· trending (default) numbers = π₯ heat
ISO: An RLVR-Native Optimization Stack (2026) Hanqing Zhu et al.
8.80 H$^2$SD: Hybrid Hindsight Self-Distillation (2026) Qiye Cai et al.
8.51 NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs (2026) Jiarong Zhao et al.
7.90 A Controlled Study of Attention-Only Transformers (2026) Henry Ndubuaku et al.
5.90 Measuring Reward-Seeking via Contrastive Belief Updates (2026) Axel H{\o}jmark et al.
4.95 How Much of the Routing Gap Is Real? Decomposing the Router-to-Oracle Gap into Reproducible Specialist Advantage and Single-Draw Label Noise (2026) Teng-Ruei Chen
4.33 Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages (2026) Lucas Pinto
4.33 TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation (2026) Andrii Balashov et al.
4.33 HiFuzz: Hierarchical Reinforcement Learning for Semantic-Aware and Adaptive CPU Fuzzing (2026) Ya Wang et al.
4.33 Open-Ended Scenario Reasoning for Specialist Model Adaptation (2026) Youcheng Zong et al.
4.33 At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics (2026) Truong Xuan Khanh
4.33 The Rank-One Corner: How Much Value Equivalence Does a Task Need from a World Model? (2026) Donna Vakalis
4.33 Can a Language Model Learn Facts Continually in Its Weights? (2026) Charles O'Neill
4.33 Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans (2026) Jasmin Thelen et al.
4.33 AAAI-26 Dual Submissions: Novel Challenges (2026) Kiri L. Wagstaff et al.
4.33 CARE-LoRA: Compressed Activation REconstruction for Memory-Efficient LoRA (2026) Gengyu Zhang et al.
4.33 When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary (2026) Jim Allchin
4.33 Constructed Reality, Contested Priors: Decoupling and the Architecture of Cognitive Relapse Under the Free Energy Principle (2026) MD Ibrahim Hossain Ridoy
4.33 Removable Defects: The Economics and Limits of Deliberate Deficiency (2026) Cheng Qian
4.33 Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs (2026) Nikita Kozodoi et al.
4.33 Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning (2026) Jing Liu et al.
4.33 Forgetful Attention: A Trainable Support-Vector Memory with Certified Selection and Exact Unlearning (2026) Vishwajith Ramesh
4.33 Speculate with Memory: Lossless Acceleration for LLM Agents (2026) Yu Li et al.
4.33 The Sound of Absence: Audio-Language Embedding Models Struggle with Negation (2026) Chun-Yi Kuan et al.
4.33 Extractable Memorization From First Principles (2026) A. Feder Cooper et al.
4.33 What Makes a Representational Prior Work? Feature Families, Label-Free Invariances, and Critical Windows in Grokking (2026) Gunner Levi Howe
4.33 Quantifying Diversity of Thought: A Predictive Law of Weighted LLM Ensemble Lift (2026) Junade Ali
4.33 Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows (2026) Jinyuan Deng et al.
4.33 Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models (2026) Chao Han et al.
4.33 One Student, Many Teachers: Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context (2026) Yingzi Ma et al.
4.33 On the Limits of Support-Preserving Alignment and Bounded Filtering (2026) Aryan Dutt et al.
4.33 A Better Start for Language Models: Domain-Conditional Position Offsets (2026) Ye Qiao
4.33 TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue (2026) Shuzhong Lai et al.
4.33 The Information Shadow: Measuring Structural Limits on What Language Models Can Learn (2026) Priyansh Srivastava et al.
4.33 Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative (2026) David G\'omez-Guill\'en et al.
4.33 Estimating Rare Events in Language Models with Proper Evaluation (2026) Nikita Y. Parulekar et al.
4.33 Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary (2026) Jan Kirin
4.33 Attacking Graph Foundation Models Through Their Shared Representation (2026) Pankaj Kumar et al.
4.33 BRIDGE: Bottleneck-Aware Regulator-Set Inference and Diagnosis for Cooperative Gene Regulatory Recovery (2026) Maryam Rahimimovassagh et al.
4.33 Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs (2026) Seunghyun Lee et al.
4.33 Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset Generator (2026) Yuxiang Ji
4.33 Exposure-Based Reinforcement Learning to Rank (2026) Harrie Oosterhuis et al.
4.33 Relative Positions Generalize, Absolute Positions Memorize: An Implicit-Bias Account of Length Generalization in Attention (2026) Subham Singh et al.
4.33 PertReason: A Knowledge-Grounded Benchmark and Framework for Cell-State-Conditioned Mechanistic Reasoning of Perturbation Effects (2026) Dongkwan Kim et al.
4.33 Elicitation without Backpropagation: Steering Model Behavior by Optimizing the Latent Posterior (2026) Garrett Baker et al.
4.33 From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning (2026) Garvit Singla et al.
4.33 ABOPD: Antibody CDR Design via On-Policy Distillation (2026) Zhuo Yang et al.
4.33 HindsightBench: A Black-Box Behavioral Audit Protocol for Parametric Hindsight in Time-Indexed LLM Decision Tasks (2026) Haozhe Jia
4.33 Breaking Feedback-Blindness: Utility-Augmented Transformer for Sequential Decision Making (2026) Yuyang Shen et al.
4.33 Circuit Claims Depend on What Is Extracted and How It Is Compared (2026) Yang Sheng et al.
4.33 Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation (2026) Mingxuan Ouyang et al.
4.33 Parallel Noising in Neural Markov Logic Networks (2026) Peter Jung et al.
4.33 AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters (2026) Yu-Yang Qian et al.
4.33 Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks (2026) Guy Stephane Waffo Dzuyo (Forvis Mazars et al.
4.33 Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information (2026) Priyank Agrawal et al.
4.33 Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code (2026) Diego Salda\~na Ulloa
4.33 ESC: Emotional Self-Correction for Reliable Vision-Language Models (2026) Tien-Huy Nguyen et al.
3.45 Rank-Order N-of-M Codes for Sparse Distributed Memory: Disentangling Representation and Learning Effects in Noise Robustness Against Contemporary Neuromorphic Architectures (2026) Joy Bose
3.45 ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy (2026) Zijun Xie et al.
3.45 To Retain or to Adapt? Generalizing Continual Learning (2026) Giulia Lanzillotta et al.
3.45