Awesome Large Language Models
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Senthooran Rajamanoharan — most-cited papers & profile · Large Language Models
← authors
·
overview
Senthooran Rajamanoharan
6
papers ·
1
citations
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Emergent Misalignment Is Easy, Narrow Misalignment Is Hard
2026 · 1 citations
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
2025
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
2025
Eliciting Secret Knowledge from Language Models
2025
Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
2024
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
2024
Top co-authors
Neel Nanda
· 5
Arthur Conmy
· 2
Bartosz Cywiński
· 2
Emil Ryd
· 2
Samuel Marks
· 2
Adam Karvonen
· 1
Anna Soligo
· 1
Caden Juang
· 1
Edward Turner
· 1
Helena Casademunt
· 1
Janos Kramar
· 1
Javier Ferrando
· 1
Topics
Training Techniques
Safety & Alignment
Fine-Tuning
Evaluation
Code
Model Architecture
Survey Paper
Concept
Ablation
CAFT