Awesome AI Agents
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β authors
Β·
overview
Loading authorβ¦
π€
Ask AI
Zhenhong Zhou β most-cited papers & profile Β· AI Agents
β authors
Β·
overview
Zhenhong Zhou
12
papers Β·
92
citations
Google Scholar β
Semantic Scholar β
OpenAlex β
Most-cited papers
How Alignment And Jailbreak Work: Explain LLM Safety Through Intermediate Hidden States
2024 Β· 92 citations
How Emotion Shapes the Behavior of LLMs and Agents: A Mechanistic Study
2026
Resource Consumption Threats in Large Language Models
2026
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
2025
Backdoor Collapse: Eliminating Unknown Threats via Known Backdoor Aggregation in Language Models
2025
Resource Consumption Red-Teaming for Large Vision-Language Models
2025
Resource Consumption Red-teaming For Large Vision-language Models
2025
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
2025
LIFEBench: Evaluating Length Instruction Following in Large Language Models
2025
Crabs: Consuming Resource Via Auto-generation For Llm-dos Attack Under Black-box Settings
2024
Course-Correction: Safety Alignment Using Synthetic Preferences
2024
On the Role of Attention Heads in Large Language Model Safety
2024
Topics
Safety & Alignment
Fine-Tuning
Model Architecture
cs.AI
cs.CL
LLM Security
cs.CR
Adversarial ML
Vulnerability Detection
Evaluation