Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
attacks
loadingβ¦
π€
Ask AI
Awesome attacks β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
attacks
17 papers tagged attacks β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
17 papers Β· trending (default)
numbers = π₯ heat
Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification
(2026)
Carter Luck et al.
5.01
Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models
(2026)
Malikeh Ehghaghi et al.
1.94
The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1
(2025)
Kaiwen Zhou et al.
1.28
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
(2025)
Salman Rahman et al.
1.28
The Aloe Family Recipe for Open and Specialized Healthcare LLMs
(2025)
Dario Garcia-Gasulla et al.
1.28
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
(2025)
Zihao Zhu et al.
1.28
Position: Privacy Is Not Just Memorization!
(2025)
Niloofar Mireshghallah et al.
1.28
Imperceptible Jailbreaking against Large Language Models
(2025)
Kuofeng Gao et al.
1.28
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
(2025)
Mikhail Terekhov et al.
1.28
Machine Text Detectors are Membership Inference Attacks
(2025)
Ryuto Koike et al.
1.28
OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation
(2025)
Xiaojun Jia et al.
1.28
CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare
(2025)
Akash Ghosh et al.
1.28
On the Reliability of Watermarks for Large Language Models
(2023)
John Kirchenbauer et al.
β
Jailbroken: How Does LLM Safety Training Fail?
(2023)
Alexander Wei et al.
β
Can LLMs Follow Simple Rules?
(2023)
Norman Mu et al.
β
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
(2024)
Liwei Jiang et al.
β
Hallucinating AI Hijacking Attack: Large Language Models and Malicious Code Recommenders
(2024)
David Noever et al.
β