Step-by-Step Reasoning Attack: Revealing 'Erased' Knowledge in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Sinha, Yash, Baser, Manit, Mandal, Murari, Divakaran, Dinil Mon, Kankanhalli, Mohan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Large Language Models for Phishing Webpage Detection and Identification
by: Lee, Jehyun, et al.
Published: (2024)
by: Lee, Jehyun, et al.
Published: (2024)
AI-based Traffic Modeling for Network Security and Privacy: Challenges Ahead
by: Divakaran, Dinil Mon
Published: (2025)
by: Divakaran, Dinil Mon
Published: (2025)
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
by: Nawal, Aditya, et al.
Published: (2026)
by: Nawal, Aditya, et al.
Published: (2026)
MATRA: Modeling the Attack Surface of Agentic AI Systems -- OpenClaw Case Study
by: Van hamme, Tim, et al.
Published: (2026)
by: Van hamme, Tim, et al.
Published: (2026)
ThinkEval: Practical Evaluation of Knowledge Leakage in LLM Editing using Thought-based Knowledge Graphs
by: Baser, Manit, et al.
Published: (2025)
by: Baser, Manit, et al.
Published: (2025)
From ML to LLM: Evaluating the Robustness of Phishing Webpage Detection Models against Adversarial Attacks
by: Kulkarni, Aditya, et al.
Published: (2024)
by: Kulkarni, Aditya, et al.
Published: (2024)
UniNet: A Unified Multi-granular Traffic Modeling Framework for Network Security
by: Wu, Binghui, et al.
Published: (2025)
by: Wu, Binghui, et al.
Published: (2025)
Involuntary Jailbreak: On Self-Prompting Attacks
by: Guo, Yangyang, et al.
Published: (2025)
by: Guo, Yangyang, et al.
Published: (2025)
RECTor: Robust and Efficient Correlation Attack on Tor
by: Wu, Binghui, et al.
Published: (2025)
by: Wu, Binghui, et al.
Published: (2025)
LLMs for Cyber Security: New Opportunities
by: Divakaran, Dinil Mon, et al.
Published: (2024)
by: Divakaran, Dinil Mon, et al.
Published: (2024)
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
by: Sinha, Yash, et al.
Published: (2024)
by: Sinha, Yash, et al.
Published: (2024)
Mitigating Bias in Machine Learning Models for Phishing Webpage Detection
by: Kulkarni, Aditya, et al.
Published: (2024)
by: Kulkarni, Aditya, et al.
Published: (2024)
ZEST: Attention-based Zero-Shot Learning for Unseen IoT Device Classification
by: Wu, Binghui, et al.
Published: (2023)
by: Wu, Binghui, et al.
Published: (2023)
Privacy Risks and Preservation Methods in Explainable Artificial Intelligence: A Scoping Review
by: Allana, Sonal, et al.
Published: (2025)
by: Allana, Sonal, et al.
Published: (2025)
When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
by: Sahoo, Devanshu, et al.
Published: (2025)
by: Sahoo, Devanshu, et al.
Published: (2025)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
by: Wang, Libo
Published: (2024)
by: Wang, Libo
Published: (2024)
AttacKG+:Boosting Attack Knowledge Graph Construction with Large Language Models
by: Zhang, Yongheng, et al.
Published: (2024)
by: Zhang, Yongheng, et al.
Published: (2024)
CLaRE-ty Amid Chaos: Quantifying Representational Entanglement to Predict Ripple Effects in LLM Editing
by: Baser, Manit, et al.
Published: (2026)
by: Baser, Manit, et al.
Published: (2026)
BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models
by: Liu, Shuaitong, et al.
Published: (2025)
by: Liu, Shuaitong, et al.
Published: (2025)
Conflicts Make Large Reasoning Models Vulnerable to Attacks
by: Liu, Honghao, et al.
Published: (2026)
by: Liu, Honghao, et al.
Published: (2026)
EagleEye: Attention to Unveil Malicious Event Sequences from Provenance Graphs
by: Gysel, Philipp, et al.
Published: (2024)
by: Gysel, Philipp, et al.
Published: (2024)
FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
by: Gong, Yichen, et al.
Published: (2023)
by: Gong, Yichen, et al.
Published: (2023)
A Survey of Attacks on Large Language Models
by: Xu, Wenrui, et al.
Published: (2025)
by: Xu, Wenrui, et al.
Published: (2025)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models
by: Truong, Vu Tuan, et al.
Published: (2026)
by: Truong, Vu Tuan, et al.
Published: (2026)
Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models
by: Wang, Shuqiang, et al.
Published: (2026)
by: Wang, Shuqiang, et al.
Published: (2026)
ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models
by: Liu, Xiaogeng, et al.
Published: (2026)
by: Liu, Xiaogeng, et al.
Published: (2026)
Membership Inference Attacks on Tokenizers of Large Language Models
by: Tong, Meng, et al.
Published: (2025)
by: Tong, Meng, et al.
Published: (2025)
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
by: Li, Jie, et al.
Published: (2024)
by: Li, Jie, et al.
Published: (2024)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
by: Park, Junyoung, et al.
Published: (2026)
by: Park, Junyoung, et al.
Published: (2026)
Recent Advances in Attack and Defense Approaches of Large Language Models
by: Cui, Jing, et al.
Published: (2024)
by: Cui, Jing, et al.
Published: (2024)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
by: Li, Xi, et al.
Published: (2024)
by: Li, Xi, et al.
Published: (2024)
TensorCommitments: A Lightweight Verifiable Inference for Language Models
by: Baser, Oguzhan, et al.
Published: (2026)
by: Baser, Oguzhan, et al.
Published: (2026)
Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey
by: Jain, Bhavuk, et al.
Published: (2026)
by: Jain, Bhavuk, et al.
Published: (2026)
SoK: Robustness in Large Language Models against Jailbreak Attacks
by: Xu, Feiyue, et al.
Published: (2026)
by: Xu, Feiyue, et al.
Published: (2026)
Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models
by: Das, Badhan Chandra, et al.
Published: (2026)
by: Das, Badhan Chandra, et al.
Published: (2026)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
by: Luo, Weidi, et al.
Published: (2025)
by: Luo, Weidi, et al.
Published: (2025)
Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
by: Hartman, Max, et al.
Published: (2026)
by: Hartman, Max, et al.
Published: (2026)
Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models
by: Yang, Fan
Published: (2025)
by: Yang, Fan
Published: (2025)
Similar Items
-
Multimodal Large Language Models for Phishing Webpage Detection and Identification
by: Lee, Jehyun, et al.
Published: (2024) -
AI-based Traffic Modeling for Network Security and Privacy: Challenges Ahead
by: Divakaran, Dinil Mon
Published: (2025) -
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
by: Nawal, Aditya, et al.
Published: (2026) -
MATRA: Modeling the Attack Surface of Agentic AI Systems -- OpenClaw Case Study
by: Van hamme, Tim, et al.
Published: (2026) -
ThinkEval: Practical Evaluation of Knowledge Leakage in LLM Editing using Thought-based Knowledge Graphs
by: Baser, Manit, et al.
Published: (2025)