OverThink: Slowdown Attacks on Reasoning LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Kumar, Abhinav, Roh, Jaechul, Naseh, Ali, Karpinska, Marzena, Iyyer, Mohit, Houmansadr, Amir, Bagdasarian, Eugene |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Throttling Web Agents Using Reasoning Gates
di: Kumar, Abhinav, et al.
Pubblicazione: (2025)
di: Kumar, Abhinav, et al.
Pubblicazione: (2025)
Backdooring Bias ($B^2$) into Stable Diffusion Models
di: Naseh, Ali, et al.
Pubblicazione: (2024)
di: Naseh, Ali, et al.
Pubblicazione: (2024)
OSLO: One-Shot Label-Only Membership Inference Attacks
di: Peng, Yuefeng, et al.
Pubblicazione: (2024)
di: Peng, Yuefeng, et al.
Pubblicazione: (2024)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
di: Naseh, Ali, et al.
Pubblicazione: (2025)
di: Naseh, Ali, et al.
Pubblicazione: (2025)
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
di: Roh, Jaechul, et al.
Pubblicazione: (2026)
di: Roh, Jaechul, et al.
Pubblicazione: (2026)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
di: Naseh, Ali, et al.
Pubblicazione: (2024)
di: Naseh, Ali, et al.
Pubblicazione: (2024)
Diffence: Fencing Membership Privacy With Diffusion Models
di: Peng, Yuefeng, et al.
Pubblicazione: (2023)
di: Peng, Yuefeng, et al.
Pubblicazione: (2023)
Multilingual and Multi-Accent Jailbreaking of Audio LLMs
di: Roh, Jaechul, et al.
Pubblicazione: (2025)
di: Roh, Jaechul, et al.
Pubblicazione: (2025)
PostMark: A Robust Blackbox Watermark for Large Language Models
di: Chang, Yapei, et al.
Pubblicazione: (2024)
di: Chang, Yapei, et al.
Pubblicazione: (2024)
Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security
di: Naseh, Ali, et al.
Pubblicazione: (2025)
di: Naseh, Ali, et al.
Pubblicazione: (2025)
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
di: Suri, Anshuman, et al.
Pubblicazione: (2025)
di: Suri, Anshuman, et al.
Pubblicazione: (2025)
Network-Level Prompt and Trait Leakage in Local Research Agents
di: Jeong, Hyejun, et al.
Pubblicazione: (2025)
di: Jeong, Hyejun, et al.
Pubblicazione: (2025)
RAIFLE: Reconstruction Attacks on Interaction-based Federated Learning with Adversarial Data Manipulation
di: Pham, Dzung, et al.
Pubblicazione: (2023)
di: Pham, Dzung, et al.
Pubblicazione: (2023)
FameBias: Embedding Manipulation Bias Attack in Text-to-Image Models
di: Roh, Jaechul, et al.
Pubblicazione: (2024)
di: Roh, Jaechul, et al.
Pubblicazione: (2024)
Identifying Models Behind Text-to-Image Leaderboards
di: Naseh, Ali, et al.
Pubblicazione: (2026)
di: Naseh, Ali, et al.
Pubblicazione: (2026)
Contextual Agent Security: A Policy for Every Purpose
di: Tsai, Lillian, et al.
Pubblicazione: (2025)
di: Tsai, Lillian, et al.
Pubblicazione: (2025)
Fake or Compromised? Making Sense of Malicious Clients in Federated Learning
di: Mozaffari, Hamid, et al.
Pubblicazione: (2024)
di: Mozaffari, Hamid, et al.
Pubblicazione: (2024)
Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation
di: Naseh, Ali, et al.
Pubblicazione: (2025)
di: Naseh, Ali, et al.
Pubblicazione: (2025)
Excessive Reasoning Attack on Reasoning LLMs
di: Si, Wai Man, et al.
Pubblicazione: (2025)
di: Si, Wai Man, et al.
Pubblicazione: (2025)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
di: Naseh, Ali, et al.
Pubblicazione: (2025)
di: Naseh, Ali, et al.
Pubblicazione: (2025)
Persistent Backdoor Attacks in Continual Learning
di: Guo, Zhen, et al.
Pubblicazione: (2024)
di: Guo, Zhen, et al.
Pubblicazione: (2024)
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
di: Roh, Jaechul, et al.
Pubblicazione: (2025)
di: Roh, Jaechul, et al.
Pubblicazione: (2025)
Can Large Language Models Really Recognize Your Name?
di: Pham, Dzung, et al.
Pubblicazione: (2025)
di: Pham, Dzung, et al.
Pubblicazione: (2025)
Self-interpreting Adversarial Images
di: Zhang, Tingwei, et al.
Pubblicazione: (2024)
di: Zhang, Tingwei, et al.
Pubblicazione: (2024)
VIDSTAMP: A Temporally-Aware Watermark for Ownership and Integrity in Video Diffusion Models
di: Teymoorianfard, Mohammadreza, et al.
Pubblicazione: (2025)
di: Teymoorianfard, Mohammadreza, et al.
Pubblicazione: (2025)
MeanSparse: Post-Training Robustness Enhancement Through Mean-Centered Feature Sparsification
di: Amini, Sajjad, et al.
Pubblicazione: (2024)
di: Amini, Sajjad, et al.
Pubblicazione: (2024)
FedSpy-LLM: Towards Scalable and Generalizable Data Reconstruction Attacks from Gradients on LLMs
di: Meerza, Syed Irfan Ali, et al.
Pubblicazione: (2026)
di: Meerza, Syed Irfan Ali, et al.
Pubblicazione: (2026)
Trusted Machine Learning Models Unlock Private Inference for Problems Currently Infeasible with Cryptography
di: Shumailov, Ilia, et al.
Pubblicazione: (2025)
di: Shumailov, Ilia, et al.
Pubblicazione: (2025)
Efficient Adversarial Training in LLMs with Continuous Attacks
di: Xhonneux, Sophie, et al.
Pubblicazione: (2024)
di: Xhonneux, Sophie, et al.
Pubblicazione: (2024)
Instruction Backdoor Attacks Against Customized LLMs
di: Zhang, Rui, et al.
Pubblicazione: (2024)
di: Zhang, Rui, et al.
Pubblicazione: (2024)
AirGapAgent: Protecting Privacy-Conscious Conversational Agents
di: Bagdasarian, Eugene, et al.
Pubblicazione: (2024)
di: Bagdasarian, Eugene, et al.
Pubblicazione: (2024)
Adaptive Backdoor Attacks with Reasonable Constraints on Graph Neural Networks
di: Dong, Xuewen, et al.
Pubblicazione: (2025)
di: Dong, Xuewen, et al.
Pubblicazione: (2025)
Taking off the Rose-Tinted Glasses: A Critical Look at Adversarial ML Through the Lens of Evasion Attacks
di: Eykholt, Kevin, et al.
Pubblicazione: (2024)
di: Eykholt, Kevin, et al.
Pubblicazione: (2024)
Membership Inference Attacks on Vision-Language-Action Models
di: Peng, Yuefeng, et al.
Pubblicazione: (2026)
di: Peng, Yuefeng, et al.
Pubblicazione: (2026)
Towards Automatic Hands-on-Keyboard Attack Detection Using LLMs in EDR Solutions
di: Portnoy, Amit, et al.
Pubblicazione: (2024)
di: Portnoy, Amit, et al.
Pubblicazione: (2024)
An Attack to Break Permutation-Based Private Third-Party Inference Schemes for LLMs
di: Thomas, Rahul, et al.
Pubblicazione: (2025)
di: Thomas, Rahul, et al.
Pubblicazione: (2025)
Ensembler: Protect Collaborative Inference Privacy from Model Inversion Attack via Selective Ensemble
di: Liu, Dancheng, et al.
Pubblicazione: (2024)
di: Liu, Dancheng, et al.
Pubblicazione: (2024)
Guided Reasoning in LLM-Driven Penetration Testing Using Structured Attack Trees
di: Nakano, Katsuaki, et al.
Pubblicazione: (2025)
di: Nakano, Katsuaki, et al.
Pubblicazione: (2025)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
di: Chaudhari, Harsh, et al.
Pubblicazione: (2026)
di: Chaudhari, Harsh, et al.
Pubblicazione: (2026)
Linkage Attacks Expose Identity Risks in Public ECG Data Sharing
di: Wang, Ziyu, et al.
Pubblicazione: (2025)
di: Wang, Ziyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Throttling Web Agents Using Reasoning Gates
di: Kumar, Abhinav, et al.
Pubblicazione: (2025) -
Backdooring Bias ($B^2$) into Stable Diffusion Models
di: Naseh, Ali, et al.
Pubblicazione: (2024) -
OSLO: One-Shot Label-Only Membership Inference Attacks
di: Peng, Yuefeng, et al.
Pubblicazione: (2024) -
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
di: Naseh, Ali, et al.
Pubblicazione: (2025) -
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
di: Roh, Jaechul, et al.
Pubblicazione: (2026)