How Catastrophic is Your LLM? Certifying Risk in Conversation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Chengxiao, Chaudhary, Isha, Hu, Qian, Ruan, Weitong, Gupta, Rahul, Singh, Gagandeep |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
by: Vega, Jason, et al.
Published: (2023)
by: Vega, Jason, et al.
Published: (2023)
Quantitative Certification of Agentic Tool Selection
by: Yeon, Jehyeok, et al.
Published: (2025)
by: Yeon, Jehyeok, et al.
Published: (2025)
Certifying Counterfactual Bias in LLMs
by: Chaudhary, Isha, et al.
Published: (2024)
by: Chaudhary, Isha, et al.
Published: (2024)
Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning
by: Xu, Yinglun, et al.
Published: (2024)
by: Xu, Yinglun, et al.
Published: (2024)
Cross-Input Certified Training for Universal Perturbations
by: Xu, Changming, et al.
Published: (2024)
by: Xu, Changming, et al.
Published: (2024)
Towards Generalized Certified Robustness with Multi-Norm Training
by: Jiang, Enyi, et al.
Published: (2024)
by: Jiang, Enyi, et al.
Published: (2024)
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
by: Nikolić, Kristina, et al.
Published: (2025)
by: Nikolić, Kristina, et al.
Published: (2025)
Certifying Knowledge Comprehension in LLMs
by: Chaudhary, Isha, et al.
Published: (2024)
by: Chaudhary, Isha, et al.
Published: (2024)
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
by: Zhao, Mengnan, et al.
Published: (2026)
by: Zhao, Mengnan, et al.
Published: (2026)
BadSampler: Harnessing the Power of Catastrophic Forgetting to Poison Byzantine-robust Federated Learning
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
Certifying LLM Safety against Adversarial Prompting
by: Kumar, Aounon, et al.
Published: (2023)
by: Kumar, Aounon, et al.
Published: (2023)
Enhancing Continual Learning for Software Vulnerability Prediction: Addressing Catastrophic Forgetting via Hybrid-Confidence-Aware Selective Replay for Temporal LLM Fine-Tuning
by: Dou, Xuhui, et al.
Published: (2026)
by: Dou, Xuhui, et al.
Published: (2026)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
by: Anurin, Andrey, et al.
Published: (2024)
by: Anurin, Andrey, et al.
Published: (2024)
PRUNE: A Patching Based Repair Framework for Certifiable Unlearning of Neural Networks
by: Li, Xuran, et al.
Published: (2025)
by: Li, Xuran, et al.
Published: (2025)
FairProof : Confidential and Certifiable Fairness for Neural Networks
by: Yadav, Chhavi, et al.
Published: (2024)
by: Yadav, Chhavi, et al.
Published: (2024)
How Not to Detect Prompt Injections with an LLM
by: Choudhary, Sarthak, et al.
Published: (2025)
by: Choudhary, Sarthak, et al.
Published: (2025)
How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy
by: Ponomareva, Natalia, et al.
Published: (2025)
by: Ponomareva, Natalia, et al.
Published: (2025)
Robust Privacy: Inference-Time Privacy through Certified Robustness
by: Jin, Jiankai, et al.
Published: (2026)
by: Jin, Jiankai, et al.
Published: (2026)
How to Backdoor the Knowledge Distillation
by: Wu, Chen, et al.
Published: (2025)
by: Wu, Chen, et al.
Published: (2025)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
Your Agent Can Defend Itself against Backdoor Attacks
by: Changjiang, Li, et al.
Published: (2025)
by: Changjiang, Li, et al.
Published: (2025)
AuditVotes: A Framework Towards More Deployable Certified Robustness for Graph Neural Networks
by: Lai, Yuni, et al.
Published: (2025)
by: Lai, Yuni, et al.
Published: (2025)
Q-AGNN: Quantum-Enhanced Attentive Graph Neural Network for Intrusion Detection
by: Chaudhary, Devashish, et al.
Published: (2026)
by: Chaudhary, Devashish, et al.
Published: (2026)
Private-RAG: Answering Multiple Queries with LLMs while Keeping Your Data Private
by: Wu, Ruihan, et al.
Published: (2025)
by: Wu, Ruihan, et al.
Published: (2025)
Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation
by: Zafar, Osama, et al.
Published: (2026)
by: Zafar, Osama, et al.
Published: (2026)
Mark Your LLM: Detecting the Misuse of Open-Source Large Language Models via Watermarking
by: Xu, Yijie, et al.
Published: (2025)
by: Xu, Yijie, et al.
Published: (2025)
One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
by: Zhang, Mohan, et al.
Published: (2025)
by: Zhang, Mohan, et al.
Published: (2025)
Formal Synthesis of Certifiably Robust Neural Lyapunov-Barrier Certificates
by: Wang, Chengxiao, et al.
Published: (2026)
by: Wang, Chengxiao, et al.
Published: (2026)
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
by: Luo, Zeren, et al.
Published: (2025)
by: Luo, Zeren, et al.
Published: (2025)
How Private is Your Attention? Bridging Privacy with In-Context Learning
by: Bonnerjee, Soham, et al.
Published: (2025)
by: Bonnerjee, Soham, et al.
Published: (2025)
Uncertainty-Aware Hardware Trojan Detection Using Multimodal Deep Learning
by: Vishwakarma, Rahul, et al.
Published: (2024)
by: Vishwakarma, Rahul, et al.
Published: (2024)
Protect Your Score: Contact Tracing With Differential Privacy Guarantees
by: Romijnders, Rob, et al.
Published: (2023)
by: Romijnders, Rob, et al.
Published: (2023)
ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
by: Liang, Jiacheng, et al.
Published: (2026)
by: Liang, Jiacheng, et al.
Published: (2026)
Your Privacy Depends on Others: Collusion Vulnerabilities in Individual Differential Privacy
by: Kaiser, Johannes, et al.
Published: (2026)
by: Kaiser, Johannes, et al.
Published: (2026)
Matching Ranks Over Probability Yields Truly Deep Safety Alignment
by: Vega, Jason, et al.
Published: (2025)
by: Vega, Jason, et al.
Published: (2025)
How to Craft Backdoors with Unlabeled Data Alone?
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Is The Watermarking Of LLM-Generated Code Robust?
by: Suresh, Tarun, et al.
Published: (2024)
by: Suresh, Tarun, et al.
Published: (2024)
Do Phone-Use Agents Respect Your Privacy?
by: Tang, Zhengyang, et al.
Published: (2026)
by: Tang, Zhengyang, et al.
Published: (2026)
Improving Your Model Ranking on Chatbot Arena by Vote Rigging
by: Min, Rui, et al.
Published: (2025)
by: Min, Rui, et al.
Published: (2025)
Similar Items
-
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
by: Vega, Jason, et al.
Published: (2023) -
Quantitative Certification of Agentic Tool Selection
by: Yeon, Jehyeok, et al.
Published: (2025) -
Certifying Counterfactual Bias in LLMs
by: Chaudhary, Isha, et al.
Published: (2024) -
Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning
by: Xu, Yinglun, et al.
Published: (2024) -
Cross-Input Certified Training for Universal Perturbations
by: Xu, Changming, et al.
Published: (2024)