Towards Building a Robust Toxicity Predictor
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bespalov, Dmitriy, Bhabesh, Sourav, Xiang, Yi, Zhou, Liutong, Qi, Yanjun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TaeBench: Improving Quality of Toxic Adversarial Examples
von: Zhu, Xuan, et al.
Veröffentlicht: (2024)
von: Zhu, Xuan, et al.
Veröffentlicht: (2024)
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice
von: Goel, Aman, et al.
Veröffentlicht: (2025)
von: Goel, Aman, et al.
Veröffentlicht: (2025)
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025)
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025)
Towards Understanding the Robustness of Sparse Autoencoders
von: Saiyed, Ahson, et al.
Veröffentlicht: (2026)
von: Saiyed, Ahson, et al.
Veröffentlicht: (2026)
Toxicity Detection towards Adaptability to Changing Perturbations
von: Kang, Hankun, et al.
Veröffentlicht: (2024)
von: Kang, Hankun, et al.
Veröffentlicht: (2024)
STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents
von: Li, Jing-Jing, et al.
Veröffentlicht: (2025)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2025)
Preference Tuning For Toxicity Mitigation Generalizes Across Languages
von: Li, Xiaochen, et al.
Veröffentlicht: (2024)
von: Li, Xiaochen, et al.
Veröffentlicht: (2024)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
von: Wahed, Muntasir, et al.
Veröffentlicht: (2025)
von: Wahed, Muntasir, et al.
Veröffentlicht: (2025)
On Adversarial Robustness of Language Models in Transfer Learning
von: Turbal, Bohdan, et al.
Veröffentlicht: (2024)
von: Turbal, Bohdan, et al.
Veröffentlicht: (2024)
Directional Embedding Smoothing for Robust Vision Language Models
von: Wang, Ye, et al.
Veröffentlicht: (2026)
von: Wang, Ye, et al.
Veröffentlicht: (2026)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
FIT to Forget: Robust Continual Unlearning for Large Language Models
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
von: Aldahoul, Nouar, et al.
Veröffentlicht: (2025)
von: Aldahoul, Nouar, et al.
Veröffentlicht: (2025)
Was it Slander? Towards Exact Inversion of Generative Language Models
von: Skapars, Adrians, et al.
Veröffentlicht: (2024)
von: Skapars, Adrians, et al.
Veröffentlicht: (2024)
PostMark: A Robust Blackbox Watermark for Large Language Models
von: Chang, Yapei, et al.
Veröffentlicht: (2024)
von: Chang, Yapei, et al.
Veröffentlicht: (2024)
SELF: A Robust Singular Value and Eigenvalue Approach for LLM Fingerprinting
von: Zhang, Hanxiu, et al.
Veröffentlicht: (2025)
von: Zhang, Hanxiu, et al.
Veröffentlicht: (2025)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing
von: Li, Yuexin, et al.
Veröffentlicht: (2026)
von: Li, Yuexin, et al.
Veröffentlicht: (2026)
Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?
von: Liang, Zi, et al.
Veröffentlicht: (2025)
von: Liang, Zi, et al.
Veröffentlicht: (2025)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
von: Poppi, Samuele, et al.
Veröffentlicht: (2024)
von: Poppi, Samuele, et al.
Veröffentlicht: (2024)
SVIP: Towards Verifiable Inference of Open-source Large Language Models
von: Sun, Yifan, et al.
Veröffentlicht: (2024)
von: Sun, Yifan, et al.
Veröffentlicht: (2024)
Beyond Gradient and Priors in Privacy Attacks: Leveraging Pooler Layer Inputs of Language Models in Federated Learning
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
Special Characters Attack: Toward Scalable Training Data Extraction From Large Language Models
von: Bai, Yang, et al.
Veröffentlicht: (2024)
von: Bai, Yang, et al.
Veröffentlicht: (2024)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models
von: Dang, Trung Cuong, et al.
Veröffentlicht: (2025)
von: Dang, Trung Cuong, et al.
Veröffentlicht: (2025)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
von: Qi, Zhenting, et al.
Veröffentlicht: (2024)
von: Qi, Zhenting, et al.
Veröffentlicht: (2024)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
Machine Unlearning of Pre-trained Large Language Models
von: Yao, Jin, et al.
Veröffentlicht: (2024)
von: Yao, Jin, et al.
Veröffentlicht: (2024)
Less is More: Understanding Word-level Textual Adversarial Attack via n-gram Frequency Descend
von: Lu, Ning, et al.
Veröffentlicht: (2023)
von: Lu, Ning, et al.
Veröffentlicht: (2023)
Soft-Label Integration for Robust Toxicity Classification
von: Cheng, Zelei, et al.
Veröffentlicht: (2024)
von: Cheng, Zelei, et al.
Veröffentlicht: (2024)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
von: Yuan, Zhuowen, et al.
Veröffentlicht: (2024)
von: Yuan, Zhuowen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TaeBench: Improving Quality of Toxic Adversarial Examples
von: Zhu, Xuan, et al.
Veröffentlicht: (2024) -
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice
von: Goel, Aman, et al.
Veröffentlicht: (2025) -
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025) -
Towards Understanding the Robustness of Sparse Autoencoders
von: Saiyed, Ahson, et al.
Veröffentlicht: (2026) -
Toxicity Detection towards Adaptability to Changing Perturbations
von: Kang, Hankun, et al.
Veröffentlicht: (2024)