Purple-teaming LLMs with Adversarial Defender Training
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Jingyan, Li, Kun, Li, Junan, Kang, Jiawen, Hu, Minda, Wu, Xixin, Meng, Helen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Machine Ethics -- Can LLMs Perform Moral Reasoning through the Lens of Moral Theories?
by: Zhou, Jingyan, et al.
Published: (2023)
by: Zhou, Jingyan, et al.
Published: (2023)
Not All Errors Are Equal: Investigation of Speech Recognition Errors in Alzheimer's Disease Detection
by: Kang, Jiawen, et al.
Published: (2024)
by: Kang, Jiawen, et al.
Published: (2024)
Devising a Set of Compact and Explainable Spoken Language Feature for Screening Alzheimer's Disease
by: Li, Junan, et al.
Published: (2024)
by: Li, Junan, et al.
Published: (2024)
On the Within-class Variation Issue in Alzheimer's Disease Detection
by: Kang, Jiawen, et al.
Published: (2024)
by: Kang, Jiawen, et al.
Published: (2024)
TreePS-RAG: Tree-based Process Supervision for Reinforcement Learning in Agentic RAG
by: Zhang, Tianhua, et al.
Published: (2026)
by: Zhang, Tianhua, et al.
Published: (2026)
Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational Answers
by: Zhang, Tianhua, et al.
Published: (2024)
by: Zhang, Tianhua, et al.
Published: (2024)
Decoding on Graphs: Faithful and Sound Reasoning on Knowledge Graphs through Generation of Well-Formed Chains
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
by: Meng, Lingwei, et al.
Published: (2024)
by: Meng, Lingwei, et al.
Published: (2024)
Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
by: Meng, Lingwei, et al.
Published: (2024)
by: Meng, Lingwei, et al.
Published: (2024)
RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning
by: Li, Kun, et al.
Published: (2025)
by: Li, Kun, et al.
Published: (2025)
Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
by: Kang, Jiawen, et al.
Published: (2024)
by: Kang, Jiawen, et al.
Published: (2024)
Injecting linguistic knowledge into BERT for Dialogue State Tracking
by: Feng, Xiaohan, et al.
Published: (2023)
by: Feng, Xiaohan, et al.
Published: (2023)
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
by: Wang, Dingdong, et al.
Published: (2025)
by: Wang, Dingdong, et al.
Published: (2025)
Generate, Discriminate, Evolve: Enhancing Context Faithfulness via Fine-Grained Sentence-Level Self-Evolution
by: Li, Kun, et al.
Published: (2025)
by: Li, Kun, et al.
Published: (2025)
Agentic Cognitive Profiling: Realigning Automated Alzheimer's Disease Detection with Clinical Construct Validity
by: Kang, Jiawen, et al.
Published: (2026)
by: Kang, Jiawen, et al.
Published: (2026)
Seamless Language Expansion: Enhancing Multilingual Mastery in Self-Supervised Models
by: Xu, Jing, et al.
Published: (2024)
by: Xu, Jing, et al.
Published: (2024)
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
by: Wu, Wenxuan, et al.
Published: (2025)
by: Wu, Wenxuan, et al.
Published: (2025)
Large Language Model-based FMRI Encoding of Language Functions for Subjects with Neurocognitive Disorder
by: Wang, Yuejiao, et al.
Published: (2024)
by: Wang, Yuejiao, et al.
Published: (2024)
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
by: Li, Yuanfan, et al.
Published: (2025)
by: Li, Yuanfan, et al.
Published: (2025)
Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching
by: Zhang, Xiaoying, et al.
Published: (2024)
by: Zhang, Xiaoying, et al.
Published: (2024)
SeRTS: Self-Rewarding Tree Search for Biomedical Retrieval-Augmented Generation
by: Hu, Minda, et al.
Published: (2024)
by: Hu, Minda, et al.
Published: (2024)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
by: Wang, Yuejiao, et al.
Published: (2024)
by: Wang, Yuejiao, et al.
Published: (2024)
MiLorE-SSL: Scaling Multilingual Capabilities in Self-Supervised Models without Forgetting
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
Lamer-SSL: Layer-aware Mixture of LoRA Experts for Continual Multilingual Expansion of Self-supervised Models without Forgetting
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
ConMax: Confidence-Maximizing Compression for Efficient Chain-of-Thought Reasoning
by: Hu, Minda, et al.
Published: (2026)
by: Hu, Minda, et al.
Published: (2026)
NILE: Internal Consistency Alignment in Large Language Models
by: Hu, Minda, et al.
Published: (2024)
by: Hu, Minda, et al.
Published: (2024)
CSSBench: Evaluating the Safety of Lightweight LLMs against Chinese-Specific Adversarial Patterns
by: Zhou, Zhenhong, et al.
Published: (2026)
by: Zhou, Zhenhong, et al.
Published: (2026)
Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
by: Zhang, Xiaoying, et al.
Published: (2024)
by: Zhang, Xiaoying, et al.
Published: (2024)
Discovery and Reinforcement of Tool-Integrated Reasoning Chains via Rollout Trees
by: Li, Kun, et al.
Published: (2026)
by: Li, Kun, et al.
Published: (2026)
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
by: Wang, Dingdong, et al.
Published: (2025)
by: Wang, Dingdong, et al.
Published: (2025)
WebCoT: Enhancing Web Agent Reasoning by Reconstructing Chain-of-Thought in Reflection, Branching, and Rollback
by: Hu, Minda, et al.
Published: (2025)
by: Hu, Minda, et al.
Published: (2025)
Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
by: Li, Weiqin, et al.
Published: (2024)
by: Li, Weiqin, et al.
Published: (2024)
Fast Adversarial Training against Textual Adversarial Attacks
by: Yang, Yichen, et al.
Published: (2024)
by: Yang, Yichen, et al.
Published: (2024)
How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue
by: Lu, Hui, et al.
Published: (2026)
by: Lu, Hui, et al.
Published: (2026)
Defending Against Social Engineering Attacks in the Age of LLMs
by: Ai, Lin, et al.
Published: (2024)
by: Ai, Lin, et al.
Published: (2024)
The Integration of Semantic and Structural Knowledge in Knowledge Graph Entity Typing
by: Li, Muzhi, et al.
Published: (2024)
by: Li, Muzhi, et al.
Published: (2024)
Intention Analysis Makes LLMs A Good Jailbreak Defender
by: Zhang, Yuqi, et al.
Published: (2024)
by: Zhang, Yuqi, et al.
Published: (2024)
Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
Similar Items
-
Rethinking Machine Ethics -- Can LLMs Perform Moral Reasoning through the Lens of Moral Theories?
by: Zhou, Jingyan, et al.
Published: (2023) -
Not All Errors Are Equal: Investigation of Speech Recognition Errors in Alzheimer's Disease Detection
by: Kang, Jiawen, et al.
Published: (2024) -
Devising a Set of Compact and Explainable Spoken Language Feature for Screening Alzheimer's Disease
by: Li, Junan, et al.
Published: (2024) -
On the Within-class Variation Issue in Alzheimer's Disease Detection
by: Kang, Jiawen, et al.
Published: (2024) -
TreePS-RAG: Tree-based Process Supervision for Reinforcement Learning in Agentic RAG
by: Zhang, Tianhua, et al.
Published: (2026)