Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, San, Lee, Gary Geunbae |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Assistive Large Language Model Agents for Socially-Aware Negotiation Dialogues
di: Hua, Yuncheng, et al.
Pubblicazione: (2024)
di: Hua, Yuncheng, et al.
Pubblicazione: (2024)
Instructional Agents: Reducing Teaching Faculty Workload through Multi-Agent Instructional Design
di: Yao, Huaiyuan, et al.
Pubblicazione: (2025)
di: Yao, Huaiyuan, et al.
Pubblicazione: (2025)
Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data
di: Lübbers, Christopher Lee
Pubblicazione: (2025)
di: Lübbers, Christopher Lee
Pubblicazione: (2025)
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
di: Lee, Yejin, et al.
Pubblicazione: (2025)
di: Lee, Yejin, et al.
Pubblicazione: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)
di: Saji, Alan, et al.
Pubblicazione: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
di: Kautsar, Muhammad Dehan Al, et al.
Pubblicazione: (2026)
di: Kautsar, Muhammad Dehan Al, et al.
Pubblicazione: (2026)
Predictive Simultaneous Interpretation: Harnessing Large Language Models for Democratizing Real-Time Multilingual Communication
di: Iida, Kurando, et al.
Pubblicazione: (2024)
di: Iida, Kurando, et al.
Pubblicazione: (2024)
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation
di: Yuan, Weikang, et al.
Pubblicazione: (2025)
di: Yuan, Weikang, et al.
Pubblicazione: (2025)
Semantic Delta: An Interpretable Signal Differentiating Human and LLMs Dialogue
di: Scantamburlo, Riccardo, et al.
Pubblicazione: (2026)
di: Scantamburlo, Riccardo, et al.
Pubblicazione: (2026)
Efficient Toxicity Detection in Gaming Chats: A Comparative Study of Embeddings, Fine-Tuned Transformers and LLMs
di: Tereshchenko, Yehor, et al.
Pubblicazione: (2025)
di: Tereshchenko, Yehor, et al.
Pubblicazione: (2025)
SADAS: A Dialogue Assistant System Towards Remediating Norm Violations in Bilingual Socio-Cultural Conversations
di: Hua, Yuncheng, et al.
Pubblicazione: (2024)
di: Hua, Yuncheng, et al.
Pubblicazione: (2024)
HR-Agent: A Task-Oriented Dialogue (TOD) LLM Agent Tailored for HR Applications
di: Xu, Weijie, et al.
Pubblicazione: (2024)
di: Xu, Weijie, et al.
Pubblicazione: (2024)
DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory
di: Zhou, Wenxuan, et al.
Pubblicazione: (2025)
di: Zhou, Wenxuan, et al.
Pubblicazione: (2025)
Bypassing DARCY Defense: Indistinguishable Universal Adversarial Triggers
di: Peng, Zuquan, et al.
Pubblicazione: (2024)
di: Peng, Zuquan, et al.
Pubblicazione: (2024)
A Practical Approach for Building Production-Grade Conversational Agents with Workflow Graphs
di: Park, Chiwan, et al.
Pubblicazione: (2025)
di: Park, Chiwan, et al.
Pubblicazione: (2025)
Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
LASTIST: LArge-Scale Target-Independent STance dataset
di: Kim, DongJae, et al.
Pubblicazione: (2025)
di: Kim, DongJae, et al.
Pubblicazione: (2025)
Proactive Agent: Shifting LLM Agents from Reactive Responses to Active Assistance
di: Lu, Yaxi, et al.
Pubblicazione: (2024)
di: Lu, Yaxi, et al.
Pubblicazione: (2024)
Fine-Tuned Large Language Models for Logical Translation: Reducing Hallucinations with Lang2Logic
di: Pan, Muyu, et al.
Pubblicazione: (2025)
di: Pan, Muyu, et al.
Pubblicazione: (2025)
HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent
di: Xu, Weijie, et al.
Pubblicazione: (2024)
di: Xu, Weijie, et al.
Pubblicazione: (2024)
Machine Unlearning for Masked Diffusion Language Models
di: Lee, Georu, et al.
Pubblicazione: (2026)
di: Lee, Georu, et al.
Pubblicazione: (2026)
Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning
di: Xu, Shuyao, et al.
Pubblicazione: (2025)
di: Xu, Shuyao, et al.
Pubblicazione: (2025)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
di: Wu, Shuai, et al.
Pubblicazione: (2026)
di: Wu, Shuai, et al.
Pubblicazione: (2026)
USTCCTSU at SemEval-2024 Task 1: Reducing Anisotropy for Cross-lingual Semantic Textual Relatedness Task
di: Li, Jianjian, et al.
Pubblicazione: (2024)
di: Li, Jianjian, et al.
Pubblicazione: (2024)
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
di: Le, Linh, et al.
Pubblicazione: (2026)
di: Le, Linh, et al.
Pubblicazione: (2026)
SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation
di: Kim, Seoyeon, et al.
Pubblicazione: (2026)
di: Kim, Seoyeon, et al.
Pubblicazione: (2026)
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
di: Wang, Xinyue, et al.
Pubblicazione: (2026)
di: Wang, Xinyue, et al.
Pubblicazione: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments
di: Gu, Yu, et al.
Pubblicazione: (2024)
di: Gu, Yu, et al.
Pubblicazione: (2024)
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
di: Siegel, Noah Y., et al.
Pubblicazione: (2025)
di: Siegel, Noah Y., et al.
Pubblicazione: (2025)
More Is Not Always Better: Cross-Component Interference in LLM Agent Scaffolding
di: Liu, Ming
Pubblicazione: (2026)
di: Liu, Ming
Pubblicazione: (2026)
Persona Inconstancy in Multi-Agent LLM Collaboration: Conformity, Confabulation, and Impersonation
di: Baltaji, Razan, et al.
Pubblicazione: (2024)
di: Baltaji, Razan, et al.
Pubblicazione: (2024)
Blessing or curse? A survey on the Impact of Generative AI on Fake News
di: Loth, Alexander, et al.
Pubblicazione: (2024)
di: Loth, Alexander, et al.
Pubblicazione: (2024)
Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams
di: Llorente-Saguer, Isaac
Pubblicazione: (2026)
di: Llorente-Saguer, Isaac
Pubblicazione: (2026)
Persuasiveness and Bias in LLM: Investigating the Impact of Persuasiveness and Reinforcement of Bias in Language Models
di: Roy, Saumya
Pubblicazione: (2025)
di: Roy, Saumya
Pubblicazione: (2025)
Self-Emotion Blended Dialogue Generation in Social Simulation Agents
di: Zhang, Qiang, et al.
Pubblicazione: (2024)
di: Zhang, Qiang, et al.
Pubblicazione: (2024)
Argumentatively Coherent Judgmental Forecasting
di: Gorur, Deniz, et al.
Pubblicazione: (2025)
di: Gorur, Deniz, et al.
Pubblicazione: (2025)
From Reading to Compressing: Exploring the Multi-document Reader for Prompt Compression
di: Choi, Eunseong, et al.
Pubblicazione: (2024)
di: Choi, Eunseong, et al.
Pubblicazione: (2024)
Low-Resource Court Judgment Summarization for Common Law Systems
di: Liu, Shuaiqi, et al.
Pubblicazione: (2024)
di: Liu, Shuaiqi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Assistive Large Language Model Agents for Socially-Aware Negotiation Dialogues
di: Hua, Yuncheng, et al.
Pubblicazione: (2024) -
Instructional Agents: Reducing Teaching Faculty Workload through Multi-Agent Instructional Design
di: Yao, Huaiyuan, et al.
Pubblicazione: (2025) -
Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data
di: Lübbers, Christopher Lee
Pubblicazione: (2025) -
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
di: Lee, Yejin, et al.
Pubblicazione: (2025) -
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)