ToxiGAN: Toxic Data Augmentation via LLM-Guided Directional Adversarial Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Peiran, Fillies, Jan, Paschke, Adrian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Hate Speech Classification with Cross-Taxonomy Dataset Integration
von: Fillies, Jan, et al.
Veröffentlicht: (2025)
von: Fillies, Jan, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Generation-based Relation Extraction
von: Efeoglu, Sefika, et al.
Veröffentlicht: (2024)
von: Efeoglu, Sefika, et al.
Veröffentlicht: (2024)
ToxiLab: How Well Do Open-Source LLMs Generate Synthetic Toxicity Data?
von: Hui, Zheng, et al.
Veröffentlicht: (2024)
von: Hui, Zheng, et al.
Veröffentlicht: (2024)
Bolster Hallucination Detection via Prompt-Guided Data Augmentation
von: Li, Wenyun, et al.
Veröffentlicht: (2025)
von: Li, Wenyun, et al.
Veröffentlicht: (2025)
Relation Extraction with Fine-Tuned Large Language Models in Retrieval Augmented Generation Frameworks
von: Efeoglu, Sefika, et al.
Veröffentlicht: (2024)
von: Efeoglu, Sefika, et al.
Veröffentlicht: (2024)
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
von: Shimabucoro, Luísa, et al.
Veröffentlicht: (2024)
von: Shimabucoro, Luísa, et al.
Veröffentlicht: (2024)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
Malinowski in the Age of AI: Can large language models create a text game based on an anthropological classic?
von: Hoffmann, Michael Peter, et al.
Veröffentlicht: (2024)
von: Hoffmann, Michael Peter, et al.
Veröffentlicht: (2024)
Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game
von: Cheng, Pengyu, et al.
Veröffentlicht: (2023)
von: Cheng, Pengyu, et al.
Veröffentlicht: (2023)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
von: Li, Miaomiao, et al.
Veröffentlicht: (2025)
von: Li, Miaomiao, et al.
Veröffentlicht: (2025)
Designing and Evaluating Malinowski's Lens: An AI-Native Educational Game for Ethnographic Learning
von: Hoffmann, Michael, et al.
Veröffentlicht: (2025)
von: Hoffmann, Michael, et al.
Veröffentlicht: (2025)
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
von: Tang, Haochun, et al.
Veröffentlicht: (2026)
von: Tang, Haochun, et al.
Veröffentlicht: (2026)
PSK@EEUCA 2026: Fine-Tuning Large Language Models with Synthetic Data Augmentation for Multi-Class Toxicity Detection in Gaming Chat
von: Pulipaka, Srikar Kashyap
Veröffentlicht: (2026)
von: Pulipaka, Srikar Kashyap
Veröffentlicht: (2026)
Lightweight Safety Guardrails via Synthetic Data and RL-guided Adversarial Training
von: Ilin, Aleksei, et al.
Veröffentlicht: (2025)
von: Ilin, Aleksei, et al.
Veröffentlicht: (2025)
TaeBench: Improving Quality of Toxic Adversarial Examples
von: Zhu, Xuan, et al.
Veröffentlicht: (2024)
von: Zhu, Xuan, et al.
Veröffentlicht: (2024)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
von: Jing, Yi, et al.
Veröffentlicht: (2026)
von: Jing, Yi, et al.
Veröffentlicht: (2026)
Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word Importance
von: Tuck, Bryan E., et al.
Veröffentlicht: (2025)
von: Tuck, Bryan E., et al.
Veröffentlicht: (2025)
Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
von: Ding, Yiwen, et al.
Veröffentlicht: (2024)
von: Ding, Yiwen, et al.
Veröffentlicht: (2024)
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
von: Kaneko, Masahiro
Veröffentlicht: (2026)
von: Kaneko, Masahiro
Veröffentlicht: (2026)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
von: Yu, Xiaodong, et al.
Veröffentlicht: (2023)
von: Yu, Xiaodong, et al.
Veröffentlicht: (2023)
GCC-Spam: Spam Detection via GAN, Contrastive Learning, and Character Similarity Networks
von: Wang, Zhijie, et al.
Veröffentlicht: (2025)
von: Wang, Zhijie, et al.
Veröffentlicht: (2025)
Data Cartography for Detecting Memorization Hotspots and Guiding Data Interventions in Generative Models
von: Patel, Laksh, et al.
Veröffentlicht: (2025)
von: Patel, Laksh, et al.
Veröffentlicht: (2025)
Enhancing NLP Robustness and Generalization through LLM-Generated Contrast Sets: A Scalable Framework for Systematic Evaluation and Adversarial Training
von: Lin, Hender
Veröffentlicht: (2025)
von: Lin, Hender
Veröffentlicht: (2025)
Forging the Forger: An Attempt to Improve Authorship Verification via Data Augmentation
von: Corbara, Silvia, et al.
Veröffentlicht: (2024)
von: Corbara, Silvia, et al.
Veröffentlicht: (2024)
LLM-Match: An Open-Sourced Patient Matching Model Based on Large Language Models and Retrieval-Augmented Generation
von: Li, Xiaodi, et al.
Veröffentlicht: (2025)
von: Li, Xiaodi, et al.
Veröffentlicht: (2025)
LLM-Based Adversarial Persuasion Attacks on Fact-Checking Systems
von: Leite, João A., et al.
Veröffentlicht: (2026)
von: Leite, João A., et al.
Veröffentlicht: (2026)
Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models
von: Jadon, Aryan, et al.
Veröffentlicht: (2025)
von: Jadon, Aryan, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Data Augmentation for Low-Resource Domain Tasks
von: Seo, Minju, et al.
Veröffentlicht: (2024)
von: Seo, Minju, et al.
Veröffentlicht: (2024)
Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation
von: Balestriero, Randall, et al.
Veröffentlicht: (2023)
von: Balestriero, Randall, et al.
Veröffentlicht: (2023)
Set-LLM: A Permutation-Invariant LLM
von: Egressy, Beni, et al.
Veröffentlicht: (2025)
von: Egressy, Beni, et al.
Veröffentlicht: (2025)
Exploring Gradient-Guided Masked Language Model to Detect Textual Adversarial Attacks
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2025)
LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
von: Huang, Wei, et al.
Veröffentlicht: (2026)
von: Huang, Wei, et al.
Veröffentlicht: (2026)
LLM Unlearning via Loss Adjustment with Only Forget Data
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
Automated Rewards via LLM-Generated Progress Functions
von: Sarukkai, Vishnu, et al.
Veröffentlicht: (2024)
von: Sarukkai, Vishnu, et al.
Veröffentlicht: (2024)
KV Cache Transform Coding for Compact Storage in LLM Inference
von: Staniszewski, Konrad, et al.
Veröffentlicht: (2025)
von: Staniszewski, Konrad, et al.
Veröffentlicht: (2025)
Preference Tuning For Toxicity Mitigation Generalizes Across Languages
von: Li, Xiaochen, et al.
Veröffentlicht: (2024)
von: Li, Xiaochen, et al.
Veröffentlicht: (2024)
GUARD: Guided Unlearning and Retention via Data Attribution for Large Language Models
von: Niu, Peizhi, et al.
Veröffentlicht: (2025)
von: Niu, Peizhi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Improving Hate Speech Classification with Cross-Taxonomy Dataset Integration
von: Fillies, Jan, et al.
Veröffentlicht: (2025) -
Retrieval-Augmented Generation-based Relation Extraction
von: Efeoglu, Sefika, et al.
Veröffentlicht: (2024) -
ToxiLab: How Well Do Open-Source LLMs Generate Synthetic Toxicity Data?
von: Hui, Zheng, et al.
Veröffentlicht: (2024) -
Bolster Hallucination Detection via Prompt-Guided Data Augmentation
von: Li, Wenyun, et al.
Veröffentlicht: (2025) -
Relation Extraction with Fine-Tuned Large Language Models in Retrieval Augmented Generation Frameworks
von: Efeoglu, Sefika, et al.
Veröffentlicht: (2024)