Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nöther, Jonathan, Singla, Adish, Radanović, Goran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025)
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025)
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
von: Nöther, Jonathan, et al.
Veröffentlicht: (2026)
von: Nöther, Jonathan, et al.
Veröffentlicht: (2026)
Policy Teaching via Data Poisoning in Learning from Human Preferences
von: Nika, Andi, et al.
Veröffentlicht: (2025)
von: Nika, Andi, et al.
Veröffentlicht: (2025)
Learning Embeddings for Sequential Tasks Using Population of Agents
von: Mahajan, Mridul, et al.
Veröffentlicht: (2023)
von: Mahajan, Mridul, et al.
Veröffentlicht: (2023)
Corruption-Robust Offline Two-Player Zero-Sum Markov Games
von: Nika, Andi, et al.
Veröffentlicht: (2024)
von: Nika, Andi, et al.
Veröffentlicht: (2024)
AgenticRed: Evolving Agentic Systems for Red-Teaming
von: Yuan, Jiayi, et al.
Veröffentlicht: (2026)
von: Yuan, Jiayi, et al.
Veröffentlicht: (2026)
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
von: Nika, Andi, et al.
Veröffentlicht: (2026)
von: Nika, Andi, et al.
Veröffentlicht: (2026)
Corruption Robust Offline Reinforcement Learning with Human Feedback
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
von: Nika, Andi, et al.
Veröffentlicht: (2024)
von: Nika, Andi, et al.
Veröffentlicht: (2024)
Proximal Curriculum with Task Correlations for Deep Reinforcement Learning
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2024)
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2024)
Hints-In-Browser: Benchmarking Language Models for Programming Feedback Generation
von: Kotalwar, Nachiket, et al.
Veröffentlicht: (2024)
von: Kotalwar, Nachiket, et al.
Veröffentlicht: (2024)
Towards Generalizable Agents in Text-Based Educational Environments: A Study of Integrating RL with LLMs
von: Radmehr, Bahar, et al.
Veröffentlicht: (2024)
von: Radmehr, Bahar, et al.
Veröffentlicht: (2024)
Reward Design for Justifiable Sequential Decision-Making
von: Sukovic, Aleksa, et al.
Veröffentlicht: (2024)
von: Sukovic, Aleksa, et al.
Veröffentlicht: (2024)
Performative Reinforcement Learning with Linear Markov Decision Process
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2025)
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2025)
Informativeness of Reward Functions in Reinforcement Learning
von: Devidze, Rati, et al.
Veröffentlicht: (2024)
von: Devidze, Rati, et al.
Veröffentlicht: (2024)
On Corruption-Robustness in Performative Reinforcement Learning
von: Pollatos, Vasilis, et al.
Veröffentlicht: (2025)
von: Pollatos, Vasilis, et al.
Veröffentlicht: (2025)
Abstractive Red-Teaming of Language Model Character
von: Rahn, Nate, et al.
Veröffentlicht: (2026)
von: Rahn, Nate, et al.
Veröffentlicht: (2026)
Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?
von: Sasnauskas, Paulius, et al.
Veröffentlicht: (2025)
von: Sasnauskas, Paulius, et al.
Veröffentlicht: (2025)
Distributionally Robust Reinforcement Learning with Human Feedback
von: Mandal, Debmalya, et al.
Veröffentlicht: (2025)
von: Mandal, Debmalya, et al.
Veröffentlicht: (2025)
Performative Reinforcement Learning in Gradually Shifting Environments
von: Rank, Ben, et al.
Veröffentlicht: (2024)
von: Rank, Ben, et al.
Veröffentlicht: (2024)
Optimal Decision Making Under Strategic Behavior
von: Tsirtsis, Stratis, et al.
Veröffentlicht: (2019)
von: Tsirtsis, Stratis, et al.
Veröffentlicht: (2019)
Formal Models of Active Learning from Contrastive Examples
von: Mansouri, Farnam, et al.
Veröffentlicht: (2025)
von: Mansouri, Farnam, et al.
Veröffentlicht: (2025)
Red-Teaming for Inducing Societal Bias in Large Language Models
von: Luo, Chu Fei, et al.
Veröffentlicht: (2024)
von: Luo, Chu Fei, et al.
Veröffentlicht: (2024)
Independent Learning in Performative Markov Potential Games
von: Sahitaj, Rilind, et al.
Veröffentlicht: (2025)
von: Sahitaj, Rilind, et al.
Veröffentlicht: (2025)
DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
Neural Task Synthesis for Visual Programming
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2023)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2023)
Inference-Time Personalized Alignment with a Few User Preference Queries
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
von: Ding, Jiale, et al.
Veröffentlicht: (2025)
von: Ding, Jiale, et al.
Veröffentlicht: (2025)
Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image Generation
von: Quaye, Jessica, et al.
Veröffentlicht: (2024)
von: Quaye, Jessica, et al.
Veröffentlicht: (2024)
Stochastic Principal-Agent Problems: Efficient Computation and Learning
von: Gan, Jiarui, et al.
Veröffentlicht: (2023)
von: Gan, Jiarui, et al.
Veröffentlicht: (2023)
Large Language Models for In-Context Student Modeling: Synthesizing Student's Behavior in Visual Programming
von: Nguyen, Manh Hung, et al.
Veröffentlicht: (2023)
von: Nguyen, Manh Hung, et al.
Veröffentlicht: (2023)
Sparse Offline Reinforcement Learning with Corruption Robustness
von: Tran, Nam Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Nam Phuong, et al.
Veröffentlicht: (2025)
LLM-Assisted Red Teaming of Diffusion Models through "Failures Are Fated, But Can Be Faded"
von: Sagar, Som, et al.
Veröffentlicht: (2024)
von: Sagar, Som, et al.
Veröffentlicht: (2024)
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
von: Hu, Kai, et al.
Veröffentlicht: (2025)
von: Hu, Kai, et al.
Veröffentlicht: (2025)
Learning Half-Spaces from Perturbed Contrastive Examples
von: Ravari, Aryan Alavi Razavi, et al.
Veröffentlicht: (2026)
von: Ravari, Aryan Alavi Razavi, et al.
Veröffentlicht: (2026)
Divergent-Convergent Thinking in Large Language Models for Creative Problem Generation
von: Nguyen, Manh Hung, et al.
Veröffentlicht: (2025)
von: Nguyen, Manh Hung, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Durable Algorithmic Recourse
von: Ceccon, Marina, et al.
Veröffentlicht: (2025)
von: Ceccon, Marina, et al.
Veröffentlicht: (2025)
Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient Influence
von: Nguyen, Quoc Minh, et al.
Veröffentlicht: (2026)
von: Nguyen, Quoc Minh, et al.
Veröffentlicht: (2026)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024)
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025) -
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
von: Nöther, Jonathan, et al.
Veröffentlicht: (2026) -
Policy Teaching via Data Poisoning in Learning from Human Preferences
von: Nika, Andi, et al.
Veröffentlicht: (2025) -
Learning Embeddings for Sequential Tasks Using Population of Agents
von: Mahajan, Mridul, et al.
Veröffentlicht: (2023) -
Corruption-Robust Offline Two-Player Zero-Sum Markov Games
von: Nika, Andi, et al.
Veröffentlicht: (2024)