Dark and Bright Side of Participatory Red-Teaming with Targets of Stereotyping for Eliciting Harmful Behaviors from Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Sieun, Jo, Yeeun, Na, Sungmin, Lim, Hyunseung, Lee, Eunchae, Choi, Yu Min, Cho, Soohyun, Hong, Hwajung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Identify Design Problems Through Questioning: Exploring Role-playing Interactions with Large Language Models to Foster Design Questioning Skills
di: Lim, Hyunseung, et al.
Pubblicazione: (2024)
di: Lim, Hyunseung, et al.
Pubblicazione: (2024)
Understanding Human-Multi-Agent Team Formation for Creative Work
di: Lim, Hyunseung, et al.
Pubblicazione: (2026)
di: Lim, Hyunseung, et al.
Pubblicazione: (2026)
A Large-Scale Real-World Evaluation of LLM-Based Virtual Teaching Assistant
di: Kweon, Sunjun, et al.
Pubblicazione: (2025)
di: Kweon, Sunjun, et al.
Pubblicazione: (2025)
Feed-O-Meter: Investigating AI-Generated Mentee Personas as Interactive Agents for Scaffolding Design Feedback Practice
di: Lim, Hyunseung, et al.
Pubblicazione: (2025)
di: Lim, Hyunseung, et al.
Pubblicazione: (2025)
PANORAMA: A Dataset and Benchmarks Capturing Decision Trails and Rationales in Patent Examination
di: Lim, Hyunseung, et al.
Pubblicazione: (2025)
di: Lim, Hyunseung, et al.
Pubblicazione: (2025)
Beyond Performance Disparities: A Three-Level Audit of Representational Harm in CelebA
di: Park, Sieun, et al.
Pubblicazione: (2026)
di: Park, Sieun, et al.
Pubblicazione: (2026)
Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews
di: Shin, Hyungyu, et al.
Pubblicazione: (2025)
di: Shin, Hyungyu, et al.
Pubblicazione: (2025)
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints
di: Nöther, Jonathan, et al.
Pubblicazione: (2025)
di: Nöther, Jonathan, et al.
Pubblicazione: (2025)
Pointwise estimates of the Bergman kernel with an exponential weight on the unit ball
di: Cho, Hong Rae, et al.
Pubblicazione: (2024)
di: Cho, Hong Rae, et al.
Pubblicazione: (2024)
The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI Relationships
di: Zhang, Renwen, et al.
Pubblicazione: (2024)
di: Zhang, Renwen, et al.
Pubblicazione: (2024)
Less Interaction But More Explanation: A Communication Perspective on Agentic AI Interfaces
di: Jang, Eunchae, et al.
Pubblicazione: (2026)
di: Jang, Eunchae, et al.
Pubblicazione: (2026)
When Scaffolding Breaks: Investigating Student Interaction with LLM-Based Writing Support in Real-Time K-12 EFL Classrooms
di: Myung, Junho, et al.
Pubblicazione: (2025)
di: Myung, Junho, et al.
Pubblicazione: (2025)
Time is Not Enough: Time-Frequency based Explanation for Time-Series Black-Box Models
di: Chung, Hyunseung, et al.
Pubblicazione: (2024)
di: Chung, Hyunseung, et al.
Pubblicazione: (2024)
Maintenance of Sinus Rhythm Is Associated With Lower Incidence of Stroke in Patients With Drug‐Refractory Atrial Fibrillation
di: Soohyun Kim, et al.
Pubblicazione: (2024)
di: Soohyun Kim, et al.
Pubblicazione: (2024)
Beyond Forgetting: Machine Unlearning Elicits Controllable Side Behaviors and Capabilities
di: Dang, Tien, et al.
Pubblicazione: (2026)
di: Dang, Tien, et al.
Pubblicazione: (2026)
Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South
di: Rastogi, Charvi, et al.
Pubblicazione: (2026)
di: Rastogi, Charvi, et al.
Pubblicazione: (2026)
How Far I'll Go: Imagining Futures of Conversational AI with People with Visual Impairments Through Design Fiction
di: Choi, Jeanne, et al.
Pubblicazione: (2025)
di: Choi, Jeanne, et al.
Pubblicazione: (2025)
Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors
di: Sicilia, Anthony, et al.
Pubblicazione: (2024)
di: Sicilia, Anthony, et al.
Pubblicazione: (2024)
Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
di: Dimino, Fabrizio, et al.
Pubblicazione: (2026)
di: Dimino, Fabrizio, et al.
Pubblicazione: (2026)
Boundary Behavior of Bisectional Curvatures for Weighted Bergman Metrics
di: Yoo, Sungmin
Pubblicazione: (2026)
di: Yoo, Sungmin
Pubblicazione: (2026)
MF-LPR$^2$: Multi-Frame License Plate Image Restoration and Recognition using Optical Flow
di: Na, Kihyun, et al.
Pubblicazione: (2025)
di: Na, Kihyun, et al.
Pubblicazione: (2025)
STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming
di: Jung, MinJae, et al.
Pubblicazione: (2026)
di: Jung, MinJae, et al.
Pubblicazione: (2026)
Unveiling the Dark Side of UV/Optical Bright Galaxies: Optically Thick Dust Absorption
di: Cheng, Yingjie, et al.
Pubblicazione: (2024)
di: Cheng, Yingjie, et al.
Pubblicazione: (2024)
Disentangling the Leadership Theory Jungle: A Reconciliation of Bright and Dark Side Leadership Theories
di: Jianyun Tang, et al.
Pubblicazione: (2026)
di: Jianyun Tang, et al.
Pubblicazione: (2026)
The Bright Side of Timed Opacity
di: André, Étienne, et al.
Pubblicazione: (2024)
di: André, Étienne, et al.
Pubblicazione: (2024)
Designing for Understanding: How Interface-Level Consent Designs Shape Attention and Understanding in Privacy Disclosures
di: Xiao, Wei, et al.
Pubblicazione: (2026)
di: Xiao, Wei, et al.
Pubblicazione: (2026)
Red Teaming Multimodal Language Models: Evaluating Harm Across Prompt Modalities and Models
di: Van Doren, Madison, et al.
Pubblicazione: (2025)
di: Van Doren, Madison, et al.
Pubblicazione: (2025)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
di: Mazeika, Mantas, et al.
Pubblicazione: (2024)
di: Mazeika, Mantas, et al.
Pubblicazione: (2024)
Constructing Everyday Well-Being: Insights from God-Saeng for Personal Informatics
di: Song, Inhwa, et al.
Pubblicazione: (2026)
di: Song, Inhwa, et al.
Pubblicazione: (2026)
Tiny Refinements Elicit Resilience: Toward Efficient Prefix-Model Against LLM Red-Teaming
di: Liu, Jiaxu, et al.
Pubblicazione: (2024)
di: Liu, Jiaxu, et al.
Pubblicazione: (2024)
NoRe: Augmenting Journaling Experience with Generative AI for Music Creation
di: Park, Joonyoung, et al.
Pubblicazione: (2025)
di: Park, Joonyoung, et al.
Pubblicazione: (2025)
Contour-Guided Query-Based Feature Fusion for Boundary-Aware and Generalizable Cardiac Ultrasound Segmentation
di: Ullah, Zahid, et al.
Pubblicazione: (2026)
di: Ullah, Zahid, et al.
Pubblicazione: (2026)
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
di: Jones, Jaylen, et al.
Pubblicazione: (2026)
di: Jones, Jaylen, et al.
Pubblicazione: (2026)
Red Teaming AI Red Teaming
di: Majumdar, Subhabrata, et al.
Pubblicazione: (2025)
di: Majumdar, Subhabrata, et al.
Pubblicazione: (2025)
MC-GenRef: Annotation-free mammography microcalcification segmentation with generative posterior refinement
di: Cho, Hyunwoo, et al.
Pubblicazione: (2026)
di: Cho, Hyunwoo, et al.
Pubblicazione: (2026)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
di: Kaunismaa, Jackson, et al.
Pubblicazione: (2026)
di: Kaunismaa, Jackson, et al.
Pubblicazione: (2026)
Why Alignment Must Precede Distillation: A Minimal Working Explanation
di: Cha, Sungmin, et al.
Pubblicazione: (2025)
di: Cha, Sungmin, et al.
Pubblicazione: (2025)
Hyperparameters in Continual Learning: A Reality Check
di: Cha, Sungmin, et al.
Pubblicazione: (2024)
di: Cha, Sungmin, et al.
Pubblicazione: (2024)
Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation
di: Cha, Sungmin, et al.
Pubblicazione: (2025)
di: Cha, Sungmin, et al.
Pubblicazione: (2025)
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency
di: Kyung, Kwanhee, et al.
Pubblicazione: (2025)
di: Kyung, Kwanhee, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Identify Design Problems Through Questioning: Exploring Role-playing Interactions with Large Language Models to Foster Design Questioning Skills
di: Lim, Hyunseung, et al.
Pubblicazione: (2024) -
Understanding Human-Multi-Agent Team Formation for Creative Work
di: Lim, Hyunseung, et al.
Pubblicazione: (2026) -
A Large-Scale Real-World Evaluation of LLM-Based Virtual Teaching Assistant
di: Kweon, Sunjun, et al.
Pubblicazione: (2025) -
Feed-O-Meter: Investigating AI-Generated Mentee Personas as Interactive Agents for Scaffolding Design Feedback Practice
di: Lim, Hyunseung, et al.
Pubblicazione: (2025) -
PANORAMA: A Dataset and Benchmarks Capturing Decision Trails and Rationales in Patent Examination
di: Lim, Hyunseung, et al.
Pubblicazione: (2025)