Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Suau, Xavier, Delobelle, Pieter, Metcalf, Katherine, Joulin, Armand, Apostoloff, Nicholas, Zappella, Luca, Rodríguez, Pau |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Controlling Language and Diffusion Models by Transporting Activations
di: Rodriguez, Pau, et al.
Pubblicazione: (2024)
di: Rodriguez, Pau, et al.
Pubblicazione: (2024)
Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
di: Khan, Falaah Arif, et al.
Pubblicazione: (2025)
di: Khan, Falaah Arif, et al.
Pubblicazione: (2025)
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
di: Rodriguez, Pau, et al.
Pubblicazione: (2025)
di: Rodriguez, Pau, et al.
Pubblicazione: (2025)
HyperTransport: Amortized Conditioning of T2I Generative Models
di: Maiorca, Valentino, et al.
Pubblicazione: (2026)
di: Maiorca, Valentino, et al.
Pubblicazione: (2026)
Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
di: Mackraz, Natalie, et al.
Pubblicazione: (2024)
di: Mackraz, Natalie, et al.
Pubblicazione: (2024)
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
di: Wang, Yinong Oliver, et al.
Pubblicazione: (2025)
di: Wang, Yinong Oliver, et al.
Pubblicazione: (2025)
GenCtrl -- A Formal Controllability Toolkit for Generative Models
di: Cheng, Emily, et al.
Pubblicazione: (2026)
di: Cheng, Emily, et al.
Pubblicazione: (2026)
Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language Models
di: Yuan, Hongbang, et al.
Pubblicazione: (2024)
di: Yuan, Hongbang, et al.
Pubblicazione: (2024)
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
di: Danieli, Federico, et al.
Pubblicazione: (2025)
di: Danieli, Federico, et al.
Pubblicazione: (2025)
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
di: Sundar, Anirudh, et al.
Pubblicazione: (2025)
di: Sundar, Anirudh, et al.
Pubblicazione: (2025)
Dynamically Scaled Activation Steering
di: Ferrando, Alex, et al.
Pubblicazione: (2025)
di: Ferrando, Alex, et al.
Pubblicazione: (2025)
ExpertLens: Activation steering features are highly interpretable
di: Fedzechkina, Masha, et al.
Pubblicazione: (2025)
di: Fedzechkina, Masha, et al.
Pubblicazione: (2025)
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models
di: Pozzobon, Luiza, et al.
Pubblicazione: (2024)
di: Pozzobon, Luiza, et al.
Pubblicazione: (2024)
Whisper Finetuning on Nepali Language
di: Rijal, Sanjay, et al.
Pubblicazione: (2024)
di: Rijal, Sanjay, et al.
Pubblicazione: (2024)
Fairness Dynamics During Training
di: Patel, Krishna, et al.
Pubblicazione: (2025)
di: Patel, Krishna, et al.
Pubblicazione: (2025)
GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace
di: Duan, Zenghao, et al.
Pubblicazione: (2025)
di: Duan, Zenghao, et al.
Pubblicazione: (2025)
CoRet: Improved Retriever for Code Editing
di: Fehr, Fabio, et al.
Pubblicazione: (2025)
di: Fehr, Fabio, et al.
Pubblicazione: (2025)
Bias after Prompting: Persistent Discrimination in Large Language Models
di: Sivakumar, Nivedha, et al.
Pubblicazione: (2025)
di: Sivakumar, Nivedha, et al.
Pubblicazione: (2025)
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
di: Santilli, Andrea, et al.
Pubblicazione: (2025)
di: Santilli, Andrea, et al.
Pubblicazione: (2025)
How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models
di: Corbo, Simone, et al.
Pubblicazione: (2025)
di: Corbo, Simone, et al.
Pubblicazione: (2025)
Time Sensitive Knowledge Editing through Efficient Finetuning
di: Ge, Xiou, et al.
Pubblicazione: (2024)
di: Ge, Xiou, et al.
Pubblicazione: (2024)
Large Language Model Adaptation for Financial Sentiment Analysis
di: Inserte, Pau Rodriguez, et al.
Pubblicazione: (2024)
di: Inserte, Pau Rodriguez, et al.
Pubblicazione: (2024)
ProbLog4Fairness: A Neurosymbolic Approach to Modeling and Mitigating Bias
di: Adriaensen, Rik, et al.
Pubblicazione: (2025)
di: Adriaensen, Rik, et al.
Pubblicazione: (2025)
DSO: Direct Steering Optimization for Bias Mitigation
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2025)
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2025)
Characterising Toxicity in Generative Large Language Models
di: Zhang, Zhiyao, et al.
Pubblicazione: (2026)
di: Zhang, Zhiyao, et al.
Pubblicazione: (2026)
Realistic Evaluation of Toxicity in Large Language Models
di: Luong, Tinh Son, et al.
Pubblicazione: (2024)
di: Luong, Tinh Son, et al.
Pubblicazione: (2024)
Preference Tuning For Toxicity Mitigation Generalizes Across Languages
di: Li, Xiaochen, et al.
Pubblicazione: (2024)
di: Li, Xiaochen, et al.
Pubblicazione: (2024)
Following the Whispers of Values: Unraveling Neural Mechanisms Behind Value-Oriented Behaviors in LLMs
di: Hu, Ling, et al.
Pubblicazione: (2025)
di: Hu, Ling, et al.
Pubblicazione: (2025)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
di: Crabbé, Jonathan, et al.
Pubblicazione: (2023)
di: Crabbé, Jonathan, et al.
Pubblicazione: (2023)
ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
di: Xu, Benfeng, et al.
Pubblicazione: (2023)
di: Xu, Benfeng, et al.
Pubblicazione: (2023)
Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense Reasoning
di: Li, Jiachun, et al.
Pubblicazione: (2024)
di: Li, Jiachun, et al.
Pubblicazione: (2024)
Large Language Models as Generalizable Policies for Embodied Tasks
di: Szot, Andrew, et al.
Pubblicazione: (2023)
di: Szot, Andrew, et al.
Pubblicazione: (2023)
Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
di: Zhu, Shien, et al.
Pubblicazione: (2025)
di: Zhu, Shien, et al.
Pubblicazione: (2025)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
di: Assogba, Yannick, et al.
Pubblicazione: (2026)
di: Assogba, Yannick, et al.
Pubblicazione: (2026)
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions
di: Xu, Jingxin, et al.
Pubblicazione: (2025)
di: Xu, Jingxin, et al.
Pubblicazione: (2025)
Unplug and Play Language Models: Decomposing Experts in Language Models at Inference Time
di: Yang, Nakyeong, et al.
Pubblicazione: (2024)
di: Yang, Nakyeong, et al.
Pubblicazione: (2024)
Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective
di: Kou, Zhiqiang, et al.
Pubblicazione: (2025)
di: Kou, Zhiqiang, et al.
Pubblicazione: (2025)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
di: Wang, Shuxun, et al.
Pubblicazione: (2025)
di: Wang, Shuxun, et al.
Pubblicazione: (2025)
Pruning General Large Language Models into Customized Expert Models
di: Zhao, Yirao, et al.
Pubblicazione: (2025)
di: Zhao, Yirao, et al.
Pubblicazione: (2025)
Mix of Experts Language Model for Named Entity Recognition
di: Chen, Xinwei, et al.
Pubblicazione: (2024)
di: Chen, Xinwei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Controlling Language and Diffusion Models by Transporting Activations
di: Rodriguez, Pau, et al.
Pubblicazione: (2024) -
Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
di: Khan, Falaah Arif, et al.
Pubblicazione: (2025) -
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
di: Rodriguez, Pau, et al.
Pubblicazione: (2025) -
HyperTransport: Amortized Conditioning of T2I Generative Models
di: Maiorca, Valentino, et al.
Pubblicazione: (2026) -
Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
di: Mackraz, Natalie, et al.
Pubblicazione: (2024)