Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Goyal, Agam, Rathi, Vedant, Yeh, William, Wang, Yian, Chen, Yuen, Sundaram, Hari |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification
von: Wang, Yian, et al.
Veröffentlicht: (2026)
von: Wang, Yian, et al.
Veröffentlicht: (2026)
From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities
von: Goyal, Agam, et al.
Veröffentlicht: (2026)
von: Goyal, Agam, et al.
Veröffentlicht: (2026)
State Contamination in Memory-Augmented LLM Agents
von: Wang, Yian, et al.
Veröffentlicht: (2026)
von: Wang, Yian, et al.
Veröffentlicht: (2026)
Social Simulacra in the Wild: AI Agent Communities on Moltbook
von: Goyal, Agam, et al.
Veröffentlicht: (2026)
von: Goyal, Agam, et al.
Veröffentlicht: (2026)
Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?
von: Lin, Fei, et al.
Veröffentlicht: (2025)
von: Lin, Fei, et al.
Veröffentlicht: (2025)
Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
von: Shimgekar, Soorya Ram, et al.
Veröffentlicht: (2026)
von: Shimgekar, Soorya Ram, et al.
Veröffentlicht: (2026)
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
DIVEBATCH: Accelerating Model Training Through Gradient-Diversity Aware Batch Size Adaptation
von: Chen, Yuen, et al.
Veröffentlicht: (2025)
von: Chen, Yuen, et al.
Veröffentlicht: (2025)
SLM-Mod: Small Language Models Surpass LLMs at Content Moderation
von: Zhan, Xianyang, et al.
Veröffentlicht: (2024)
von: Zhan, Xianyang, et al.
Veröffentlicht: (2024)
Breaking Token Into Concepts: Exploring Extreme Compression in Token Representation Via Compositional Shared Semantics
von: R V, Kavin, et al.
Veröffentlicht: (2025)
von: R V, Kavin, et al.
Veröffentlicht: (2025)
ArgCMV: An Argument Summarization Benchmark for the LLM-era
von: Gurjar, Omkar, et al.
Veröffentlicht: (2025)
von: Gurjar, Omkar, et al.
Veröffentlicht: (2025)
Detecting Early and Implicit Suicidal Ideation via Longitudinal and Information Environment Signals on Social Media
von: Shimgekar, Soorya Ram, et al.
Veröffentlicht: (2025)
von: Shimgekar, Soorya Ram, et al.
Veröffentlicht: (2025)
MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governance
von: Goyal, Agam, et al.
Veröffentlicht: (2025)
von: Goyal, Agam, et al.
Veröffentlicht: (2025)
SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
von: Härle, Ruben, et al.
Veröffentlicht: (2024)
von: Härle, Ruben, et al.
Veröffentlicht: (2024)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025)
von: Chanin, David, et al.
Veröffentlicht: (2025)
Answer Bubbles: Information Exposure in AI-Mediated Search
von: Huang, Michelle, et al.
Veröffentlicht: (2026)
von: Huang, Michelle, et al.
Veröffentlicht: (2026)
On the Robustness of Knowledge Editing for Detoxification
von: Dong, Ming, et al.
Veröffentlicht: (2026)
von: Dong, Ming, et al.
Veröffentlicht: (2026)
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
von: Wu, Xuansheng, et al.
Veröffentlicht: (2025)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2025)
Cross-Lingual Transfer of Debiasing and Detoxification in Multilingual LLMs: An Extensive Investigation
von: Neplenbroek, Vera, et al.
Veröffentlicht: (2024)
von: Neplenbroek, Vera, et al.
Veröffentlicht: (2024)
CEV-LM: Controlled Edit Vector Language Model for Shaping Natural Language Generations
von: Moorjani, Samraj, et al.
Veröffentlicht: (2024)
von: Moorjani, Samraj, et al.
Veröffentlicht: (2024)
Optical Context Compression Is Just (Bad) Autoencoding
von: Lee, Ivan Yee, et al.
Veröffentlicht: (2025)
von: Lee, Ivan Yee, et al.
Veröffentlicht: (2025)
SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs
von: Deng, Boyi, et al.
Veröffentlicht: (2025)
von: Deng, Boyi, et al.
Veröffentlicht: (2025)
Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders
von: Xuan, Richmond Sin Jing, et al.
Veröffentlicht: (2025)
von: Xuan, Richmond Sin Jing, et al.
Veröffentlicht: (2025)
Attribute Controlled Fine-tuning for Large Language Models: A Case Study on Detoxification
von: Meng, Tao, et al.
Veröffentlicht: (2024)
von: Meng, Tao, et al.
Veröffentlicht: (2024)
ylmmcl at Multilingual Text Detoxification 2025: Lexicon-Guided Detoxification and Classifier-Gated Rewriting
von: Lai-Lopez, Nicole, et al.
Veröffentlicht: (2025)
von: Lai-Lopez, Nicole, et al.
Veröffentlicht: (2025)
Detoxification for LLM: From Dataset Itself
von: Shao, Wei, et al.
Veröffentlicht: (2026)
von: Shao, Wei, et al.
Veröffentlicht: (2026)
Algorithmic Cultivation: How Social Media Feeds Shape User Language
von: Pal, Olivia, et al.
Veröffentlicht: (2026)
von: Pal, Olivia, et al.
Veröffentlicht: (2026)
The Hidden Toll of Social Media News: Causal Effects on Psychosocial Wellbeing
von: Pal, Olivia, et al.
Veröffentlicht: (2026)
von: Pal, Olivia, et al.
Veröffentlicht: (2026)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators
von: Moskovskiy, Daniil, et al.
Veröffentlicht: (2025)
von: Moskovskiy, Daniil, et al.
Veröffentlicht: (2025)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
von: Zhang, Ruikang, et al.
Veröffentlicht: (2026)
von: Zhang, Ruikang, et al.
Veröffentlicht: (2026)
Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders
von: Bail, Mathis Le, et al.
Veröffentlicht: (2025)
von: Bail, Mathis Le, et al.
Veröffentlicht: (2025)
SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
Constrain Alignment with Sparse Autoencoders
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
Sparse Autoencoder Insights on Voice Embeddings
von: Pluth, Daniel, et al.
Veröffentlicht: (2025)
von: Pluth, Daniel, et al.
Veröffentlicht: (2025)
Sparse Autoencoders for Hypothesis Generation
von: Movva, Rajiv, et al.
Veröffentlicht: (2025)
von: Movva, Rajiv, et al.
Veröffentlicht: (2025)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
von: Jørgensen, Mikkel Godsk, et al.
Veröffentlicht: (2026)
von: Jørgensen, Mikkel Godsk, et al.
Veröffentlicht: (2026)
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
von: Yi, Zihao, et al.
Veröffentlicht: (2025)
von: Yi, Zihao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification
von: Wang, Yian, et al.
Veröffentlicht: (2026) -
From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities
von: Goyal, Agam, et al.
Veröffentlicht: (2026) -
State Contamination in Memory-Augmented LLM Agents
von: Wang, Yian, et al.
Veröffentlicht: (2026) -
Social Simulacra in the Wild: AI Agent Communities on Moltbook
von: Goyal, Agam, et al.
Veröffentlicht: (2026) -
Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?
von: Lin, Fei, et al.
Veröffentlicht: (2025)