Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Beniwal, Himanshu, Singh, Mayank |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Remember This Event That Year? Assessing Temporal Information and Reasoning in Large Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2024)
por: Beniwal, Himanshu, et al.
Publicado: (2024)
Cross-lingual Editing in Multilingual Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2024)
por: Beniwal, Himanshu, et al.
Publicado: (2024)
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
por: Sheth, Rajvee, et al.
Publicado: (2025)
por: Sheth, Rajvee, et al.
Publicado: (2025)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
por: Yadav, Ankit, et al.
Publicado: (2024)
por: Yadav, Ankit, et al.
Publicado: (2024)
Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
por: Beniwal, Himanshu, et al.
Publicado: (2025)
por: Beniwal, Himanshu, et al.
Publicado: (2025)
UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
por: Beniwal, Himanshu, et al.
Publicado: (2025)
por: Beniwal, Himanshu, et al.
Publicado: (2025)
COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework
por: Sheth, Rajvee, et al.
Publicado: (2024)
por: Sheth, Rajvee, et al.
Publicado: (2024)
One Instruction Does Not Fit All: How Well Do Embeddings Align Personas and Instructions in Low-Resource Indian Languages?
por: Shah, Arya, et al.
Publicado: (2026)
por: Shah, Arya, et al.
Publicado: (2026)
TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks
por: Kapoor, Vansh, et al.
Publicado: (2026)
por: Kapoor, Vansh, et al.
Publicado: (2026)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
por: Aravindan, Ashwath Vaithinathan, et al.
Publicado: (2026)
por: Aravindan, Ashwath Vaithinathan, et al.
Publicado: (2026)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
por: Cho, Hakaze, et al.
Publicado: (2025)
por: Cho, Hakaze, et al.
Publicado: (2025)
Error Taxonomy-Guided Prompt Optimization
por: Singh, Mayank, et al.
Publicado: (2026)
por: Singh, Mayank, et al.
Publicado: (2026)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
por: Kim, Geonhee, et al.
Publicado: (2024)
por: Kim, Geonhee, et al.
Publicado: (2024)
LOLAMEME: Logic, Language, Memory, Mechanistic Framework
por: Desai, Jay, et al.
Publicado: (2024)
por: Desai, Jay, et al.
Publicado: (2024)
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
por: Zhou, Hanhan, et al.
Publicado: (2026)
por: Zhou, Hanhan, et al.
Publicado: (2026)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
por: Wang, Xu, et al.
Publicado: (2026)
por: Wang, Xu, et al.
Publicado: (2026)
Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation
por: Balestriero, Randall, et al.
Publicado: (2023)
por: Balestriero, Randall, et al.
Publicado: (2023)
Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
por: Bozoukov, Matthew, et al.
Publicado: (2025)
por: Bozoukov, Matthew, et al.
Publicado: (2025)
Where Do Reasoning Models Refuse?
por: Yamaguchi, Kureha, et al.
Publicado: (2025)
por: Yamaguchi, Kureha, et al.
Publicado: (2025)
BiSup: Bidirectional Quantization Error Suppression for Large Language Models
por: Zou, Minghui, et al.
Publicado: (2024)
por: Zou, Minghui, et al.
Publicado: (2024)
Mechanistic?
por: Saphra, Naomi, et al.
Publicado: (2024)
por: Saphra, Naomi, et al.
Publicado: (2024)
Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression
por: Huang, Jiameng, et al.
Publicado: (2025)
por: Huang, Jiameng, et al.
Publicado: (2025)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
por: Yan, Tianyi Lorena, et al.
Publicado: (2025)
por: Yan, Tianyi Lorena, et al.
Publicado: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
por: Mishra, Anurag
Publicado: (2025)
por: Mishra, Anurag
Publicado: (2025)
Selective Self-Rehearsal: A Fine-Tuning Approach to Improve Generalization in Large Language Models
por: Gupta, Sonam, et al.
Publicado: (2024)
por: Gupta, Sonam, et al.
Publicado: (2024)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
por: Pan, Bowen, et al.
Publicado: (2024)
por: Pan, Bowen, et al.
Publicado: (2024)
How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
por: Verdini, Francesco, et al.
Publicado: (2024)
por: Verdini, Francesco, et al.
Publicado: (2024)
Multi-Attribute Steering of Language Models via Targeted Intervention
por: Nguyen, Duy, et al.
Publicado: (2025)
por: Nguyen, Duy, et al.
Publicado: (2025)
LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
por: Zou, Kaijian, et al.
Publicado: (2025)
por: Zou, Kaijian, et al.
Publicado: (2025)
Chem42: a Family of chemical Language Models for Target-aware Ligand Generation
por: Singh, Aahan, et al.
Publicado: (2025)
por: Singh, Aahan, et al.
Publicado: (2025)
Personas within Parameters: Fine-Tuning Small Language Models with Low-Rank Adapters to Mimic User Behaviors
por: Thakur, Himanshu, et al.
Publicado: (2025)
por: Thakur, Himanshu, et al.
Publicado: (2025)
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
por: Alzahrani, Norah, et al.
Publicado: (2024)
por: Alzahrani, Norah, et al.
Publicado: (2024)
Does Pre-trained Language Model Actually Infer Unseen Links in Knowledge Graph Completion?
por: Sakai, Yusuke, et al.
Publicado: (2023)
por: Sakai, Yusuke, et al.
Publicado: (2023)
Parallax: Parameterized Local Linear Attention for Language Modeling
por: Zuo, Yifei, et al.
Publicado: (2026)
por: Zuo, Yifei, et al.
Publicado: (2026)
Locally Coherent Parallel Decoding in Diffusion Language Models
por: Hersche, Michael, et al.
Publicado: (2026)
por: Hersche, Michael, et al.
Publicado: (2026)
PSK@EEUCA 2026: Fine-Tuning Large Language Models with Synthetic Data Augmentation for Multi-Class Toxicity Detection in Gaming Chat
por: Pulipaka, Srikar Kashyap
Publicado: (2026)
por: Pulipaka, Srikar Kashyap
Publicado: (2026)
Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits
por: Mann, Logan, et al.
Publicado: (2026)
por: Mann, Logan, et al.
Publicado: (2026)
Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework
por: Guan, Zihan, et al.
Publicado: (2026)
por: Guan, Zihan, et al.
Publicado: (2026)
Does Visual Rendering Bypass Tokenization? Investigating Script-Tokenizer Misalignment in Pixel-Based Language Models
por: Susanto, Lucky, et al.
Publicado: (2026)
por: Susanto, Lucky, et al.
Publicado: (2026)
When Does a Language Model Commit? A Finite-Answer Theory of Pre-Verbalization Commitment
por: Zhang, Long, et al.
Publicado: (2026)
por: Zhang, Long, et al.
Publicado: (2026)
Ejemplares similares
-
Remember This Event That Year? Assessing Temporal Information and Reasoning in Large Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2024) -
Cross-lingual Editing in Multilingual Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2024) -
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
por: Sheth, Rajvee, et al.
Publicado: (2025) -
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
por: Yadav, Ankit, et al.
Publicado: (2024) -
Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
por: Beniwal, Himanshu, et al.
Publicado: (2025)