Understanding and Mitigating Dataset Corruption in LLM Steering
Fuente:
arXiv
Guardado en:
| Autores principales: | Anderson, Cullen, Oozeer, Narmeen, Namjoo, Foad, Ogasawara, Remy, Abdullah, Amirali, Phillips, Jeff M. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
por: Oozeer, Narmeen, et al.
Publicado: (2025)
por: Oozeer, Narmeen, et al.
Publicado: (2025)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
por: Roytburg, Dani, et al.
Publicado: (2025)
por: Roytburg, Dani, et al.
Publicado: (2025)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
por: Roytburg, Dani, et al.
Publicado: (2026)
por: Roytburg, Dani, et al.
Publicado: (2026)
Approximating Human Preferences Using a Multi-Judge Learned System
por: Sprejer, Eitán, et al.
Publicado: (2025)
por: Sprejer, Eitán, et al.
Publicado: (2025)
Bilinear Convolution Decomposition for Causal RL Interpretability
por: Oozeer, Narmeen, et al.
Publicado: (2024)
por: Oozeer, Narmeen, et al.
Publicado: (2024)
Efficient and Stable Multi-Dimensional Kolmogorov-Smirnov Distance
por: Jacobs, Peter Matthew, et al.
Publicado: (2025)
por: Jacobs, Peter Matthew, et al.
Publicado: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
por: Heyman, Geert, et al.
Publicado: (2026)
por: Heyman, Geert, et al.
Publicado: (2026)
DreamReader: An Interpretability Toolkit for Text-to-Image Models
por: Prakash, Nirmalendu, et al.
Publicado: (2026)
por: Prakash, Nirmalendu, et al.
Publicado: (2026)
Mitigating Overthinking in Large Reasoning Models via Manifold Steering
por: Huang, Yao, et al.
Publicado: (2025)
por: Huang, Yao, et al.
Publicado: (2025)
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
por: Siddique, Zara, et al.
Publicado: (2025)
por: Siddique, Zara, et al.
Publicado: (2025)
Steer LLM Latents for Hallucination Detection
por: Park, Seongheon, et al.
Publicado: (2025)
por: Park, Seongheon, et al.
Publicado: (2025)
Spectral Superposition: A Theory of Feature Geometry
por: Ivanov, Georgi, et al.
Publicado: (2026)
por: Ivanov, Georgi, et al.
Publicado: (2026)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
por: Li, Miaomiao, et al.
Publicado: (2025)
por: Li, Miaomiao, et al.
Publicado: (2025)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
por: Zhang, Zhenyu, et al.
Publicado: (2025)
por: Zhang, Zhenyu, et al.
Publicado: (2025)
Brain-Grounded Axes for Reading and Steering LLM States
por: Andric, Sandro
Publicado: (2025)
por: Andric, Sandro
Publicado: (2025)
Distribution-Aware Feature Selection for SAEs
por: Oozeer, Narmeen, et al.
Publicado: (2025)
por: Oozeer, Narmeen, et al.
Publicado: (2025)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
por: Yu, Zhuohao, et al.
Publicado: (2025)
por: Yu, Zhuohao, et al.
Publicado: (2025)
That's Deprecated! Understanding, Detecting, and Steering Knowledge Conflicts in Language Models for Code Generation
por: Bae, Jaesung, et al.
Publicado: (2025)
por: Bae, Jaesung, et al.
Publicado: (2025)
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
por: Braun, Joschka
Publicado: (2026)
por: Braun, Joschka
Publicado: (2026)
HyperSteer: Activation Steering at Scale with Hypernetworks
por: Sun, Jiuding, et al.
Publicado: (2025)
por: Sun, Jiuding, et al.
Publicado: (2025)
Understanding and Mitigating the Uncertainty in Zero-Shot Translation
por: Wang, Wenxuan, et al.
Publicado: (2022)
por: Wang, Wenxuan, et al.
Publicado: (2022)
Understanding and Mitigating Tokenization Bias in Language Models
por: Phan, Buu, et al.
Publicado: (2024)
por: Phan, Buu, et al.
Publicado: (2024)
Compositional Steering of Large Language Models with Steering Tokens
por: Radevski, Gorjan, et al.
Publicado: (2026)
por: Radevski, Gorjan, et al.
Publicado: (2026)
Mitigating LLM Hallucinations via Conformal Abstention
por: Yadkori, Yasin Abbasi, et al.
Publicado: (2024)
por: Yadkori, Yasin Abbasi, et al.
Publicado: (2024)
On Mitigating Code LLM Hallucinations with API Documentation
por: Jain, Nihal, et al.
Publicado: (2024)
por: Jain, Nihal, et al.
Publicado: (2024)
FairFlow: Mitigating Dataset Biases through Undecided Learning
por: Cheng, Jiali, et al.
Publicado: (2025)
por: Cheng, Jiali, et al.
Publicado: (2025)
Directional Attractors in LLM Reasoning: How Similarity Retrieval Steers Iterative Summarization Based Reasoning
por: Tekin, Cagatay, et al.
Publicado: (2025)
por: Tekin, Cagatay, et al.
Publicado: (2025)
Sycophancy as compositions of Atomic Psychometric Traits
por: Jain, Shreyans, et al.
Publicado: (2025)
por: Jain, Shreyans, et al.
Publicado: (2025)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
por: Hua, Zhenglin, et al.
Publicado: (2025)
por: Hua, Zhenglin, et al.
Publicado: (2025)
Understanding LLM Embeddings for Regression
por: Tang, Eric, et al.
Publicado: (2024)
por: Tang, Eric, et al.
Publicado: (2024)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
por: Yang, Jinming, et al.
Publicado: (2026)
por: Yang, Jinming, et al.
Publicado: (2026)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
por: Siu, Vincent, et al.
Publicado: (2025)
por: Siu, Vincent, et al.
Publicado: (2025)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
por: Cheng, Stephen, et al.
Publicado: (2026)
por: Cheng, Stephen, et al.
Publicado: (2026)
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
por: Han, Pengrui, et al.
Publicado: (2026)
por: Han, Pengrui, et al.
Publicado: (2026)
Activation Space Interventions Can Be Transferred Between Large Language Models
por: Oozeer, Narmeen, et al.
Publicado: (2025)
por: Oozeer, Narmeen, et al.
Publicado: (2025)
Large Language Model Unlearning via Embedding-Corrupted Prompts
por: Liu, Chris Yuhao, et al.
Publicado: (2024)
por: Liu, Chris Yuhao, et al.
Publicado: (2024)
Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
por: Ethayarajh, Kawin, et al.
Publicado: (2021)
por: Ethayarajh, Kawin, et al.
Publicado: (2021)
Optimal Brain Iterative Merging: Mitigating Interference in LLM Merging
por: Wang, Zhixiang, et al.
Publicado: (2025)
por: Wang, Zhixiang, et al.
Publicado: (2025)
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
por: Sharma, Kartik, et al.
Publicado: (2026)
por: Sharma, Kartik, et al.
Publicado: (2026)
LLM Unlearning Without an Expert Curated Dataset
por: Zhu, Xiaoyuan, et al.
Publicado: (2025)
por: Zhu, Xiaoyuan, et al.
Publicado: (2025)
Ejemplares similares
-
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
por: Oozeer, Narmeen, et al.
Publicado: (2025) -
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
por: Roytburg, Dani, et al.
Publicado: (2025) -
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
por: Roytburg, Dani, et al.
Publicado: (2026) -
Approximating Human Preferences Using a Multi-Judge Learned System
por: Sprejer, Eitán, et al.
Publicado: (2025) -
Bilinear Convolution Decomposition for Causal RL Interpretability
por: Oozeer, Narmeen, et al.
Publicado: (2024)