SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
Fuente:
arXiv
Guardado en:
| Autores principales: | Batra, Shourya, Tillman, Pierce, Gaggar, Samarth, Kesineni, Shashank, Zhu, Kevin, Dev, Sunishchal, Panda, Ashwinee, Sharma, Vasu, Chaudhary, Maheep |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
por: Egbuna, Nathan, et al.
Publicado: (2025)
por: Egbuna, Nathan, et al.
Publicado: (2025)
FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
por: Swaroop, Anand, et al.
Publicado: (2025)
por: Swaroop, Anand, et al.
Publicado: (2025)
Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
por: Chaturvedi, Isha, et al.
Publicado: (2025)
por: Chaturvedi, Isha, et al.
Publicado: (2025)
Evaluation Awareness Scales Predictably in Open-Weights Large Language Models
por: Chaudhary, Maheep, et al.
Publicado: (2025)
por: Chaudhary, Maheep, et al.
Publicado: (2025)
Broken Chains: The Cost of Incomplete Reasoning in LLMs
por: Su, Ian, et al.
Publicado: (2026)
por: Su, Ian, et al.
Publicado: (2026)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
por: Patel, Dev, et al.
Publicado: (2025)
por: Patel, Dev, et al.
Publicado: (2025)
Hydra: A Modular Architecture for Efficient Long-Context Reasoning
por: Chaudhary, Siddharth, et al.
Publicado: (2025)
por: Chaudhary, Siddharth, et al.
Publicado: (2025)
In-Context Environments Induce Evaluation-Awareness in Language Models
por: Chaudhary, Maheep
Publicado: (2026)
por: Chaudhary, Maheep
Publicado: (2026)
PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
por: Vuddanti, Sri Vatsa, et al.
Publicado: (2025)
por: Vuddanti, Sri Vatsa, et al.
Publicado: (2025)
Modeling and Predicting Multi-Turn Answer Instability in Large Language Models
por: He, Jiahang, et al.
Publicado: (2025)
por: He, Jiahang, et al.
Publicado: (2025)
Activation Steering for Chain-of-Thought Compression
por: Azizi, Seyedarmin, et al.
Publicado: (2025)
por: Azizi, Seyedarmin, et al.
Publicado: (2025)
Inference-Time Chain-of-Thought Pruning with Latent Informativeness Signals
por: Li, Sophie, et al.
Publicado: (2025)
por: Li, Sophie, et al.
Publicado: (2025)
Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
por: More, Abhishek, et al.
Publicado: (2025)
por: More, Abhishek, et al.
Publicado: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
por: Chaudhary, Maheep, et al.
Publicado: (2025)
por: Chaudhary, Maheep, et al.
Publicado: (2025)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
por: Chaudhary, Maheep, et al.
Publicado: (2024)
por: Chaudhary, Maheep, et al.
Publicado: (2024)
DuoLens: A Framework for Robust Detection of Machine-Generated Multilingual Text and Code
por: Agrawal, Shriyansh, et al.
Publicado: (2025)
por: Agrawal, Shriyansh, et al.
Publicado: (2025)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
por: Thomas, Rohan Subramanian, et al.
Publicado: (2026)
por: Thomas, Rohan Subramanian, et al.
Publicado: (2026)
Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
por: Nunez, Jeanmely Rojas, et al.
Publicado: (2026)
por: Nunez, Jeanmely Rojas, et al.
Publicado: (2026)
A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy
por: O'Brien, Claire, et al.
Publicado: (2026)
por: O'Brien, Claire, et al.
Publicado: (2026)
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
por: Zaman, Kerem, et al.
Publicado: (2025)
por: Zaman, Kerem, et al.
Publicado: (2025)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
por: Alagharu, Rishab, et al.
Publicado: (2026)
por: Alagharu, Rishab, et al.
Publicado: (2026)
COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
por: Sahay, Kenji, et al.
Publicado: (2025)
por: Sahay, Kenji, et al.
Publicado: (2025)
AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
por: Rana, Manik, et al.
Publicado: (2025)
por: Rana, Manik, et al.
Publicado: (2025)
@GrokSet: multi-party Human-LLM Interactions in Social Media
por: Migliarini, Matteo, et al.
Publicado: (2026)
por: Migliarini, Matteo, et al.
Publicado: (2026)
Freespace twistronics for optical supertopologies
por: Dev, Vasu, et al.
Publicado: (2025)
por: Dev, Vasu, et al.
Publicado: (2025)
Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
por: Afonin, Nikita, et al.
Publicado: (2025)
por: Afonin, Nikita, et al.
Publicado: (2025)
Decoding Answers Before Chain-of-Thought: Evidence from Pre-CoT Probes and Activation Steering
por: Cox, Kyle, et al.
Publicado: (2026)
por: Cox, Kyle, et al.
Publicado: (2026)
Data Augmentation for NeRFs in the Low Data Limit
por: Gaggar, Ayush, et al.
Publicado: (2025)
por: Gaggar, Ayush, et al.
Publicado: (2025)
GeoSteer: Faithful Chain-of-Thought Steering via Latent Manifold Gradients
por: Kazama, Kentaro, et al.
Publicado: (2026)
por: Kazama, Kentaro, et al.
Publicado: (2026)
Evaluating Chain-of-Thought Reasoning through Reusability and Verifiability
por: Aggarwal, Shashank, et al.
Publicado: (2026)
por: Aggarwal, Shashank, et al.
Publicado: (2026)
Melody or Machine: Detecting Synthetic Music with Dual-Stream Contrastive Learning
por: Batra, Arnesh, et al.
Publicado: (2025)
por: Batra, Arnesh, et al.
Publicado: (2025)
Deep Thinking by Markov Chain of Continuous Thoughts
por: Liu, Jiayu, et al.
Publicado: (2025)
por: Liu, Jiayu, et al.
Publicado: (2025)
Safer Reasoning Traces: Measuring and Mitigating Chain-of-Thought Leakage in LLMs
por: Ahrend, Patrick, et al.
Publicado: (2026)
por: Ahrend, Patrick, et al.
Publicado: (2026)
PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage
por: Nakka, Krishna Kanth, et al.
Publicado: (2025)
por: Nakka, Krishna Kanth, et al.
Publicado: (2025)
How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
por: Chen, Xi, et al.
Publicado: (2025)
por: Chen, Xi, et al.
Publicado: (2025)
LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
por: Zhang, Juzheng, et al.
Publicado: (2025)
por: Zhang, Juzheng, et al.
Publicado: (2025)
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
por: Shah, Siddhant Bikram, et al.
Publicado: (2024)
por: Shah, Siddhant Bikram, et al.
Publicado: (2024)
Propagation of circular Airy derivative beams in complex media
por: Kumari, Anita, et al.
Publicado: (2024)
por: Kumari, Anita, et al.
Publicado: (2024)
Chain-of-Sanitized-Thoughts: Plugging PII Leakage in CoT of Large Reasoning Models
por: Das, Arghyadeep, et al.
Publicado: (2026)
por: Das, Arghyadeep, et al.
Publicado: (2026)
Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow
por: Moore, Kyle, et al.
Publicado: (2024)
por: Moore, Kyle, et al.
Publicado: (2024)
Ejemplares similares
-
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
por: Egbuna, Nathan, et al.
Publicado: (2025) -
FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
por: Swaroop, Anand, et al.
Publicado: (2025) -
Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
por: Chaturvedi, Isha, et al.
Publicado: (2025) -
Evaluation Awareness Scales Predictably in Open-Weights Large Language Models
por: Chaudhary, Maheep, et al.
Publicado: (2025) -
Broken Chains: The Cost of Incomplete Reasoning in LLMs
por: Su, Ian, et al.
Publicado: (2026)