ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Anand, Nikhil, Somasundaram, Shwetha, Phukan, Anirudh, Saxena, Apoorv, Mukherjee, Koyel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PLD+: Accelerating LLM inference by leveraging Language Model Artifacts
von: Somasundaram, Shwetha, et al.
Veröffentlicht: (2024)
von: Somasundaram, Shwetha, et al.
Veröffentlicht: (2024)
Peering into the Mind of Language Models: An Approach for Attribution in Contextual Question Answering
von: Phukan, Anirudh, et al.
Veröffentlicht: (2024)
von: Phukan, Anirudh, et al.
Veröffentlicht: (2024)
Towards Optimizing the Costs of LLM Usage
von: Shekhar, Shivanshu, et al.
Veröffentlicht: (2024)
von: Shekhar, Shivanshu, et al.
Veröffentlicht: (2024)
RCStat: A Statistical Framework for using Relative Contextualization in Transformers
von: Mahapatra, Debabrata, et al.
Veröffentlicht: (2025)
von: Mahapatra, Debabrata, et al.
Veröffentlicht: (2025)
PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization
von: Jaisankar, Vijay, et al.
Veröffentlicht: (2024)
von: Jaisankar, Vijay, et al.
Veröffentlicht: (2024)
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
von: Sharma, Kartik, et al.
Veröffentlicht: (2026)
von: Sharma, Kartik, et al.
Veröffentlicht: (2026)
Multi-property Steering of Large Language Models with Dynamic Activation Composition
von: Scalena, Daniel, et al.
Veröffentlicht: (2024)
von: Scalena, Daniel, et al.
Veröffentlicht: (2024)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
FaithLM: Towards Faithful Explanations for Large Language Models
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
Endogenous Resistance to Activation Steering in Language Models
von: McKenzie, Alex, et al.
Veröffentlicht: (2026)
von: McKenzie, Alex, et al.
Veröffentlicht: (2026)
Compositional Steering of Large Language Models with Steering Tokens
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
Listen to the Context: Towards Faithful Large Language Models for Retrieval Augmented Generation on Climate Questions
von: Thulke, David, et al.
Veröffentlicht: (2025)
von: Thulke, David, et al.
Veröffentlicht: (2025)
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
Improving Instruction-Following in Language Models through Activation Steering
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
von: Bigelow, Eric, et al.
Veröffentlicht: (2025)
von: Bigelow, Eric, et al.
Veröffentlicht: (2025)
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
von: Sundar, Anirudh, et al.
Veröffentlicht: (2025)
von: Sundar, Anirudh, et al.
Veröffentlicht: (2025)
HyperSteer: Activation Steering at Scale with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
von: Matton, Katie, et al.
Veröffentlicht: (2025)
von: Matton, Katie, et al.
Veröffentlicht: (2025)
Steering Large Language Models for Machine Translation Personalization
von: Scalena, Daniel, et al.
Veröffentlicht: (2025)
von: Scalena, Daniel, et al.
Veröffentlicht: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
von: Lamb, Tom A., et al.
Veröffentlicht: (2024)
von: Lamb, Tom A., et al.
Veröffentlicht: (2024)
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
von: Morlat, Geoffroy, et al.
Veröffentlicht: (2025)
von: Morlat, Geoffroy, et al.
Veröffentlicht: (2025)
FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models
von: Weng, Zixuan, et al.
Veröffentlicht: (2026)
von: Weng, Zixuan, et al.
Veröffentlicht: (2026)
Programming Refusal with Conditional Activation Steering
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
SAKE: Steering Activations for Knowledge Editing
von: Scialanga, Marco, et al.
Veröffentlicht: (2025)
von: Scialanga, Marco, et al.
Veröffentlicht: (2025)
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
von: Sun, Jingyi, et al.
Veröffentlicht: (2026)
von: Sun, Jingyi, et al.
Veröffentlicht: (2026)
Steering Large Language Model Activations in Sparse Spaces
von: Bayat, Reza, et al.
Veröffentlicht: (2025)
von: Bayat, Reza, et al.
Veröffentlicht: (2025)
Word Embeddings Are Steers for Language Models
von: Han, Chi, et al.
Veröffentlicht: (2023)
von: Han, Chi, et al.
Veröffentlicht: (2023)
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
von: Achintalwar, Swapnaja, et al.
Veröffentlicht: (2024)
von: Achintalwar, Swapnaja, et al.
Veröffentlicht: (2024)
A Data-Centric Approach To Generate Faithful and High Quality Patient Summaries with Large Language Models
von: Hegselmann, Stefan, et al.
Veröffentlicht: (2024)
von: Hegselmann, Stefan, et al.
Veröffentlicht: (2024)
Extracting Unlearned Information from LLMs with Activation Steering
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
Steering Llama 2 via Contrastive Activation Addition
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
When Should Models Change Their Minds? Contextual Belief Management in Large Language Models
von: Xu, Haoming, et al.
Veröffentlicht: (2026)
von: Xu, Haoming, et al.
Veröffentlicht: (2026)
Large Language Models are Miscalibrated In-Context Learners
von: Li, Chengzu, et al.
Veröffentlicht: (2023)
von: Li, Chengzu, et al.
Veröffentlicht: (2023)
PrivacyMind: Large Language Models Can Be Contextual Privacy Protection Learners
von: Xiao, Yijia, et al.
Veröffentlicht: (2023)
von: Xiao, Yijia, et al.
Veröffentlicht: (2023)
Spectral Editing of Activations for Large Language Model Alignment
von: Qiu, Yifu, et al.
Veröffentlicht: (2024)
von: Qiu, Yifu, et al.
Veröffentlicht: (2024)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
Polynomial Composition Activations: Unleashing the Dynamics of Large Language Models
von: Zhuo, Zhijian, et al.
Veröffentlicht: (2024)
von: Zhuo, Zhijian, et al.
Veröffentlicht: (2024)
The Role of Diversity in In-Context Learning for Large Language Models
von: Xiao, Wenyang, et al.
Veröffentlicht: (2025)
von: Xiao, Wenyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PLD+: Accelerating LLM inference by leveraging Language Model Artifacts
von: Somasundaram, Shwetha, et al.
Veröffentlicht: (2024) -
Peering into the Mind of Language Models: An Approach for Attribution in Contextual Question Answering
von: Phukan, Anirudh, et al.
Veröffentlicht: (2024) -
Towards Optimizing the Costs of LLM Usage
von: Shekhar, Shivanshu, et al.
Veröffentlicht: (2024) -
RCStat: A Statistical Framework for using Relative Contextualization in Transformers
von: Mahapatra, Debabrata, et al.
Veröffentlicht: (2025) -
PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization
von: Jaisankar, Vijay, et al.
Veröffentlicht: (2024)