Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yilin, Wang, Heng, Bai, Yuyang, Luo, Minnan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Knowledge Card: Filling LLMs' Knowledge Gaps with Plug-in Specialized Language Models
von: Feng, Shangbin, et al.
Veröffentlicht: (2023)
von: Feng, Shangbin, et al.
Veröffentlicht: (2023)
ExpertSteer: Intervening in LLMs through Expert Knowledge
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations
von: Wang, Haoran, et al.
Veröffentlicht: (2026)
von: Wang, Haoran, et al.
Veröffentlicht: (2026)
Bot Meets Shortcut: How Can LLMs Aid in Handling Unknown Invariance OOD Scenarios?
von: Zheng, Shiyan, et al.
Veröffentlicht: (2025)
von: Zheng, Shiyan, et al.
Veröffentlicht: (2025)
DELL: Generating Reactions and Explanations for LLM-Based Misinformation Detection
von: Wan, Herun, et al.
Veröffentlicht: (2024)
von: Wan, Herun, et al.
Veröffentlicht: (2024)
Contextual Linear Activation Steering of Language Models
von: Hsu, Brandon, et al.
Veröffentlicht: (2026)
von: Hsu, Brandon, et al.
Veröffentlicht: (2026)
On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMs
von: Wan, Herun, et al.
Veröffentlicht: (2024)
von: Wan, Herun, et al.
Veröffentlicht: (2024)
Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
von: Wang, Linlin, et al.
Veröffentlicht: (2025)
von: Wang, Linlin, et al.
Veröffentlicht: (2025)
What Does the Bot Say? Opportunities and Risks of Large Language Models in Social Media Bot Detection
von: Feng, Shangbin, et al.
Veröffentlicht: (2024)
von: Feng, Shangbin, et al.
Veröffentlicht: (2024)
Steering LLMs for Culturally Localized Generation
von: Khanuja, Simran, et al.
Veröffentlicht: (2026)
von: Khanuja, Simran, et al.
Veröffentlicht: (2026)
Exploiting Contextual Knowledge in LLMs through V-usable Information based Layer Enhancement
von: Yuan, Xiaowei, et al.
Veröffentlicht: (2025)
von: Yuan, Xiaowei, et al.
Veröffentlicht: (2025)
LMBot: Distilling Graph Knowledge into Language Model for Graph-less Deployment in Twitter Bot Detection
von: Cai, Zijian, et al.
Veröffentlicht: (2023)
von: Cai, Zijian, et al.
Veröffentlicht: (2023)
Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models
von: Cheng, Sitao, et al.
Veröffentlicht: (2024)
von: Cheng, Sitao, et al.
Veröffentlicht: (2024)
Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models
von: Chrabąszcz, Maciej, et al.
Veröffentlicht: (2025)
von: Chrabąszcz, Maciej, et al.
Veröffentlicht: (2025)
Unveiling the Hidden: Movie Genre and User Bias in Spoiler Detection
von: Zhang, Haokai, et al.
Veröffentlicht: (2025)
von: Zhang, Haokai, et al.
Veröffentlicht: (2025)
KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and Unlearning
von: Luo, Yinyi, et al.
Veröffentlicht: (2025)
von: Luo, Yinyi, et al.
Veröffentlicht: (2025)
Tuning Language Models by Proxy
von: Liu, Alisa, et al.
Veröffentlicht: (2024)
von: Liu, Alisa, et al.
Veröffentlicht: (2024)
Speech LLMs are Contextual Reasoning Transcribers
von: Deng, Keqi, et al.
Veröffentlicht: (2026)
von: Deng, Keqi, et al.
Veröffentlicht: (2026)
The Anatomy of an Edit: Mechanism-Guided Activation Steering for Knowledge Editing
von: Cao, Yuan, et al.
Veröffentlicht: (2026)
von: Cao, Yuan, et al.
Veröffentlicht: (2026)
WisdoM: Improving Multimodal Sentiment Analysis by Fusing Contextual World Knowledge
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
Self-Improving Model Steering
von: Zhu, Rongyi, et al.
Veröffentlicht: (2025)
von: Zhu, Rongyi, et al.
Veröffentlicht: (2025)
HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring
von: Su, Zhixiong, et al.
Veröffentlicht: (2025)
von: Su, Zhixiong, et al.
Veröffentlicht: (2025)
SteerConf: Steering LLMs for Confidence Elicitation
von: Zhou, Ziang, et al.
Veröffentlicht: (2025)
von: Zhou, Ziang, et al.
Veröffentlicht: (2025)
Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors
von: Wang, Weixuan, et al.
Veröffentlicht: (2024)
von: Wang, Weixuan, et al.
Veröffentlicht: (2024)
GuessBench: Sensemaking Multimodal Creativity in the Wild
von: Zhu, Zifeng, et al.
Veröffentlicht: (2025)
von: Zhu, Zifeng, et al.
Veröffentlicht: (2025)
FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering
von: Li, Yichen, et al.
Veröffentlicht: (2025)
von: Li, Yichen, et al.
Veröffentlicht: (2025)
Steering When Necessary: Flexible Steering Large Language Models with Backtracking
von: Cheng, Zifeng, et al.
Veröffentlicht: (2025)
von: Cheng, Zifeng, et al.
Veröffentlicht: (2025)
KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language Models
von: Bai, Yuyang, et al.
Veröffentlicht: (2023)
von: Bai, Yuyang, et al.
Veröffentlicht: (2023)
Lifelong Knowledge Editing for LLMs with Retrieval-Augmented Continuous Prompt Learning
von: Chen, Qizhou, et al.
Veröffentlicht: (2024)
von: Chen, Qizhou, et al.
Veröffentlicht: (2024)
TABES: Trajectory-Aware Backward-on-Entropy Steering for Masked Diffusion Models
von: Saini, Shreshth, et al.
Veröffentlicht: (2026)
von: Saini, Shreshth, et al.
Veröffentlicht: (2026)
Contrastive Prompting Enhances Sentence Embeddings in LLMs through Inference-Time Steering
von: Cheng, Zifeng, et al.
Veröffentlicht: (2025)
von: Cheng, Zifeng, et al.
Veröffentlicht: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
That's Deprecated! Understanding, Detecting, and Steering Knowledge Conflicts in Language Models for Code Generation
von: Bae, Jaesung, et al.
Veröffentlicht: (2025)
von: Bae, Jaesung, et al.
Veröffentlicht: (2025)
Forecasting Downstream Performance of LLMs With Proxy Metrics
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
Infusing Knowledge into Large Language Models with Contextual Prompts
von: Vasisht, Kinshuk, et al.
Veröffentlicht: (2024)
von: Vasisht, Kinshuk, et al.
Veröffentlicht: (2024)
Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation Detection
von: Wan, Herun, et al.
Veröffentlicht: (2025)
von: Wan, Herun, et al.
Veröffentlicht: (2025)
Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' Memory
von: Tao, Xingjian, et al.
Veröffentlicht: (2024)
von: Tao, Xingjian, et al.
Veröffentlicht: (2024)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Knowledge Card: Filling LLMs' Knowledge Gaps with Plug-in Specialized Language Models
von: Feng, Shangbin, et al.
Veröffentlicht: (2023) -
ExpertSteer: Intervening in LLMs through Expert Knowledge
von: Wang, Weixuan, et al.
Veröffentlicht: (2025) -
Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations
von: Wang, Haoran, et al.
Veröffentlicht: (2026) -
Bot Meets Shortcut: How Can LLMs Aid in Handling Unknown Invariance OOD Scenarios?
von: Zheng, Shiyan, et al.
Veröffentlicht: (2025) -
DELL: Generating Reactions and Explanations for LLM-Based Misinformation Detection
von: Wan, Herun, et al.
Veröffentlicht: (2024)