Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Yu, Devoto, Alessio, Hong, Giwon, Du, Xiaotang, Gema, Aryo Pradipta, Wang, Hongru, He, Xuanli, Wong, Kam-Fai, Minervini, Pasquale |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Analysing the Residual Stream of Language Models Under Knowledge Conflicts
di: Zhao, Yu, et al.
Pubblicazione: (2024)
di: Zhao, Yu, et al.
Pubblicazione: (2024)
GRADA: Graph-based Reranking against Adversarial Documents Attack
di: Zheng, Jingjie, et al.
Pubblicazione: (2025)
di: Zheng, Jingjie, et al.
Pubblicazione: (2025)
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
di: Saxena, Rohit, et al.
Pubblicazione: (2025)
di: Saxena, Rohit, et al.
Pubblicazione: (2025)
The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models
di: Hong, Giwon, et al.
Pubblicazione: (2024)
di: Hong, Giwon, et al.
Pubblicazione: (2024)
Self-Training Large Language Models for Tool-Use Without Demonstrations
di: Luo, Ne, et al.
Pubblicazione: (2025)
di: Luo, Ne, et al.
Pubblicazione: (2025)
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2023)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2023)
Are We Done with MMLU?
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks
di: Kwan, Wai-Chung, et al.
Pubblicazione: (2026)
di: Kwan, Wai-Chung, et al.
Pubblicazione: (2026)
Edinburgh Clinical NLP at MEDIQA-CORR 2024: Guiding Large Language Models with Hints
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
di: Leang, Joshua Ong Jun, et al.
Pubblicazione: (2025)
di: Leang, Joshua Ong Jun, et al.
Pubblicazione: (2025)
Analyzing LLM Instruction Optimization for Tabular Fact Verification
di: Du, Xiaotang, et al.
Pubblicazione: (2026)
di: Du, Xiaotang, et al.
Pubblicazione: (2026)
Noiser: Bounded Input Perturbations for Attributing Large Language Models
di: Madani, Mohammad Reza Ghasemi, et al.
Pubblicazione: (2025)
di: Madani, Mohammad Reza Ghasemi, et al.
Pubblicazione: (2025)
A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression
di: Devoto, Alessio, et al.
Pubblicazione: (2024)
di: Devoto, Alessio, et al.
Pubblicazione: (2024)
DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinations
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
di: Murphy, Alexander, et al.
Pubblicazione: (2025)
di: Murphy, Alexander, et al.
Pubblicazione: (2025)
Same Answer, Different Representations: Hidden instability in VLMs
di: Wani, Farooq Ahmad, et al.
Pubblicazione: (2026)
di: Wani, Farooq Ahmad, et al.
Pubblicazione: (2026)
An Auditing Test To Detect Behavioral Shift in Language Models
di: Richter, Leo, et al.
Pubblicazione: (2024)
di: Richter, Leo, et al.
Pubblicazione: (2024)
Adaptive Layer Selection for Efficient Vision Transformer Fine-Tuning
di: Devoto, Alessio, et al.
Pubblicazione: (2024)
di: Devoto, Alessio, et al.
Pubblicazione: (2024)
Adaptive Computation Modules: Granular Conditional Computation For Efficient Inference
di: Wójcik, Bartosz, et al.
Pubblicazione: (2023)
di: Wójcik, Bartosz, et al.
Pubblicazione: (2023)
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
di: Shilov, Igor, et al.
Pubblicazione: (2025)
di: Shilov, Igor, et al.
Pubblicazione: (2025)
Mixtures of In-Context Learners
di: Hong, Giwon, et al.
Pubblicazione: (2024)
di: Hong, Giwon, et al.
Pubblicazione: (2024)
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning
di: Leang, Joshua Ong Jun, et al.
Pubblicazione: (2024)
di: Leang, Joshua Ong Jun, et al.
Pubblicazione: (2024)
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
di: Rajani, Neel, et al.
Pubblicazione: (2025)
di: Rajani, Neel, et al.
Pubblicazione: (2025)
Enhancing Long Document Long Form Summarisation with Self-Planning
di: Du, Xiaotang, et al.
Pubblicazione: (2025)
di: Du, Xiaotang, et al.
Pubblicazione: (2025)
Using Natural Language Explanations to Improve Robustness of In-context Learning
di: He, Xuanli, et al.
Pubblicazione: (2023)
di: He, Xuanli, et al.
Pubblicazione: (2023)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
di: Cao, Jingtao, et al.
Pubblicazione: (2024)
di: Cao, Jingtao, et al.
Pubblicazione: (2024)
Conditional computation in neural networks: principles and research trends
di: Scardapane, Simone, et al.
Pubblicazione: (2024)
di: Scardapane, Simone, et al.
Pubblicazione: (2024)
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
di: He, Xuanli, et al.
Pubblicazione: (2024)
di: He, Xuanli, et al.
Pubblicazione: (2024)
SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models
di: He, Zirui, et al.
Pubblicazione: (2025)
di: He, Zirui, et al.
Pubblicazione: (2025)
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
di: Hägele, Alexander, et al.
Pubblicazione: (2026)
di: Hägele, Alexander, et al.
Pubblicazione: (2026)
Compositional Steering of Large Language Models with Steering Tokens
di: Radevski, Gorjan, et al.
Pubblicazione: (2026)
di: Radevski, Gorjan, et al.
Pubblicazione: (2026)
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
di: Attimonelli, Matteo, et al.
Pubblicazione: (2026)
di: Attimonelli, Matteo, et al.
Pubblicazione: (2026)
A Survey of the Evolution of Language Model-Based Dialogue Systems: Data, Task and Models
di: Wang, Hongru, et al.
Pubblicazione: (2023)
di: Wang, Hongru, et al.
Pubblicazione: (2023)
Universal Properties of Activation Sparsity in Modern Large Language Models
di: Szatkowski, Filip, et al.
Pubblicazione: (2025)
di: Szatkowski, Filip, et al.
Pubblicazione: (2025)
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
di: Godey, Nathan, et al.
Pubblicazione: (2025)
di: Godey, Nathan, et al.
Pubblicazione: (2025)
HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds
di: Wu, Honghan, et al.
Pubblicazione: (2026)
di: Wu, Honghan, et al.
Pubblicazione: (2026)
UniRetriever: Multi-task Candidates Selection for Various Context-Adaptive Conversational Retrieval
di: Wang, Hongru, et al.
Pubblicazione: (2024)
di: Wang, Hongru, et al.
Pubblicazione: (2024)
OSPC: Detecting Harmful Memes with Large Language Model as a Catalyst
di: Cao, Jingtao, et al.
Pubblicazione: (2024)
di: Cao, Jingtao, et al.
Pubblicazione: (2024)
Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human Perception
di: Lin, Luyang, et al.
Pubblicazione: (2024)
di: Lin, Luyang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Analysing the Residual Stream of Language Models Under Knowledge Conflicts
di: Zhao, Yu, et al.
Pubblicazione: (2024) -
GRADA: Graph-based Reranking against Adversarial Documents Attack
di: Zheng, Jingjie, et al.
Pubblicazione: (2025) -
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024) -
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
di: Saxena, Rohit, et al.
Pubblicazione: (2025) -
The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models
di: Hong, Giwon, et al.
Pubblicazione: (2024)