Multi-property Steering of Large Language Models with Dynamic Activation Composition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Scalena, Daniel, Sarti, Gabriele, Nissim, Malvina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Steering Large Language Models for Machine Translation Personalization
von: Scalena, Daniel, et al.
Veröffentlicht: (2025)
von: Scalena, Daniel, et al.
Veröffentlicht: (2025)
A gentle push funziona benissimo: making instructed models in Italian via contrastive activation steering
von: Scalena, Daniel, et al.
Veröffentlicht: (2024)
von: Scalena, Daniel, et al.
Veröffentlicht: (2024)
Non Verbis, Sed Rebus: Large Language Models are Weak Solvers of Italian Rebuses
von: Sarti, Gabriele, et al.
Veröffentlicht: (2024)
von: Sarti, Gabriele, et al.
Veröffentlicht: (2024)
Quantifying the Plausibility of Context Reliance in Neural Machine Translation
von: Sarti, Gabriele, et al.
Veröffentlicht: (2023)
von: Sarti, Gabriele, et al.
Veröffentlicht: (2023)
Compositional Steering of Large Language Models with Steering Tokens
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
von: Sarti, Gabriele, et al.
Veröffentlicht: (2025)
von: Sarti, Gabriele, et al.
Veröffentlicht: (2025)
IT5: Text-to-text Pretraining for Italian Language Understanding and Generation
von: Sarti, Gabriele, et al.
Veröffentlicht: (2022)
von: Sarti, Gabriele, et al.
Veröffentlicht: (2022)
Polynomial Composition Activations: Unleashing the Dynamics of Large Language Models
von: Zhuo, Zhijian, et al.
Veröffentlicht: (2024)
von: Zhuo, Zhijian, et al.
Veröffentlicht: (2024)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
Endogenous Resistance to Activation Steering in Language Models
von: McKenzie, Alex, et al.
Veröffentlicht: (2026)
von: McKenzie, Alex, et al.
Veröffentlicht: (2026)
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
von: Sharma, Kartik, et al.
Veröffentlicht: (2026)
von: Sharma, Kartik, et al.
Veröffentlicht: (2026)
Improving Instruction-Following in Language Models through Activation Steering
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
von: Bigelow, Eric, et al.
Veröffentlicht: (2025)
von: Bigelow, Eric, et al.
Veröffentlicht: (2025)
HyperSteer: Activation Steering at Scale with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
TACLer: Tailored Curriculum Reinforcement Learning for Efficient Reasoning
von: Lai, Huiyuan, et al.
Veröffentlicht: (2026)
von: Lai, Huiyuan, et al.
Veröffentlicht: (2026)
Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation
von: Qi, Jirui, et al.
Veröffentlicht: (2024)
von: Qi, Jirui, et al.
Veröffentlicht: (2024)
Steer Like the LLM: Activation Steering that Mimics Prompting
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models
von: Weng, Zixuan, et al.
Veröffentlicht: (2026)
von: Weng, Zixuan, et al.
Veröffentlicht: (2026)
Multi-Attribute Steering of Language Models via Targeted Intervention
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
On-device System of Compositional Multi-tasking in Large Language Models
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2025)
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2025)
Efficient Compositional Multi-tasking for On-device Large Language Models
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2025)
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2025)
Programming Refusal with Conditional Activation Steering
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
SAKE: Steering Activations for Knowledge Editing
von: Scialanga, Marco, et al.
Veröffentlicht: (2025)
von: Scialanga, Marco, et al.
Veröffentlicht: (2025)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
Word Embeddings Are Steers for Language Models
von: Han, Chi, et al.
Veröffentlicht: (2023)
von: Han, Chi, et al.
Veröffentlicht: (2023)
Extracting Unlearned Information from LLMs with Activation Steering
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
Steering Llama 2 via Contrastive Activation Addition
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
Spectral Editing of Activations for Large Language Model Alignment
von: Qiu, Yifu, et al.
Veröffentlicht: (2024)
von: Qiu, Yifu, et al.
Veröffentlicht: (2024)
Understanding Subword Compositionality of Large Language Models
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
The Compositional Architecture of Regret in Large Language Models
von: Cui, Xiangxiang, et al.
Veröffentlicht: (2025)
von: Cui, Xiangxiang, et al.
Veröffentlicht: (2025)
Practising responsibility: Ethics in NLP as a hands-on course
von: Nissim, Malvina, et al.
Veröffentlicht: (2025)
von: Nissim, Malvina, et al.
Veröffentlicht: (2025)
Mitigating Overthinking in Large Reasoning Models via Manifold Steering
von: Huang, Yao, et al.
Veröffentlicht: (2025)
von: Huang, Yao, et al.
Veröffentlicht: (2025)
Large Language Models for Controllable Multi-property Multi-objective Molecule Optimization
von: Dey, Vishal, et al.
Veröffentlicht: (2025)
von: Dey, Vishal, et al.
Veröffentlicht: (2025)
Dialogue Action Tokens: Steering Language Models in Goal-Directed Dialogue with a Multi-Turn Planner
von: Li, Kenneth, et al.
Veröffentlicht: (2024)
von: Li, Kenneth, et al.
Veröffentlicht: (2024)
Learning Dynamics in Continual Pre-Training for Large Language Models
von: Wang, Xingjin, et al.
Veröffentlicht: (2025)
von: Wang, Xingjin, et al.
Veröffentlicht: (2025)
Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models
von: Bello, Femi, et al.
Veröffentlicht: (2025)
von: Bello, Femi, et al.
Veröffentlicht: (2025)
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
von: Han, Pengrui, et al.
Veröffentlicht: (2026)
von: Han, Pengrui, et al.
Veröffentlicht: (2026)
Do Large Language Models Know How Much They Know?
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Steering Large Language Models for Machine Translation Personalization
von: Scalena, Daniel, et al.
Veröffentlicht: (2025) -
A gentle push funziona benissimo: making instructed models in Italian via contrastive activation steering
von: Scalena, Daniel, et al.
Veröffentlicht: (2024) -
Non Verbis, Sed Rebus: Large Language Models are Weak Solvers of Italian Rebuses
von: Sarti, Gabriele, et al.
Veröffentlicht: (2024) -
Quantifying the Plausibility of Context Reliance in Neural Machine Translation
von: Sarti, Gabriele, et al.
Veröffentlicht: (2023) -
Compositional Steering of Large Language Models with Steering Tokens
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)