Extending Activation Steering to Broad Skills and Multiple Behaviours
Fuente:
arXiv
Salvato in:
| Autori principali: | van der Weij, Teun, Poesio, Massimo, Schoots, Nandi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
di: van der Weij, Teun, et al.
Pubblicazione: (2024)
di: van der Weij, Teun, et al.
Pubblicazione: (2024)
The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation
di: Pavlovic, Maja, et al.
Pubblicazione: (2024)
di: Pavlovic, Maja, et al.
Pubblicazione: (2024)
An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration
di: Pavlovic, Maja, et al.
Pubblicazione: (2026)
di: Pavlovic, Maja, et al.
Pubblicazione: (2026)
A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents
di: Arghal, Raghu, et al.
Pubblicazione: (2026)
di: Arghal, Raghu, et al.
Pubblicazione: (2026)
Improving Academic Skills Assessment with NLP and Ensemble Learning
di: Huang, Xinyi, et al.
Pubblicazione: (2024)
di: Huang, Xinyi, et al.
Pubblicazione: (2024)
HyperSteer: Activation Steering at Scale with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
di: Heyman, Geert, et al.
Pubblicazione: (2026)
di: Heyman, Geert, et al.
Pubblicazione: (2026)
Dissecting Language Models: Machine Unlearning via Selective Pruning
di: Pochinkov, Nicholas, et al.
Pubblicazione: (2024)
di: Pochinkov, Nicholas, et al.
Pubblicazione: (2024)
Programming Refusal with Conditional Activation Steering
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
SAKE: Steering Activations for Knowledge Editing
di: Scialanga, Marco, et al.
Pubblicazione: (2025)
di: Scialanga, Marco, et al.
Pubblicazione: (2025)
Sentiment Analysis of Airbnb Reviews: Exploring Their Impact on Acceptance Rates and Pricing Across Multiple U.S. Regions
di: Safari, Ali
Pubblicazione: (2025)
di: Safari, Ali
Pubblicazione: (2025)
Endogenous Resistance to Activation Steering in Language Models
di: McKenzie, Alex, et al.
Pubblicazione: (2026)
di: McKenzie, Alex, et al.
Pubblicazione: (2026)
Toward Preference-aligned Large Language Models via Residual-based Model Steering
di: La Cava, Lucio, et al.
Pubblicazione: (2025)
di: La Cava, Lucio, et al.
Pubblicazione: (2025)
AI-Mediated Communication Can Steer Collective Opinion
di: Tsirtsis, Stratis, et al.
Pubblicazione: (2026)
di: Tsirtsis, Stratis, et al.
Pubblicazione: (2026)
Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior
di: Ardebili, Ali Aghazadeh, et al.
Pubblicazione: (2026)
di: Ardebili, Ali Aghazadeh, et al.
Pubblicazione: (2026)
Extracting Unlearned Information from LLMs with Activation Steering
di: Seyitoğlu, Atakan, et al.
Pubblicazione: (2024)
di: Seyitoğlu, Atakan, et al.
Pubblicazione: (2024)
Steering Llama 2 via Contrastive Activation Addition
di: Panickssery, Nina, et al.
Pubblicazione: (2023)
di: Panickssery, Nina, et al.
Pubblicazione: (2023)
Improving Instruction-Following in Language Models through Activation Steering
di: Stolfo, Alessandro, et al.
Pubblicazione: (2024)
di: Stolfo, Alessandro, et al.
Pubblicazione: (2024)
How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare
di: Schroeder, Daniel Thilo, et al.
Pubblicazione: (2025)
di: Schroeder, Daniel Thilo, et al.
Pubblicazione: (2025)
Multi-property Steering of Large Language Models with Dynamic Activation Composition
di: Scalena, Daniel, et al.
Pubblicazione: (2024)
di: Scalena, Daniel, et al.
Pubblicazione: (2024)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
di: Bigelow, Eric, et al.
Pubblicazione: (2025)
di: Bigelow, Eric, et al.
Pubblicazione: (2025)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
di: Soo, Samuel, et al.
Pubblicazione: (2025)
di: Soo, Samuel, et al.
Pubblicazione: (2025)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
di: Anand, Nikhil, et al.
Pubblicazione: (2026)
di: Anand, Nikhil, et al.
Pubblicazione: (2026)
Understanding The Effect Of Temperature On Alignment With Human Opinions
di: Pavlovic, Maja, et al.
Pubblicazione: (2024)
di: Pavlovic, Maja, et al.
Pubblicazione: (2024)
CarbonScaling: Extending Neural Scaling Laws for Carbon Footprint in Large Language Models
di: Jiang, Lei, et al.
Pubblicazione: (2025)
di: Jiang, Lei, et al.
Pubblicazione: (2025)
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
Reliability Estimation of News Media Sources: Birds of a Feather Flock Together
di: Burdisso, Sergio, et al.
Pubblicazione: (2024)
di: Burdisso, Sergio, et al.
Pubblicazione: (2024)
How Far Are We From AGI: Are LLMs All We Need?
di: Feng, Tao, et al.
Pubblicazione: (2024)
di: Feng, Tao, et al.
Pubblicazione: (2024)
Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
di: Liu, Ryan, et al.
Pubblicazione: (2024)
di: Liu, Ryan, et al.
Pubblicazione: (2024)
Laissez-Faire Harms: Algorithmic Biases in Generative Language Models
di: Shieh, Evan, et al.
Pubblicazione: (2024)
di: Shieh, Evan, et al.
Pubblicazione: (2024)
Deep Learning Based Amharic Chatbot for FAQs in Universities
di: Hailu, Goitom Ybrah, et al.
Pubblicazione: (2024)
di: Hailu, Goitom Ybrah, et al.
Pubblicazione: (2024)
Incorporating Graph Attention Mechanism into Geometric Problem Solving Based on Deep Reinforcement Learning
di: Zhong, Xiuqin, et al.
Pubblicazione: (2024)
di: Zhong, Xiuqin, et al.
Pubblicazione: (2024)
GreedLlama: Performance of Financial Value-Aligned Large Language Models in Moral Reasoning
di: Yu, Jeffy, et al.
Pubblicazione: (2024)
di: Yu, Jeffy, et al.
Pubblicazione: (2024)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
di: Kumar, Abhishek, et al.
Pubblicazione: (2024)
di: Kumar, Abhishek, et al.
Pubblicazione: (2024)
LLM-Assisted Content Conditional Debiasing for Fair Text Embedding
di: Deng, Wenlong, et al.
Pubblicazione: (2024)
di: Deng, Wenlong, et al.
Pubblicazione: (2024)
Explainable Artificial Intelligence: A Survey of Needs, Techniques, Applications, and Future Direction
di: Mersha, Melkamu, et al.
Pubblicazione: (2024)
di: Mersha, Melkamu, et al.
Pubblicazione: (2024)
Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
di: Zhou, Han, et al.
Pubblicazione: (2024)
di: Zhou, Han, et al.
Pubblicazione: (2024)
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
di: Mohamed, Youssef, et al.
Pubblicazione: (2024)
di: Mohamed, Youssef, et al.
Pubblicazione: (2024)
Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
di: Chien, Jennifer, et al.
Pubblicazione: (2024)
di: Chien, Jennifer, et al.
Pubblicazione: (2024)
Leveraging Social Determinants of Health in Alzheimer's Research Using LLM-Augmented Literature Mining and Knowledge Graphs
di: Shang, Tianqi, et al.
Pubblicazione: (2024)
di: Shang, Tianqi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
di: van der Weij, Teun, et al.
Pubblicazione: (2024) -
The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation
di: Pavlovic, Maja, et al.
Pubblicazione: (2024) -
An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration
di: Pavlovic, Maja, et al.
Pubblicazione: (2026) -
A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents
di: Arghal, Raghu, et al.
Pubblicazione: (2026) -
Improving Academic Skills Assessment with NLP and Ensemble Learning
di: Huang, Xinyi, et al.
Pubblicazione: (2024)