Salvato in:
| Autori principali: | Lu, Dawn, Rimsky, Nina |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2402.00402 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Steering Llama 2 via Contrastive Activation Addition
di: Panickssery, Nina, et al.
Pubblicazione: (2023)
di: Panickssery, Nina, et al.
Pubblicazione: (2023)
AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
di: Fathullah, Yassir, et al.
Pubblicazione: (2023)
di: Fathullah, Yassir, et al.
Pubblicazione: (2023)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
di: Siu, Vincent, et al.
Pubblicazione: (2025)
di: Siu, Vincent, et al.
Pubblicazione: (2025)
Steering Awareness: Detecting Activation Steering from Within
di: Rivera, Joshua Fonseca, et al.
Pubblicazione: (2025)
di: Rivera, Joshua Fonseca, et al.
Pubblicazione: (2025)
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
di: Kang, Diancheng, et al.
Pubblicazione: (2026)
di: Kang, Diancheng, et al.
Pubblicazione: (2026)
Forbidden Facts: An Investigation of Competing Objectives in Llama-2
di: Wang, Tony T., et al.
Pubblicazione: (2023)
di: Wang, Tony T., et al.
Pubblicazione: (2023)
HyperSteer: Activation Steering at Scale with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
TinyLlama: An Open-Source Small Language Model
di: Zhang, Peiyuan, et al.
Pubblicazione: (2024)
di: Zhang, Peiyuan, et al.
Pubblicazione: (2024)
MGH Radiology Llama: A Llama 3 70B Model for Radiology
di: Shi, Yucheng, et al.
Pubblicazione: (2024)
di: Shi, Yucheng, et al.
Pubblicazione: (2024)
Steering Towards Fairness: Mitigating Political Bias in LLMs
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
di: Ackerman, Christopher, et al.
Pubblicazione: (2024)
di: Ackerman, Christopher, et al.
Pubblicazione: (2024)
Steer Like the LLM: Activation Steering that Mimics Prompting
di: Heyman, Geert, et al.
Pubblicazione: (2026)
di: Heyman, Geert, et al.
Pubblicazione: (2026)
Activation Scaling for Steering and Interpreting Language Models
di: Stoehr, Niklas, et al.
Pubblicazione: (2024)
di: Stoehr, Niklas, et al.
Pubblicazione: (2024)
Fusion Steering: Prompt-Specific Activation Control
di: Chang, Waldemar, et al.
Pubblicazione: (2025)
di: Chang, Waldemar, et al.
Pubblicazione: (2025)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
di: Siu, Vincent, et al.
Pubblicazione: (2025)
di: Siu, Vincent, et al.
Pubblicazione: (2025)
When Wording Steers the Evaluation: Framing Bias in LLM judges
di: Hwang, Yerin, et al.
Pubblicazione: (2026)
di: Hwang, Yerin, et al.
Pubblicazione: (2026)
Cross-Lingual Activation Steering for Multilingual Language Models
di: Pokharel, Rhitabrat, et al.
Pubblicazione: (2026)
di: Pokharel, Rhitabrat, et al.
Pubblicazione: (2026)
Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
di: Nadeem, Afrozah, et al.
Pubblicazione: (2026)
di: Nadeem, Afrozah, et al.
Pubblicazione: (2026)
Programming Refusal with Conditional Activation Steering
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
SAKE: Steering Activations for Knowledge Editing
di: Scialanga, Marco, et al.
Pubblicazione: (2025)
di: Scialanga, Marco, et al.
Pubblicazione: (2025)
Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
di: Wannan, et al.
Pubblicazione: (2025)
di: Wannan, et al.
Pubblicazione: (2025)
DESTEIN: Navigating Detoxification of Language Models via Universal Steering Pairs and Head-wise Activation Fusion
di: Li, Yu, et al.
Pubblicazione: (2024)
di: Li, Yu, et al.
Pubblicazione: (2024)
Open Llama2 Model for the Lithuanian Language
di: Nakvosas, Artūras, et al.
Pubblicazione: (2024)
di: Nakvosas, Artūras, et al.
Pubblicazione: (2024)
Endogenous Resistance to Activation Steering in Language Models
di: McKenzie, Alex, et al.
Pubblicazione: (2026)
di: McKenzie, Alex, et al.
Pubblicazione: (2026)
Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes
di: Masoudian, Shahed, et al.
Pubblicazione: (2025)
di: Masoudian, Shahed, et al.
Pubblicazione: (2025)
On Effects of Steering Latent Representation for Large Language Model Unlearning
di: Huu-Tien, Dang, et al.
Pubblicazione: (2024)
di: Huu-Tien, Dang, et al.
Pubblicazione: (2024)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models
di: Guang, Jiahui, et al.
Pubblicazione: (2026)
di: Guang, Jiahui, et al.
Pubblicazione: (2026)
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
di: Rodriguez, Pau, et al.
Pubblicazione: (2025)
di: Rodriguez, Pau, et al.
Pubblicazione: (2025)
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
di: Siddique, Zara, et al.
Pubblicazione: (2025)
di: Siddique, Zara, et al.
Pubblicazione: (2025)
CoSteer: Collaborative Decoding-Time Personalization via Local Delta Steering
di: Lv, Hang, et al.
Pubblicazione: (2025)
di: Lv, Hang, et al.
Pubblicazione: (2025)
Extracting Unlearned Information from LLMs with Activation Steering
di: Seyitoğlu, Atakan, et al.
Pubblicazione: (2024)
di: Seyitoğlu, Atakan, et al.
Pubblicazione: (2024)
The Llama 3 Herd of Models
di: Grattafiori, Aaron, et al.
Pubblicazione: (2024)
di: Grattafiori, Aaron, et al.
Pubblicazione: (2024)
Take Care of Your Prompt Bias! Investigating and Mitigating Prompt Bias in Factual Knowledge Extraction
di: Xu, Ziyang, et al.
Pubblicazione: (2024)
di: Xu, Ziyang, et al.
Pubblicazione: (2024)
EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering
di: Xu, Haolei, et al.
Pubblicazione: (2025)
di: Xu, Haolei, et al.
Pubblicazione: (2025)
JetMoE: Reaching Llama2 Performance with 0.1M Dollars
di: Shen, Yikang, et al.
Pubblicazione: (2024)
di: Shen, Yikang, et al.
Pubblicazione: (2024)
Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
di: Valentino, Marco, et al.
Pubblicazione: (2025)
di: Valentino, Marco, et al.
Pubblicazione: (2025)
Llama-Nemotron: Efficient Reasoning Models
di: Bercovich, Akhiad, et al.
Pubblicazione: (2025)
di: Bercovich, Akhiad, et al.
Pubblicazione: (2025)
Extending Activation Steering to Broad Skills and Multiple Behaviours
di: van der Weij, Teun, et al.
Pubblicazione: (2024)
di: van der Weij, Teun, et al.
Pubblicazione: (2024)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
di: Urchs, Stefanie, et al.
Pubblicazione: (2023)
di: Urchs, Stefanie, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Steering Llama 2 via Contrastive Activation Addition
di: Panickssery, Nina, et al.
Pubblicazione: (2023) -
AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
di: Fathullah, Yassir, et al.
Pubblicazione: (2023) -
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
di: Siu, Vincent, et al.
Pubblicazione: (2025) -
Steering Awareness: Detecting Activation Steering from Within
di: Rivera, Joshua Fonseca, et al.
Pubblicazione: (2025) -
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
di: Kang, Diancheng, et al.
Pubblicazione: (2026)