Scaling laws for activation steering with Llama 2 models and refusal mechanisms
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ali, Sheikh Abdur Raheem, Xu, Justin, Yang, Ivory, Li, Jasmine Xinze, Arslan, Ayse, Benham, Clark |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning
von: Shabalin, Stepan, et al.
Veröffentlicht: (2025)
von: Shabalin, Stepan, et al.
Veröffentlicht: (2025)
Scaling laws in wearable human activity recognition
von: Hoddes, Tom, et al.
Veröffentlicht: (2025)
von: Hoddes, Tom, et al.
Veröffentlicht: (2025)
NushuRescue: Revitalization of the Endangered Nushu Language with AI
von: Yang, Ivory, et al.
Veröffentlicht: (2024)
von: Yang, Ivory, et al.
Veröffentlicht: (2024)
Rethinking harmless refusals when fine-tuning foundation models
von: Pop, Florin, et al.
Veröffentlicht: (2024)
von: Pop, Florin, et al.
Veröffentlicht: (2024)
AntiPhishStack: LSTM-based Stacked Generalization Model for Optimized Phishing URL Detection
von: Aslam, Saba, et al.
Veröffentlicht: (2024)
von: Aslam, Saba, et al.
Veröffentlicht: (2024)
Scaling-laws for Large Time-series Models
von: Edwards, Thomas D. P., et al.
Veröffentlicht: (2024)
von: Edwards, Thomas D. P., et al.
Veröffentlicht: (2024)
A gentle push funziona benissimo: making instructed models in Italian via contrastive activation steering
von: Scalena, Daniel, et al.
Veröffentlicht: (2024)
von: Scalena, Daniel, et al.
Veröffentlicht: (2024)
Mechanistic interpretability for steering vision-language-action models
von: Häon, Bear, et al.
Veröffentlicht: (2025)
von: Häon, Bear, et al.
Veröffentlicht: (2025)
A New Branch-and-Bound Pruning Framework for $\ell_0$-Regularized Problems
von: Guyard, Theo, et al.
Veröffentlicht: (2024)
von: Guyard, Theo, et al.
Veröffentlicht: (2024)
Llama 3 Meets MoE: Efficient Upcycling
von: Vavre, Aditya, et al.
Veröffentlicht: (2024)
von: Vavre, Aditya, et al.
Veröffentlicht: (2024)
NeuroSep-CP-LCB: A Deep Learning-based Contextual Multi-armed Bandit Algorithm with Uncertainty Quantification for Early Sepsis Prediction
von: Zhou, Anni, et al.
Veröffentlicht: (2025)
von: Zhou, Anni, et al.
Veröffentlicht: (2025)
Toward universal steering and monitoring of AI models
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025)
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025)
Scaling and context steer LLMs along the same computational path as the human brain
von: Raugel, Joséphine, et al.
Veröffentlicht: (2025)
von: Raugel, Joséphine, et al.
Veröffentlicht: (2025)
MAST: A Multi-fidelity Augmented Surrogate model via Spatial Trust-weighting
von: Nasr, Ahmed Mohamed Eisa, et al.
Veröffentlicht: (2026)
von: Nasr, Ahmed Mohamed Eisa, et al.
Veröffentlicht: (2026)
Zebra-Llama: Towards Extremely Efficient Hybrid Models
von: Yang, Mingyu, et al.
Veröffentlicht: (2025)
von: Yang, Mingyu, et al.
Veröffentlicht: (2025)
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
von: He, Zhengfu, et al.
Veröffentlicht: (2024)
von: He, Zhengfu, et al.
Veröffentlicht: (2024)
Open Llama2 Model for the Lithuanian Language
von: Nakvosas, Artūras, et al.
Veröffentlicht: (2024)
von: Nakvosas, Artūras, et al.
Veröffentlicht: (2024)
Decoding-Time Debiasing via Process Reward Models: From Controlled Fill-in to Open-Ended Generation
von: Khan, Muneeb Ur Raheem
Veröffentlicht: (2026)
von: Khan, Muneeb Ur Raheem
Veröffentlicht: (2026)
Robust LLM safeguarding via refusal feature adversarial training
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
Scaling laws for decoding images from brain activity
von: Banville, Hubert, et al.
Veröffentlicht: (2025)
von: Banville, Hubert, et al.
Veröffentlicht: (2025)
Sepsyn-OLCP: An Online Learning-based Framework for Early Sepsis Prediction with Uncertainty Quantification using Conformal Prediction
von: Zhou, Anni, et al.
Veröffentlicht: (2025)
von: Zhou, Anni, et al.
Veröffentlicht: (2025)
TeenyTinyLlama: open-source tiny language models trained in Brazilian Portuguese
von: Corrêa, Nicholas Kluge, et al.
Veröffentlicht: (2024)
von: Corrêa, Nicholas Kluge, et al.
Veröffentlicht: (2024)
ProgressGym: Alignment with a Millennium of Moral Progress
von: Qiu, Tianyi, et al.
Veröffentlicht: (2024)
von: Qiu, Tianyi, et al.
Veröffentlicht: (2024)
Steering Llama 2 via Contrastive Activation Addition
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
Physics-integrated generative modeling using attentive planar normalizing flow based variational autoencoder
von: Akhtar, Sheikh Waqas
Veröffentlicht: (2024)
von: Akhtar, Sheikh Waqas
Veröffentlicht: (2024)
The Mamba in the Llama: Distilling and Accelerating Hybrid Models
von: Wang, Junxiong, et al.
Veröffentlicht: (2024)
von: Wang, Junxiong, et al.
Veröffentlicht: (2024)
Patterns and Mechanisms of Contrastive Activation Engineering
von: Hao, Yixiong, et al.
Veröffentlicht: (2025)
von: Hao, Yixiong, et al.
Veröffentlicht: (2025)
Temporally Unified Adversarial Perturbations for Time Series Forecasting
von: Su, Ruixian, et al.
Veröffentlicht: (2026)
von: Su, Ruixian, et al.
Veröffentlicht: (2026)
Llama-Nemotron: Efficient Reasoning Models
von: Bercovich, Akhiad, et al.
Veröffentlicht: (2025)
von: Bercovich, Akhiad, et al.
Veröffentlicht: (2025)
Electron flow matching for generative reaction mechanism prediction obeying conservation laws
von: Joung, Joonyoung F., et al.
Veröffentlicht: (2025)
von: Joung, Joonyoung F., et al.
Veröffentlicht: (2025)
Can sparse autoencoders be used to decompose and interpret steering vectors?
von: Mayne, Harry, et al.
Veröffentlicht: (2024)
von: Mayne, Harry, et al.
Veröffentlicht: (2024)
A Policy Gradient-Based Sequence-to-Sequence Method for Time Series Prediction
von: Sima, Qi, et al.
Veröffentlicht: (2024)
von: Sima, Qi, et al.
Veröffentlicht: (2024)
Clinician input steers frontier AI models toward both accurate and harmful decisions
von: Lopez, Ivan, et al.
Veröffentlicht: (2026)
von: Lopez, Ivan, et al.
Veröffentlicht: (2026)
Forbidden Facts: An Investigation of Competing Objectives in Llama-2
von: Wang, Tony T., et al.
Veröffentlicht: (2023)
von: Wang, Tony T., et al.
Veröffentlicht: (2023)
Llama-Affinity: A Predictive Antibody Antigen Binding Model Integrating Antibody Sequences with Llama3 Backbone Architecture
von: Hossain, Delower, et al.
Veröffentlicht: (2025)
von: Hossain, Delower, et al.
Veröffentlicht: (2025)
Efficient semantic uncertainty quantification in language models via diversity-steered sampling
von: Park, Ji Won, et al.
Veröffentlicht: (2025)
von: Park, Ji Won, et al.
Veröffentlicht: (2025)
Scaling laws for amplitude surrogates
von: Bahl, Henning, et al.
Veröffentlicht: (2026)
von: Bahl, Henning, et al.
Veröffentlicht: (2026)
BanglaLlama: LLaMA for Bangla Language
von: Zehady, Abdullah Khan, et al.
Veröffentlicht: (2024)
von: Zehady, Abdullah Khan, et al.
Veröffentlicht: (2024)
Scaling laws for learning with real and surrogate data
von: Jain, Ayush, et al.
Veröffentlicht: (2024)
von: Jain, Ayush, et al.
Veröffentlicht: (2024)
SynLlama: Generating Synthesizable Molecules and Their Analogs with Large Language Models
von: Sun, Kunyang, et al.
Veröffentlicht: (2025)
von: Sun, Kunyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning
von: Shabalin, Stepan, et al.
Veröffentlicht: (2025) -
Scaling laws in wearable human activity recognition
von: Hoddes, Tom, et al.
Veröffentlicht: (2025) -
NushuRescue: Revitalization of the Endangered Nushu Language with AI
von: Yang, Ivory, et al.
Veröffentlicht: (2024) -
Rethinking harmless refusals when fine-tuning foundation models
von: Pop, Florin, et al.
Veröffentlicht: (2024) -
AntiPhishStack: LSTM-based Stacked Generalization Model for Optimized Phishing URL Detection
von: Aslam, Saba, et al.
Veröffentlicht: (2024)