Extracting and Steering Emotion Representations in Small Language Models: A Methodological Comparison
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Jeong, Jihoon |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Shared Emotion Geometry Across Small Language Models: A Cross-Architecture Study of Representation, Behavior, and Methodological Confounds
par: Jeong, Jihoon
Publié: (2026)
par: Jeong, Jihoon
Publié: (2026)
Model Medicine: A Clinical Framework for Understanding, Diagnosing, and Treating AI Models
par: Jeong, Jihoon
Publié: (2026)
par: Jeong, Jihoon
Publié: (2026)
MTI: A Behavior-Based Temperament Profiling System for AI Agents
par: Jeong, Jihoon
Publié: (2026)
par: Jeong, Jihoon
Publié: (2026)
M-CARE: Standardized Clinical Case Reporting for AI Model Behavioral Disorders, with a 20-Case Atlas and Experimental Validation
par: Jeong, Jihoon
Publié: (2026)
par: Jeong, Jihoon
Publié: (2026)
On Effects of Steering Latent Representation for Large Language Model Unlearning
par: Huu-Tien, Dang, et autres
Publié: (2024)
par: Huu-Tien, Dang, et autres
Publié: (2024)
Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models
par: Bello, Femi, et autres
Publié: (2025)
par: Bello, Femi, et autres
Publié: (2025)
Steering Risk Preferences in Large Language Models by Aligning Behavioral and Neural Representations
par: Zhu, Jian-Qiao, et autres
Publié: (2025)
par: Zhu, Jian-Qiao, et autres
Publié: (2025)
Self-Steering Language Models
par: Grand, Gabriel, et autres
Publié: (2025)
par: Grand, Gabriel, et autres
Publié: (2025)
Steering When Necessary: Flexible Steering Large Language Models with Backtracking
par: Cheng, Zifeng, et autres
Publié: (2025)
par: Cheng, Zifeng, et autres
Publié: (2025)
On the Limitations of Steering in Language Model Alignment
par: Niranjan, Chebrolu, et autres
Publié: (2025)
par: Niranjan, Chebrolu, et autres
Publié: (2025)
Language-Specific Representation of Emotion-Concept Knowledge Causally Supports Emotion Inference
par: Li, Ming, et autres
Publié: (2023)
par: Li, Ming, et autres
Publié: (2023)
T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models
par: Kang, Minki, et autres
Publié: (2025)
par: Kang, Minki, et autres
Publié: (2025)
Probing Ethical Framework Representations in Large Language Models: Structure, Entanglement, and Methodological Challenges
par: Xu, Weilun, et autres
Publié: (2026)
par: Xu, Weilun, et autres
Publié: (2026)
Probing and Steering Evaluation Awareness of Language Models
par: Nguyen, Jord, et autres
Publié: (2025)
par: Nguyen, Jord, et autres
Publié: (2025)
Activation Scaling for Steering and Interpreting Language Models
par: Stoehr, Niklas, et autres
Publié: (2024)
par: Stoehr, Niklas, et autres
Publié: (2024)
CogSteer: Cognition-Inspired Selective Layer Intervention for Efficiently Steering Large Language Models
par: Wang, Xintong, et autres
Publié: (2024)
par: Wang, Xintong, et autres
Publié: (2024)
Compositional Steering of Large Language Models with Steering Tokens
par: Radevski, Gorjan, et autres
Publié: (2026)
par: Radevski, Gorjan, et autres
Publié: (2026)
Cross-Lingual Activation Steering for Multilingual Language Models
par: Pokharel, Rhitabrat, et autres
Publié: (2026)
par: Pokharel, Rhitabrat, et autres
Publié: (2026)
Steering Large Language Models to Evaluate and Amplify Creativity
par: Olson, Matthew Lyle, et autres
Publié: (2024)
par: Olson, Matthew Lyle, et autres
Publié: (2024)
Prompt-Based Value Steering of Large Language Models
par: Abbo, Giulio Antonio, et autres
Publié: (2025)
par: Abbo, Giulio Antonio, et autres
Publié: (2025)
Extracting Unlearned Information from LLMs with Activation Steering
par: Seyitoğlu, Atakan, et autres
Publié: (2024)
par: Seyitoğlu, Atakan, et autres
Publié: (2024)
Beyond Semantics: Measuring Fine-Grained Emotion Preservation in Small Language Model-Based Machine Translation
par: Wisniewski, Dawid, et autres
Publié: (2026)
par: Wisniewski, Dawid, et autres
Publié: (2026)
AI Steerability 360: A Toolkit for Steering Large Language Models
par: Miehling, Erik, et autres
Publié: (2026)
par: Miehling, Erik, et autres
Publié: (2026)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
par: Cheng, Stephen, et autres
Publié: (2026)
par: Cheng, Stephen, et autres
Publié: (2026)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
par: Siu, Vincent, et autres
Publié: (2025)
par: Siu, Vincent, et autres
Publié: (2025)
Steering Language Models Before They Speak: Logit-Level Interventions
par: An, Hyeseon, et autres
Publié: (2026)
par: An, Hyeseon, et autres
Publié: (2026)
DLM-SWAI: Steering Diffusion Language Models Before They Unmask
par: An, Hyeseon, et autres
Publié: (2026)
par: An, Hyeseon, et autres
Publié: (2026)
Steering Evaluation-Aware Language Models to Act Like They Are Deployed
par: Hua, Tim Tian, et autres
Publié: (2025)
par: Hua, Tim Tian, et autres
Publié: (2025)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
par: Siu, Vincent, et autres
Publié: (2025)
par: Siu, Vincent, et autres
Publié: (2025)
Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods
par: Wolf, Yotam, et autres
Publié: (2024)
par: Wolf, Yotam, et autres
Publié: (2024)
Under Pressure: Emotional Framing Induces Measurable Behavioral Shifts and Structured Internal Geometry in Small Language Models
par: Usman, Rana Muhammad
Publié: (2026)
par: Usman, Rana Muhammad
Publié: (2026)
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models
par: Dawes, Cutter, et autres
Publié: (2026)
par: Dawes, Cutter, et autres
Publié: (2026)
The Impact of Steering Large Language Models with Persona Vectors in Educational Applications
par: Wu, Yongchao, et autres
Publié: (2026)
par: Wu, Yongchao, et autres
Publié: (2026)
Word Embeddings Are Steers for Language Models
par: Han, Chi, et autres
Publié: (2023)
par: Han, Chi, et autres
Publié: (2023)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
par: Zhao, Haiyan, et autres
Publié: (2025)
par: Zhao, Haiyan, et autres
Publié: (2025)
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
par: Liu, Zheyuan, et autres
Publié: (2025)
par: Liu, Zheyuan, et autres
Publié: (2025)
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
par: Yu, Haeun, et autres
Publié: (2025)
par: Yu, Haeun, et autres
Publié: (2025)
Investigating Bias Representations in Llama 2 Chat via Activation Steering
par: Lu, Dawn, et autres
Publié: (2024)
par: Lu, Dawn, et autres
Publié: (2024)
Does Model Size Matter? A Comparison of Small and Large Language Models for Requirements Classification
par: Zadenoori, Mohammad Amin, et autres
Publié: (2025)
par: Zadenoori, Mohammad Amin, et autres
Publié: (2025)
Understanding Emotion in Discourse: Recognition Insights and Linguistic Patterns for Generation
par: Jeong, Cheonkam, et autres
Publié: (2026)
par: Jeong, Cheonkam, et autres
Publié: (2026)
Documents similaires
-
Shared Emotion Geometry Across Small Language Models: A Cross-Architecture Study of Representation, Behavior, and Methodological Confounds
par: Jeong, Jihoon
Publié: (2026) -
Model Medicine: A Clinical Framework for Understanding, Diagnosing, and Treating AI Models
par: Jeong, Jihoon
Publié: (2026) -
MTI: A Behavior-Based Temperament Profiling System for AI Agents
par: Jeong, Jihoon
Publié: (2026) -
M-CARE: Standardized Clinical Case Reporting for AI Model Behavioral Disorders, with a 20-Case Atlas and Experimental Validation
par: Jeong, Jihoon
Publié: (2026) -
On Effects of Steering Latent Representation for Large Language Model Unlearning
par: Huu-Tien, Dang, et autres
Publié: (2024)