Concept Heterogeneity-aware Representation Steering
Fuente:
arXiv
Saved in:
| Main Authors: | Abdullaev, Laziz U., Wong, Noelle Y. L., Lee, Ryan T. Z., Jiang, Shiqi, Nguyen, Khoi N. M., Nguyen, Tan M. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Blessing and Curse of Dimensionality in Safety Alignment
by: Teo, Rachel S. Y., et al.
Published: (2025)
by: Teo, Rachel S. Y., et al.
Published: (2025)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
by: Nguyen-Nhat, Minh-Khoi, et al.
Published: (2025)
by: Nguyen-Nhat, Minh-Khoi, et al.
Published: (2025)
Elliptical Attention
by: Nielsen, Stefan K., et al.
Published: (2024)
by: Nielsen, Stefan K., et al.
Published: (2024)
Transformer Meets Twicing: Harnessing Unattended Residual Information
by: Abdullaev, Laziz, et al.
Published: (2025)
by: Abdullaev, Laziz, et al.
Published: (2025)
Revisiting Transformers with Insights from Image Filtering and Boosting
by: Abdullaev, Laziz U., et al.
Published: (2025)
by: Abdullaev, Laziz U., et al.
Published: (2025)
Angular Steering: Behavior Control via Rotation in Activation Space
by: Vu, Hieu M., et al.
Published: (2025)
by: Vu, Hieu M., et al.
Published: (2025)
Tight Clusters Make Specialized Experts
by: Nielsen, Stefan K., et al.
Published: (2025)
by: Nielsen, Stefan K., et al.
Published: (2025)
Spherical Tree-Sliced Wasserstein Distance
by: Tran, Viet-Hoang, et al.
Published: (2025)
by: Tran, Viet-Hoang, et al.
Published: (2025)
Distance-Based Tree-Sliced Wasserstein Distance
by: Tran, Hoang V., et al.
Published: (2025)
by: Tran, Hoang V., et al.
Published: (2025)
Minimizing Collateral Damage in Activation Steering
by: Nguyen, Tam, et al.
Published: (2026)
by: Nguyen, Tam, et al.
Published: (2026)
Tree-Sliced Wasserstein Distance: A Geometric Perspective
by: Tran, Viet-Hoang, et al.
Published: (2024)
by: Tran, Viet-Hoang, et al.
Published: (2024)
Revisiting LARS for Large Batch Training Generalization of Neural Networks
by: Do, Khoi, et al.
Published: (2023)
by: Do, Khoi, et al.
Published: (2023)
Towards Layer-Wise Personalized Federated Learning: Adaptive Layer Disentanglement via Conflicting Gradients
by: Nguyen, Minh Duong, et al.
Published: (2024)
by: Nguyen, Minh Duong, et al.
Published: (2024)
How Homogenizing the Channel-wise Magnitude Can Enhance EEG Classification Model?
by: Ngo, Huyen, et al.
Published: (2024)
by: Ngo, Huyen, et al.
Published: (2024)
Learning Reconfigurable Representations for Multimodal Federated Learning with Missing Data
by: Nguyen, Duong M., et al.
Published: (2025)
by: Nguyen, Duong M., et al.
Published: (2025)
Multiple-Input Variational Auto-Encoder for Anomaly Detection in Heterogeneous Data
by: Dinh, Phai Vu, et al.
Published: (2025)
by: Dinh, Phai Vu, et al.
Published: (2025)
On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
Latency-aware Multimodal Federated Learning over UAV Networks
by: Shaon, Shaba, et al.
Published: (2025)
by: Shaon, Shaba, et al.
Published: (2025)
From Coupled Oscillators to Graph Neural Networks: Reducing Over-smoothing via a Kuramoto Model-based Approach
by: Nguyen, Tuan, et al.
Published: (2023)
by: Nguyen, Tuan, et al.
Published: (2023)
Test-time Diverse Reasoning by Riemannian Activation Steering
by: Khanh, Ly Tran Ho, et al.
Published: (2025)
by: Khanh, Ly Tran Ho, et al.
Published: (2025)
Probabilistic Federated Learning on Uncertain and Heterogeneous Data with Model Personalization
by: Rahman, Ratun, et al.
Published: (2026)
by: Rahman, Ratun, et al.
Published: (2026)
Momentum Contrastive Learning with Enhanced Negative Sampling and Hard Negative Filtering
by: Hoang, Duy, et al.
Published: (2025)
by: Hoang, Duy, et al.
Published: (2025)
Emotions as Ambiguity-aware Ordinal Representations
by: Wu, Jingyao, et al.
Published: (2025)
by: Wu, Jingyao, et al.
Published: (2025)
BarrierSteer: LLM Safety via Learning Barrier Steering
by: Tran, Thanh Q., et al.
Published: (2026)
by: Tran, Thanh Q., et al.
Published: (2026)
Lookahead Pathology in Monte-Carlo Tree Search
by: Nguyen, Khoi P. N., et al.
Published: (2022)
by: Nguyen, Khoi P. N., et al.
Published: (2022)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
by: Teo, Rachel S. Y., et al.
Published: (2025)
by: Teo, Rachel S. Y., et al.
Published: (2025)
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
by: Teo, Rachel S. Y., et al.
Published: (2024)
by: Teo, Rachel S. Y., et al.
Published: (2024)
Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
by: Teo, Rachel S. Y., et al.
Published: (2024)
by: Teo, Rachel S. Y., et al.
Published: (2024)
Leveraging Multi-facet Paths for Heterogeneous Graph Representation Learning
by: Kim, Jongwoo, et al.
Published: (2024)
by: Kim, Jongwoo, et al.
Published: (2024)
Multi-Attribute Steering of Language Models via Targeted Intervention
by: Nguyen, Duy, et al.
Published: (2025)
by: Nguyen, Duy, et al.
Published: (2025)
Revisiting Kernel Attention with Correlated Gaussian Process Representation
by: Bui, Long Minh, et al.
Published: (2025)
by: Bui, Long Minh, et al.
Published: (2025)
HyperMono: A Monotonicity-aware Approach to Hyper-Relational Knowledge Representation
by: Hu, Zhiwei, et al.
Published: (2024)
by: Hu, Zhiwei, et al.
Published: (2024)
Conformalized Neural Networks for Federated Uncertainty Quantification under Dual Heterogeneity
by: Nguyen, Quang-Huy, et al.
Published: (2026)
by: Nguyen, Quang-Huy, et al.
Published: (2026)
Learning Strategy Representation for Imitation Learning in Multi-Agent Games
by: Lei, Shiqi, et al.
Published: (2024)
by: Lei, Shiqi, et al.
Published: (2024)
Scalable Heterogeneous Graph Learning via Heterogeneous-aware Orthogonal Prototype Experts
by: Zhou, Wei, et al.
Published: (2026)
by: Zhou, Wei, et al.
Published: (2026)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
by: Cheng, Stephen, et al.
Published: (2026)
by: Cheng, Stephen, et al.
Published: (2026)
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
by: Jiang, Xinyan, et al.
Published: (2026)
by: Jiang, Xinyan, et al.
Published: (2026)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
by: Casademunt, Helena, et al.
Published: (2025)
by: Casademunt, Helena, et al.
Published: (2025)
Generative Conditional Distributions by Neural (Entropic) Optimal Transport
by: Nguyen, Bao, et al.
Published: (2024)
by: Nguyen, Bao, et al.
Published: (2024)
Similar Items
-
The Blessing and Curse of Dimensionality in Safety Alignment
by: Teo, Rachel S. Y., et al.
Published: (2025) -
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
by: Nguyen-Nhat, Minh-Khoi, et al.
Published: (2025) -
Elliptical Attention
by: Nielsen, Stefan K., et al.
Published: (2024) -
Transformer Meets Twicing: Harnessing Unattended Residual Information
by: Abdullaev, Laziz, et al.
Published: (2025) -
Revisiting Transformers with Insights from Image Filtering and Boosting
by: Abdullaev, Laziz U., et al.
Published: (2025)