Gespeichert in:
| Hauptverfasser: | Mohan, Vamshi Sunku, Gupta, Kaustubh, Das, Aneesha, Singh, Chandan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.22719 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Fog Computing for Security‐Aware Resource Allocation in Narrowband Internet of Things
von: Vamshi Sunku Mohan, et al.
Veröffentlicht: (2024)
von: Vamshi Sunku Mohan, et al.
Veröffentlicht: (2024)
VALOR: Value-Aware Revenue Uplift Modeling with Treatment-Gated Representation for B2B Sales
von: Guduguntla, Vamshi, et al.
Veröffentlicht: (2026)
von: Guduguntla, Vamshi, et al.
Veröffentlicht: (2026)
Bayesian Concept Bottleneck Models with LLM Priors
von: Feng, Jean, et al.
Veröffentlicht: (2024)
von: Feng, Jean, et al.
Veröffentlicht: (2024)
Principal Prototype Analysis on Manifold for Interpretable Reinforcement Learning
von: Vamshi, Bodla Krishna, et al.
Veröffentlicht: (2026)
von: Vamshi, Bodla Krishna, et al.
Veröffentlicht: (2026)
Assessing the Operational Viability of Foundation Models for Time Series Forecasting
von: Soni, Kavin, et al.
Veröffentlicht: (2026)
von: Soni, Kavin, et al.
Veröffentlicht: (2026)
Latent Space Disentanglement via Activation Steering for Interpretable Attribute Control in Symbolic Music Generation
von: Prokopiou, Ioannis, et al.
Veröffentlicht: (2026)
von: Prokopiou, Ioannis, et al.
Veröffentlicht: (2026)
Steering Large Language Model Activations in Sparse Spaces
von: Bayat, Reza, et al.
Veröffentlicht: (2025)
von: Bayat, Reza, et al.
Veröffentlicht: (2025)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
von: Bodla, Krishna Vamshi, et al.
Veröffentlicht: (2025)
von: Bodla, Krishna Vamshi, et al.
Veröffentlicht: (2025)
Angular Steering: Behavior Control via Rotation in Activation Space
von: Vu, Hieu M., et al.
Veröffentlicht: (2025)
von: Vu, Hieu M., et al.
Veröffentlicht: (2025)
Steering Conceptual Bias via Transformer Latent-Subspace Activation
von: Sharma, Vansh, et al.
Veröffentlicht: (2025)
von: Sharma, Vansh, et al.
Veröffentlicht: (2025)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
Subspace Control: Turning Constrained Model Steering into Controllable Spectral Optimization
von: Huang, Yancheng, et al.
Veröffentlicht: (2026)
von: Huang, Yancheng, et al.
Veröffentlicht: (2026)
Interpretable Reward Modeling with Active Concept Bottlenecks
von: Laguna, Sonia, et al.
Veröffentlicht: (2025)
von: Laguna, Sonia, et al.
Veröffentlicht: (2025)
Cross-Layer Subspace Coupling for LLM Compression: A Unifying Framework and Its Empirical Limits
von: Khilar, Snigdha Chandan
Veröffentlicht: (2026)
von: Khilar, Snigdha Chandan
Veröffentlicht: (2026)
Characterizing the Behavior of Training Mamba-based State Space Models on GPUs
von: Baruah, Trinayan, et al.
Veröffentlicht: (2025)
von: Baruah, Trinayan, et al.
Veröffentlicht: (2025)
Interpretable Prognostics with Concept Bottleneck Models
von: Forest, Florent, et al.
Veröffentlicht: (2024)
von: Forest, Florent, et al.
Veröffentlicht: (2024)
Discovering and Steering Interpretable Concepts in Large Generative Music Models
von: Singh, Nikhil, et al.
Veröffentlicht: (2025)
von: Singh, Nikhil, et al.
Veröffentlicht: (2025)
D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
von: Raina, Samarth, et al.
Veröffentlicht: (2025)
von: Raina, Samarth, et al.
Veröffentlicht: (2025)
CBMAS: Cognitive Behavioral Modeling via Activation Steering
von: Ismail, Ahmed H., et al.
Veröffentlicht: (2026)
von: Ismail, Ahmed H., et al.
Veröffentlicht: (2026)
Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs
von: Zhang, Qingru, et al.
Veröffentlicht: (2023)
von: Zhang, Qingru, et al.
Veröffentlicht: (2023)
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
von: Huang, Xinting, et al.
Veröffentlicht: (2025)
von: Huang, Xinting, et al.
Veröffentlicht: (2025)
Amnesia: Adversarial Semantic Layer Specific Activation Steering in Large Language Models
von: Raza, Ali, et al.
Veröffentlicht: (2026)
von: Raza, Ali, et al.
Veröffentlicht: (2026)
When the Coffee Feature Activates on Coffins: An Analysis of Feature Extraction and Steering for Mechanistic Interpretability
von: Ronge, Raphael, et al.
Veröffentlicht: (2026)
von: Ronge, Raphael, et al.
Veröffentlicht: (2026)
GSS: Gated Subspace Steering for Selective Memorization Mitigation in LLMs
von: Zhang, Xuanqi, et al.
Veröffentlicht: (2026)
von: Zhang, Xuanqi, et al.
Veröffentlicht: (2026)
Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study
von: Ponkshe, Kaustubh, et al.
Veröffentlicht: (2025)
von: Ponkshe, Kaustubh, et al.
Veröffentlicht: (2025)
Rethinking Interpretability in the Era of Large Language Models
von: Singh, Chandan, et al.
Veröffentlicht: (2024)
von: Singh, Chandan, et al.
Veröffentlicht: (2024)
Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing
von: Wang, Peihao, et al.
Veröffentlicht: (2024)
von: Wang, Peihao, et al.
Veröffentlicht: (2024)
Steering Language Models With Activation Engineering
von: Turner, Alexander Matt, et al.
Veröffentlicht: (2023)
von: Turner, Alexander Matt, et al.
Veröffentlicht: (2023)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
Interpretable Next-token Prediction via the Generalized Induction Head
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
von: Ghosh, Shaona, et al.
Veröffentlicht: (2025)
von: Ghosh, Shaona, et al.
Veröffentlicht: (2025)
Towards Reasonable Concept Bottleneck Models
von: Kalampalikis, Nektarios, et al.
Veröffentlicht: (2025)
von: Kalampalikis, Nektarios, et al.
Veröffentlicht: (2025)
Depth-Wise Activation Steering for Honest Language Models
von: Góral, Gracjan, et al.
Veröffentlicht: (2025)
von: Góral, Gracjan, et al.
Veröffentlicht: (2025)
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
von: Hemadri, Raghu Vamshi, et al.
Veröffentlicht: (2025)
von: Hemadri, Raghu Vamshi, et al.
Veröffentlicht: (2025)
Activation Steering with a Feedback Controller
von: Nguyen, Dung V., et al.
Veröffentlicht: (2025)
von: Nguyen, Dung V., et al.
Veröffentlicht: (2025)
HyperSteer: Activation Steering at Scale with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
Understanding In-context Learning of Addition via Activation Subspaces
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
Decoupled-Value Attention for Prior-Data Fitted Networks: GP Inference for Physical Equations
von: Sharma, Kaustubh, et al.
Veröffentlicht: (2025)
von: Sharma, Kaustubh, et al.
Veröffentlicht: (2025)
Dynamically Scaled Activation Steering
von: Ferrando, Alex, et al.
Veröffentlicht: (2025)
von: Ferrando, Alex, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leveraging Fog Computing for Security‐Aware Resource Allocation in Narrowband Internet of Things
von: Vamshi Sunku Mohan, et al.
Veröffentlicht: (2024) -
VALOR: Value-Aware Revenue Uplift Modeling with Treatment-Gated Representation for B2B Sales
von: Guduguntla, Vamshi, et al.
Veröffentlicht: (2026) -
Bayesian Concept Bottleneck Models with LLM Priors
von: Feng, Jean, et al.
Veröffentlicht: (2024) -
Principal Prototype Analysis on Manifold for Interpretable Reinforcement Learning
von: Vamshi, Bodla Krishna, et al.
Veröffentlicht: (2026) -
Assessing the Operational Viability of Foundation Models for Time Series Forecasting
von: Soni, Kavin, et al.
Veröffentlicht: (2026)