How a Bit Becomes a Story: Semantic Steering via Differentiable Fault Injection
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Haider, Zafaryab, Rahman, Md Hafizur, Moeykens, Shane, Devabhaktuni, Vijay, Chakraborty, Prabuddha |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
HARP: Measuring Harm Amplification in Multi-Agent LLM Systems
par: Rahman, Md Hafizur, et autres
Publié: (2026)
par: Rahman, Md Hafizur, et autres
Publié: (2026)
LeMo-NADe: Multi-Parameter Neural Architecture Discovery with LLMs
par: Rahman, Md Hafizur, et autres
Publié: (2024)
par: Rahman, Md Hafizur, et autres
Publié: (2024)
ILASH: A Predictive Neural Architecture Search Framework for Multi-Task Applications
par: Rahman, Md Hafizur, et autres
Publié: (2024)
par: Rahman, Md Hafizur, et autres
Publié: (2024)
The Laminar Flow Hypothesis: Detecting Jailbreaks via Semantic Turbulence in Large Language Models
par: Rahman, Md. Hasib Ur
Publié: (2025)
par: Rahman, Md. Hasib Ur
Publié: (2025)
Optimizing the Privacy-Utility Balance using Synthetic Data and Configurable Perturbation Pipelines
par: Sharma, Anantha, et autres
Publié: (2025)
par: Sharma, Anantha, et autres
Publié: (2025)
False Data Injection Attack Detection in Edge-based Smart Metering Networks with Federated Learning
par: Uddin, Md Raihan, et autres
Publié: (2024)
par: Uddin, Md Raihan, et autres
Publié: (2024)
X-DFS: Explainable Artificial Intelligence Guided Design-for-Security Solution Space Exploration
par: Mahfuz, Tanzim, et autres
Publié: (2024)
par: Mahfuz, Tanzim, et autres
Publié: (2024)
GoldenTransformer: A Modular Fault Injection Framework for Transformer Robustness Research
par: Howard, Luke
Publié: (2025)
par: Howard, Luke
Publié: (2025)
Bypassing Prompt Injection Detectors through Evasive Injections
par: Rahman, Md Jahedur, et autres
Publié: (2026)
par: Rahman, Md Jahedur, et autres
Publié: (2026)
Agile Story-Point Estimation: Is RAG a Better Way to Go?
par: Maha, Lamyea, et autres
Publié: (2026)
par: Maha, Lamyea, et autres
Publié: (2026)
DASH: A Meta-Attack Framework for Synthesizing Effective and Stealthy Adversarial Examples
par: Nafi, Abdullah Al Nomaan, et autres
Publié: (2025)
par: Nafi, Abdullah Al Nomaan, et autres
Publié: (2025)
BarrierSteer: LLM Safety via Learning Barrier Steering
par: Tran, Thanh Q., et autres
Publié: (2026)
par: Tran, Thanh Q., et autres
Publié: (2026)
Characterizing the Fault Response of the Intel Neural Compute Stick 2 Under Single-Pulse Electromagnetic Fault Injection
par: Kučerák, Štefan, et autres
Publié: (2026)
par: Kučerák, Štefan, et autres
Publié: (2026)
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
par: Jiang, Xinyan, et autres
Publié: (2026)
par: Jiang, Xinyan, et autres
Publié: (2026)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
par: Lee, Banseok, et autres
Publié: (2025)
par: Lee, Banseok, et autres
Publié: (2025)
Real Faults in Deep Learning Fault Benchmarks: How Real Are They?
par: Jahangirova, Gunel, et autres
Publié: (2024)
par: Jahangirova, Gunel, et autres
Publié: (2024)
ExecTune: Effective Steering of Black-Box LLMs with Guide Models
par: Lingam, Vijay, et autres
Publié: (2026)
par: Lingam, Vijay, et autres
Publié: (2026)
LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment
par: Lee, Banseok, et autres
Publié: (2026)
par: Lee, Banseok, et autres
Publié: (2026)
How Not to Detect Prompt Injections with an LLM
par: Choudhary, Sarthak, et autres
Publié: (2025)
par: Choudhary, Sarthak, et autres
Publié: (2025)
Relative Positioning Based Code Chunking Method For Rich Context Retrieval In Repository Level Code Completion Task With Code Language Model
par: Rahman, Imranur, et autres
Publié: (2025)
par: Rahman, Imranur, et autres
Publié: (2025)
Human-aligned Chess with a Bit of Search
par: Zhang, Yiming, et autres
Publié: (2024)
par: Zhang, Yiming, et autres
Publié: (2024)
When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity
par: Feuer, Benjamin, et autres
Publié: (2025)
par: Feuer, Benjamin, et autres
Publié: (2025)
HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs
par: Wang, Guoan, et autres
Publié: (2026)
par: Wang, Guoan, et autres
Publié: (2026)
BitFlipScope: Scalable Fault Localization and Recovery for Bit-Flip Corruptions in LLMs
par: Karamat, Muhammad Zeeshan, et autres
Publié: (2025)
par: Karamat, Muhammad Zeeshan, et autres
Publié: (2025)
Steering LLMs via Scalable Interactive Oversight
par: Zhou, Enyu, et autres
Publié: (2026)
par: Zhou, Enyu, et autres
Publié: (2026)
Strategic Fusion Optimizes Transformer Compression
par: Rahman, Md Shoaibur
Publié: (2025)
par: Rahman, Md Shoaibur
Publié: (2025)
CBMAS: Cognitive Behavioral Modeling via Activation Steering
par: Ismail, Ahmed H., et autres
Publié: (2026)
par: Ismail, Ahmed H., et autres
Publié: (2026)
The M-factor: A Novel Metric for Evaluating Neural Architecture Search in Resource-Constrained Environments
par: Thudumu, Srikanth, et autres
Publié: (2025)
par: Thudumu, Srikanth, et autres
Publié: (2025)
GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
par: Handa, Divij, et autres
Publié: (2025)
par: Handa, Divij, et autres
Publié: (2025)
Bit-Identical Medical Deep Learning via Structured Orthogonal Initialization
par: Shkolnikov, Yakov Pyotr
Publié: (2026)
par: Shkolnikov, Yakov Pyotr
Publié: (2026)
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
par: Hedström, Anna, et autres
Publié: (2025)
par: Hedström, Anna, et autres
Publié: (2025)
MidSteer: Optimal Affine Framework for Steering Generative Models
par: Gaintseva, Tatiana, et autres
Publié: (2026)
par: Gaintseva, Tatiana, et autres
Publié: (2026)
Domain Transfer Becomes Identifiable via a Single Alignment
par: Shrestha, Sagar, et autres
Publié: (2026)
par: Shrestha, Sagar, et autres
Publié: (2026)
Feature-Aware Anisotropic Local Differential Privacy for Utility-Preserving Graph Representation Learning in Metal Additive Manufacturing
par: Islam, MD Shafikul, et autres
Publié: (2026)
par: Islam, MD Shafikul, et autres
Publié: (2026)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
par: Lee, Deokjae, et autres
Publié: (2025)
par: Lee, Deokjae, et autres
Publié: (2025)
Understanding Reasoning in Thinking Language Models via Steering Vectors
par: Venhoff, Constantin, et autres
Publié: (2025)
par: Venhoff, Constantin, et autres
Publié: (2025)
Angular Steering: Behavior Control via Rotation in Activation Space
par: Vu, Hieu M., et autres
Publié: (2025)
par: Vu, Hieu M., et autres
Publié: (2025)
Preemptive Detection and Steering of LLM Misalignment via Latent Reachability
par: Karnik, Sathwik, et autres
Publié: (2025)
par: Karnik, Sathwik, et autres
Publié: (2025)
Flow-Aware GNN for Transmission Network Reconfiguration via Substation Breaker Optimization
par: Meng, Dekang, et autres
Publié: (2025)
par: Meng, Dekang, et autres
Publié: (2025)
Automating Agent Hijacking via Structural Template Injection
par: Deng, Xinhao, et autres
Publié: (2026)
par: Deng, Xinhao, et autres
Publié: (2026)
Documents similaires
-
HARP: Measuring Harm Amplification in Multi-Agent LLM Systems
par: Rahman, Md Hafizur, et autres
Publié: (2026) -
LeMo-NADe: Multi-Parameter Neural Architecture Discovery with LLMs
par: Rahman, Md Hafizur, et autres
Publié: (2024) -
ILASH: A Predictive Neural Architecture Search Framework for Multi-Task Applications
par: Rahman, Md Hafizur, et autres
Publié: (2024) -
The Laminar Flow Hypothesis: Detecting Jailbreaks via Semantic Turbulence in Large Language Models
par: Rahman, Md. Hasib Ur
Publié: (2025) -
Optimizing the Privacy-Utility Balance using Synthetic Data and Configurable Perturbation Pipelines
par: Sharma, Anantha, et autres
Publié: (2025)