Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Llorente-Saguer, Isaac |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Geometry of Harmful Intent: Training-Free Anomaly Detection via Angular Deviation in LLM Residual Streams
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
IntentGrasp: A Comprehensive Benchmark for Intent Understanding
von: Yin, Yuwei, et al.
Veröffentlicht: (2026)
von: Yin, Yuwei, et al.
Veröffentlicht: (2026)
SWI: Speaking with Intent in Large Language Models
von: Yin, Yuwei, et al.
Veröffentlicht: (2025)
von: Yin, Yuwei, et al.
Veröffentlicht: (2025)
Large Language Model (LLM) Bias Index -- LLMBI
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes
von: Liu, Ming
Veröffentlicht: (2026)
von: Liu, Ming
Veröffentlicht: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
von: Bhandari, Pranav, et al.
Veröffentlicht: (2026)
von: Bhandari, Pranav, et al.
Veröffentlicht: (2026)
Why Geometric Continuity Emerges in Deep Neural Networks: Residual Connections and Rotational Symmetry Breaking
von: Jeong, Kyungwon, et al.
Veröffentlicht: (2026)
von: Jeong, Kyungwon, et al.
Veröffentlicht: (2026)
CodingTeachLLM: Empowering LLM's Coding Ability via AST Prior Knowledge
von: Chen, Zhangquan, et al.
Veröffentlicht: (2024)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2024)
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
von: Ghosh, Shubhra, et al.
Veröffentlicht: (2025)
von: Ghosh, Shubhra, et al.
Veröffentlicht: (2025)
SocraSynth: Multi-LLM Reasoning with Conditional Statistics
von: Chang, Edward Y.
Veröffentlicht: (2024)
von: Chang, Edward Y.
Veröffentlicht: (2024)
RAC: Efficient LLM Factuality Correction with Retrieval Augmentation
von: Li, Changmao, et al.
Veröffentlicht: (2024)
von: Li, Changmao, et al.
Veröffentlicht: (2024)
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning
von: Chen, Zizhe, et al.
Veröffentlicht: (2026)
von: Chen, Zizhe, et al.
Veröffentlicht: (2026)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning
von: Saparkhan, Raman, et al.
Veröffentlicht: (2026)
von: Saparkhan, Raman, et al.
Veröffentlicht: (2026)
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2026)
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2026)
Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM
von: Codefuse, et al.
Veröffentlicht: (2025)
von: Codefuse, et al.
Veröffentlicht: (2025)
How Many Parameters Does Your Task Really Need? Task Specific Pruning with LLM-Sieve
von: Reda, Waleed, et al.
Veröffentlicht: (2025)
von: Reda, Waleed, et al.
Veröffentlicht: (2025)
Geometric Evolution Maps: Extracting Stable Concept Probes from Transformer Residual Streams
von: Henry, James
Veröffentlicht: (2026)
von: Henry, James
Veröffentlicht: (2026)
AMEL: Accumulated Message Effects on LLM Judgments
von: Temkit, Sid-Ali
Veröffentlicht: (2026)
von: Temkit, Sid-Ali
Veröffentlicht: (2026)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning
von: Pan, Muyu, et al.
Veröffentlicht: (2026)
von: Pan, Muyu, et al.
Veröffentlicht: (2026)
NRR-Phi: Text-to-State Mapping for Ambiguity Preservation in LLM Inference
von: Saito, Kei
Veröffentlicht: (2026)
von: Saito, Kei
Veröffentlicht: (2026)
Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
Sleepless Nights, Sugary Days: Creating Synthetic Users with Health Conditions for Realistic Coaching Agent Interactions
von: Yun, Taedong, et al.
Veröffentlicht: (2025)
von: Yun, Taedong, et al.
Veröffentlicht: (2025)
Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning
von: Xu, Shuyao, et al.
Veröffentlicht: (2025)
von: Xu, Shuyao, et al.
Veröffentlicht: (2025)
The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior
von: Larsen, Erik
Veröffentlicht: (2025)
von: Larsen, Erik
Veröffentlicht: (2025)
Lossless Prompt Compression via Dictionary-Encoding and In-Context Learning: Enabling Cost-Effective LLM Analysis of Repetitive Data
von: de Campos, Andresa Rodrigues, et al.
Veröffentlicht: (2026)
von: de Campos, Andresa Rodrigues, et al.
Veröffentlicht: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
Metaphors are a Source of Cross-Domain Misalignment of Large Reasoning Models
von: Hu, Zhibo, et al.
Veröffentlicht: (2026)
von: Hu, Zhibo, et al.
Veröffentlicht: (2026)
Dealing with Annotator Disagreement in Hate Speech Classification
von: Dehghan, Somaiyeh, et al.
Veröffentlicht: (2025)
von: Dehghan, Somaiyeh, et al.
Veröffentlicht: (2025)
Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies
von: Hong, Chunsan, et al.
Veröffentlicht: (2025)
von: Hong, Chunsan, et al.
Veröffentlicht: (2025)
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2023)
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2023)
Identity as Attractor: Geometric Evidence for Persistent Agent Architecture in LLM Activation Space
von: Vasilenko, Vladimir
Veröffentlicht: (2026)
von: Vasilenko, Vladimir
Veröffentlicht: (2026)
DariMis: Harm-Aware Modeling for Dari Misinformation Detection on YouTube
von: Baktash, Jawid Ahmad, et al.
Veröffentlicht: (2026)
von: Baktash, Jawid Ahmad, et al.
Veröffentlicht: (2026)
The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs
von: Oliveira, Rafael C. T.
Veröffentlicht: (2026)
von: Oliveira, Rafael C. T.
Veröffentlicht: (2026)
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
von: Gao, Yutong, et al.
Veröffentlicht: (2026)
von: Gao, Yutong, et al.
Veröffentlicht: (2026)
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
von: Vieira, Inês, et al.
Veröffentlicht: (2026)
von: Vieira, Inês, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
The Geometry of Harmful Intent: Training-Free Anomaly Detection via Angular Deviation in LLM Residual Streams
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026) -
IntentGrasp: A Comprehensive Benchmark for Intent Understanding
von: Yin, Yuwei, et al.
Veröffentlicht: (2026) -
SWI: Speaking with Intent in Large Language Models
von: Yin, Yuwei, et al.
Veröffentlicht: (2025) -
Large Language Model (LLM) Bias Index -- LLMBI
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023) -
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)