Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
Fuente:
arXiv
Saved in:
| Main Authors: | Sim, Woo Seob, Park, Yu Rang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Information Geometry of Softmax: Probing and Steering
by: Park, Kiho, et al.
Published: (2026)
by: Park, Kiho, et al.
Published: (2026)
Universal Conceptual Structure in Neural Translation: Probing NLLB-200's Multilingual Geometry
by: Mathewson, Kyle Elliott
Published: (2026)
by: Mathewson, Kyle Elliott
Published: (2026)
The Linear Representation Hypothesis and the Geometry of Large Language Models
by: Park, Kiho, et al.
Published: (2023)
by: Park, Kiho, et al.
Published: (2023)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
by: Yang, Yifei, et al.
Published: (2024)
by: Yang, Yifei, et al.
Published: (2024)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
by: Taniguchi, Rei, et al.
Published: (2026)
by: Taniguchi, Rei, et al.
Published: (2026)
The Geometry of Categorical and Hierarchical Concepts in Large Language Models
by: Park, Kiho, et al.
Published: (2024)
by: Park, Kiho, et al.
Published: (2024)
Integrating Locality-Aware Attention with Transformers for General Geometry PDEs
by: Koh, Minsu, et al.
Published: (2025)
by: Koh, Minsu, et al.
Published: (2025)
LiteSearch: Efficacious Tree Search for LLM
by: Wang, Ante, et al.
Published: (2024)
by: Wang, Ante, et al.
Published: (2024)
LPCD: Unified Framework from Layer-Wise to Submodule Quantization
by: Ichikawa, Yuma, et al.
Published: (2025)
by: Ichikawa, Yuma, et al.
Published: (2025)
Hidden Heroes and Gradient Bloats: Layer-Wise Redundancy Inverts Attribution in Transformers
by: Ye, Donald
Published: (2026)
by: Ye, Donald
Published: (2026)
The Geometries of Truth Are Orthogonal Across Tasks
by: Azizian, Waiss, et al.
Published: (2025)
by: Azizian, Waiss, et al.
Published: (2025)
Geometry of Decision Making in Language Models
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
Concept Layers: Enhancing Interpretability and Intervenability via LLM Conceptualization
by: Bidusa, Or Raphael, et al.
Published: (2025)
by: Bidusa, Or Raphael, et al.
Published: (2025)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
by: Lei, Ge, et al.
Published: (2025)
by: Lei, Ge, et al.
Published: (2025)
Cross-Layer Discrete Concept Discovery for Interpreting Language Models
by: Garg, Ankur, et al.
Published: (2025)
by: Garg, Ankur, et al.
Published: (2025)
Mousse: Rectifying the Geometry of Muon with Curvature-Aware Preconditioning
by: Zhang, Yechen, et al.
Published: (2026)
by: Zhang, Yechen, et al.
Published: (2026)
The Geometry of Meaning: Perfect Spacetime Representations of Hierarchical Structures
by: Anabalon, Andres, et al.
Published: (2025)
by: Anabalon, Andres, et al.
Published: (2025)
Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
by: Wang, Dianyun, et al.
Published: (2025)
by: Wang, Dianyun, et al.
Published: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
by: Li, Xing, et al.
Published: (2025)
by: Li, Xing, et al.
Published: (2025)
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
by: Sun, Lihao, et al.
Published: (2026)
by: Sun, Lihao, et al.
Published: (2026)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
by: Wollschläger, Tom, et al.
Published: (2025)
by: Wollschläger, Tom, et al.
Published: (2025)
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
by: Li, Melody Zixuan, et al.
Published: (2025)
by: Li, Melody Zixuan, et al.
Published: (2025)
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
by: Sinha, Shiven, et al.
Published: (2024)
by: Sinha, Shiven, et al.
Published: (2024)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
by: He, Junhui, et al.
Published: (2024)
by: He, Junhui, et al.
Published: (2024)
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
by: Schiekiera, Louis, et al.
Published: (2026)
by: Schiekiera, Louis, et al.
Published: (2026)
mini-vec2vec: Scaling Universal Geometry Alignment with Linear Transformations
by: Dar, Guy
Published: (2025)
by: Dar, Guy
Published: (2025)
Deep Language Geometry: Constructing a Metric Space from LLM Weights
by: Shamrai, Maksym, et al.
Published: (2025)
by: Shamrai, Maksym, et al.
Published: (2025)
Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation
by: Balestriero, Randall, et al.
Published: (2023)
by: Balestriero, Randall, et al.
Published: (2023)
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
by: Asif, Sadia, et al.
Published: (2026)
by: Asif, Sadia, et al.
Published: (2026)
Tracing Representation Progression: Analyzing and Enhancing Layer-Wise Similarity
by: Jiang, Jiachen, et al.
Published: (2024)
by: Jiang, Jiachen, et al.
Published: (2024)
Adaptive Margin RLHF via Preference over Preferences
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
MarginSel : Max-Margin Demonstration Selection for LLMs
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
Mixture-of-Experts as Soft Clustering: A Dual Jacobian-PCA Spectral Geometry Perspective
by: Liu, Feilong
Published: (2026)
by: Liu, Feilong
Published: (2026)
The Information of Large Language Model Geometry
by: Tan, Zhiquan, et al.
Published: (2024)
by: Tan, Zhiquan, et al.
Published: (2024)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
by: Achtibat, Reduan, et al.
Published: (2024)
by: Achtibat, Reduan, et al.
Published: (2024)
The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering
by: Zhou, Yefan, et al.
Published: (2026)
by: Zhou, Yefan, et al.
Published: (2026)
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
by: Weng, Jiaqi, et al.
Published: (2025)
by: Weng, Jiaqi, et al.
Published: (2025)
The Geometry of Reasoning: Flowing Logics in Representation Space
by: Zhou, Yufa, et al.
Published: (2025)
by: Zhou, Yufa, et al.
Published: (2025)
Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models
by: Sim, Shamus, et al.
Published: (2024)
by: Sim, Shamus, et al.
Published: (2024)
Similar Items
-
The Information Geometry of Softmax: Probing and Steering
by: Park, Kiho, et al.
Published: (2026) -
Universal Conceptual Structure in Neural Translation: Probing NLLB-200's Multilingual Geometry
by: Mathewson, Kyle Elliott
Published: (2026) -
The Linear Representation Hypothesis and the Geometry of Large Language Models
by: Park, Kiho, et al.
Published: (2023) -
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
by: Yang, Yifei, et al.
Published: (2024) -
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
by: Taniguchi, Rei, et al.
Published: (2026)