Geometric Evolution Maps: Extracting Stable Concept Probes from Transformer Residual Streams
Fuente:
arXiv
Guardado en:
| Autor principal: | Henry, James |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
por: Henry, James
Publicado: (2026)
por: Henry, James
Publicado: (2026)
A Practical Guide to Streaming Continual Learning
por: Cossu, Andrea, et al.
Publicado: (2026)
por: Cossu, Andrea, et al.
Publicado: (2026)
Why Geometric Continuity Emerges in Deep Neural Networks: Residual Connections and Rotational Symmetry Breaking
por: Jeong, Kyungwon, et al.
Publicado: (2026)
por: Jeong, Kyungwon, et al.
Publicado: (2026)
cPNN: Continuous Progressive Neural Networks for Evolving Streaming Time Series
por: Giannini, Federico, et al.
Publicado: (2026)
por: Giannini, Federico, et al.
Publicado: (2026)
Don't Look Back in Anger: MAGIC Net for Streaming Continual Learning with Temporal Dependence
por: Giannini, Federico, et al.
Publicado: (2026)
por: Giannini, Federico, et al.
Publicado: (2026)
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
por: Giannini, Federico, et al.
Publicado: (2026)
por: Giannini, Federico, et al.
Publicado: (2026)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
por: Sarkar, Nilesh, et al.
Publicado: (2026)
por: Sarkar, Nilesh, et al.
Publicado: (2026)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
por: Borobia, Hector, et al.
Publicado: (2026)
por: Borobia, Hector, et al.
Publicado: (2026)
ProbeScale: Probing Analysis to Optimize Neural Scaling Laws for Efficient Small Language Model Inference
por: Das, Sourav
Publicado: (2026)
por: Das, Sourav
Publicado: (2026)
NeuronSpark: A Spiking Neural Network Language Model with Selective State Space Dynamics
por: Tang, Zhengzheng
Publicado: (2026)
por: Tang, Zhengzheng
Publicado: (2026)
ProactBench: Beyond What The User Asked For
por: Harfi, Sepehr, et al.
Publicado: (2026)
por: Harfi, Sepehr, et al.
Publicado: (2026)
Streaming Continual Learning for Unified Adaptive Intelligence in Dynamic Environments
por: Giannini, Federico, et al.
Publicado: (2026)
por: Giannini, Federico, et al.
Publicado: (2026)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
por: Wang, Zhen, et al.
Publicado: (2025)
por: Wang, Zhen, et al.
Publicado: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
por: Yang, Yibo
Publicado: (2025)
por: Yang, Yibo
Publicado: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
por: Basu, Abhinaba
Publicado: (2026)
por: Basu, Abhinaba
Publicado: (2026)
mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters
por: Mutlu, Abdulvahap, et al.
Publicado: (2026)
por: Mutlu, Abdulvahap, et al.
Publicado: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
por: Guo, Dongxin, et al.
Publicado: (2026)
por: Guo, Dongxin, et al.
Publicado: (2026)
Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
por: Garg, Saloni, et al.
Publicado: (2026)
por: Garg, Saloni, et al.
Publicado: (2026)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
por: Asperti, Andrea, et al.
Publicado: (2025)
por: Asperti, Andrea, et al.
Publicado: (2025)
Harnessing non-adversarial robustness in large language models
por: Zhou, Qinghua, et al.
Publicado: (2026)
por: Zhou, Qinghua, et al.
Publicado: (2026)
Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
por: Li, Yin
Publicado: (2025)
por: Li, Yin
Publicado: (2025)
ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
por: Chen, Yihong, et al.
Publicado: (2022)
por: Chen, Yihong, et al.
Publicado: (2022)
Extracting Sentence Embeddings from Pretrained Transformer Models
por: Stankevičius, Lukas, et al.
Publicado: (2024)
por: Stankevičius, Lukas, et al.
Publicado: (2024)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
por: Keeman, Michael
Publicado: (2026)
por: Keeman, Michael
Publicado: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
por: Mathew, Aby Mammen
Publicado: (2026)
por: Mathew, Aby Mammen
Publicado: (2026)
Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers
por: Nayak, Nikhil, et al.
Publicado: (2026)
por: Nayak, Nikhil, et al.
Publicado: (2026)
DreamNet: A Multimodal Framework for Semantic and Emotional Analysis of Sleep Narratives
por: Panchagnula, Tapasvi
Publicado: (2025)
por: Panchagnula, Tapasvi
Publicado: (2025)
DRO-InstructZero: Distributionally Robust Prompt Optimization for Large Language Models
por: Li, Yangyang
Publicado: (2025)
por: Li, Yangyang
Publicado: (2025)
CircuitProbe: Predicting Reasoning Circuits in Transformers via Stability Zone Detection
por: Panuganti, Rajkiran
Publicado: (2026)
por: Panuganti, Rajkiran
Publicado: (2026)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
por: Alpay, Faruk, et al.
Publicado: (2026)
por: Alpay, Faruk, et al.
Publicado: (2026)
A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
por: Zhao, Zhilong, et al.
Publicado: (2025)
por: Zhao, Zhilong, et al.
Publicado: (2025)
Reconstructing 12-Lead ECG from 3-Lead ECG using Variational Autoencoder to Improve Cardiac Disease Detection of Wearable ECG Devices
por: Guan, Xinyan, et al.
Publicado: (2025)
por: Guan, Xinyan, et al.
Publicado: (2025)
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
por: Marmoret, Axel, et al.
Publicado: (2025)
por: Marmoret, Axel, et al.
Publicado: (2025)
RPRA: Predicting an LLM-Judge for Efficient but Performant Inference
por: Ashley, Dylan R., et al.
Publicado: (2026)
por: Ashley, Dylan R., et al.
Publicado: (2026)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
por: Fu, Tianyu, et al.
Publicado: (2025)
por: Fu, Tianyu, et al.
Publicado: (2025)
Modularity in Transformers: Investigating Neuron Separability & Specialization
por: Pochinkov, Nicholas, et al.
Publicado: (2024)
por: Pochinkov, Nicholas, et al.
Publicado: (2024)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
por: Lee, Wooin, et al.
Publicado: (2026)
por: Lee, Wooin, et al.
Publicado: (2026)
The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
por: Wang, Zhixiang
Publicado: (2025)
por: Wang, Zhixiang
Publicado: (2025)
Lost or Hidden? A Concept-Level Forgetting in Supervised Continual Learning
por: Filus, Katarzyna, et al.
Publicado: (2026)
por: Filus, Katarzyna, et al.
Publicado: (2026)
Ejemplares similares
-
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
por: Henry, James
Publicado: (2026) -
A Practical Guide to Streaming Continual Learning
por: Cossu, Andrea, et al.
Publicado: (2026) -
Why Geometric Continuity Emerges in Deep Neural Networks: Residual Connections and Rotational Symmetry Breaking
por: Jeong, Kyungwon, et al.
Publicado: (2026) -
cPNN: Continuous Progressive Neural Networks for Evolving Streaming Time Series
por: Giannini, Federico, et al.
Publicado: (2026) -
Don't Look Back in Anger: MAGIC Net for Streaming Continual Learning with Temporal Dependence
por: Giannini, Federico, et al.
Publicado: (2026)