From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Schiekiera, Louis, Zimmer, Max, Roux, Christophe, Pokutta, Sebastian, Günther, Fritz |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On the Byzantine-Resilience of Distillation-Based Federated Learning
por: Roux, Christophe, et al.
Publicado: (2024)
por: Roux, Christophe, et al.
Publicado: (2024)
The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning
por: Zimmer, Max, et al.
Publicado: (2026)
por: Zimmer, Max, et al.
Publicado: (2026)
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
por: Wagner, Moritz, et al.
Publicado: (2025)
por: Wagner, Moritz, et al.
Publicado: (2025)
SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
por: Zimmer, Max, et al.
Publicado: (2025)
por: Zimmer, Max, et al.
Publicado: (2025)
PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs
por: Zimmer, Max, et al.
Publicado: (2023)
por: Zimmer, Max, et al.
Publicado: (2023)
Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging
por: Zimmer, Max, et al.
Publicado: (2023)
por: Zimmer, Max, et al.
Publicado: (2023)
Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
por: Roux, Christophe, et al.
Publicado: (2025)
por: Roux, Christophe, et al.
Publicado: (2025)
Neural Sum-of-Squares: Certifying the Nonnegativity of Polynomials with Transformers
por: Pelleriti, Nico, et al.
Publicado: (2025)
por: Pelleriti, Nico, et al.
Publicado: (2025)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
por: Mezentsev, Gleb, et al.
Publicado: (2025)
por: Mezentsev, Gleb, et al.
Publicado: (2025)
Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
por: Bianchi, Owen, et al.
Publicado: (2025)
por: Bianchi, Owen, et al.
Publicado: (2025)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
por: Karvonen, Adam, et al.
Publicado: (2025)
por: Karvonen, Adam, et al.
Publicado: (2025)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
por: Lei, Ge, et al.
Publicado: (2025)
por: Lei, Ge, et al.
Publicado: (2025)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
por: Turk, Matt
Publicado: (2026)
por: Turk, Matt
Publicado: (2026)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
por: Wu, Xuansheng, et al.
Publicado: (2023)
por: Wu, Xuansheng, et al.
Publicado: (2023)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
por: Sun, Yu, et al.
Publicado: (2024)
por: Sun, Yu, et al.
Publicado: (2024)
When Chain-of-Thought Fails, the Solution Hides in the Hidden States
por: Mehrafarin, Houman, et al.
Publicado: (2026)
por: Mehrafarin, Houman, et al.
Publicado: (2026)
Semantic Refinement with LLMs for Graph Representations
por: Thapaliya, Safal, et al.
Publicado: (2025)
por: Thapaliya, Safal, et al.
Publicado: (2025)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
por: Baumann, Joachim, et al.
Publicado: (2025)
por: Baumann, Joachim, et al.
Publicado: (2025)
Extracting Unlearned Information from LLMs with Activation Steering
por: Seyitoğlu, Atakan, et al.
Publicado: (2024)
por: Seyitoğlu, Atakan, et al.
Publicado: (2024)
Hidden State Poisoning Attacks against Mamba-based Language Models
por: Mercier, Alexandre Le, et al.
Publicado: (2026)
por: Mercier, Alexandre Le, et al.
Publicado: (2026)
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
por: Flamant, Cedric, et al.
Publicado: (2026)
por: Flamant, Cedric, et al.
Publicado: (2026)
The Geometries of Truth Are Orthogonal Across Tasks
por: Azizian, Waiss, et al.
Publicado: (2025)
por: Azizian, Waiss, et al.
Publicado: (2025)
Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
por: Bozoukov, Matthew, et al.
Publicado: (2025)
por: Bozoukov, Matthew, et al.
Publicado: (2025)
Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs
por: Yang, Hongming, et al.
Publicado: (2025)
por: Yang, Hongming, et al.
Publicado: (2025)
The Remarkable Robustness of LLMs: Stages of Inference?
por: Lad, Vedang, et al.
Publicado: (2024)
por: Lad, Vedang, et al.
Publicado: (2024)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
por: Shen, Xuan, et al.
Publicado: (2023)
por: Shen, Xuan, et al.
Publicado: (2023)
LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
por: Lugoloobi, William, et al.
Publicado: (2026)
por: Lugoloobi, William, et al.
Publicado: (2026)
Interpretability Guarantees with Merlin-Arthur Classifiers
por: Wäldchen, Stephan, et al.
Publicado: (2022)
por: Wäldchen, Stephan, et al.
Publicado: (2022)
Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
por: Guo, Jizhou, et al.
Publicado: (2025)
por: Guo, Jizhou, et al.
Publicado: (2025)
Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs
por: Ovadia, Oded, et al.
Publicado: (2023)
por: Ovadia, Oded, et al.
Publicado: (2023)
Hypertokens: Holographic Associative Memory in Tokenized LLMs
por: Augeri, Christopher James
Publicado: (2025)
por: Augeri, Christopher James
Publicado: (2025)
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks
por: Soni, Nikita, et al.
Publicado: (2025)
por: Soni, Nikita, et al.
Publicado: (2025)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
por: Kossen, Jannik, et al.
Publicado: (2024)
por: Kossen, Jannik, et al.
Publicado: (2024)
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
por: Yang, Zhuonan, et al.
Publicado: (2026)
por: Yang, Zhuonan, et al.
Publicado: (2026)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
por: Arbabi, Alireza, et al.
Publicado: (2025)
por: Arbabi, Alireza, et al.
Publicado: (2025)
Sectoral Coupling in Linguistic State Space
por: Dumbrava, Sebastian
Publicado: (2025)
por: Dumbrava, Sebastian
Publicado: (2025)
Do Multilingual LLMs Think In English?
por: Schut, Lisa, et al.
Publicado: (2025)
por: Schut, Lisa, et al.
Publicado: (2025)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
por: Sheshadri, Abhay, et al.
Publicado: (2024)
por: Sheshadri, Abhay, et al.
Publicado: (2024)
Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
por: Huang, Bingning, et al.
Publicado: (2025)
por: Huang, Bingning, et al.
Publicado: (2025)
Capturing Temporal Dynamics in Large-Scale Canopy Tree Height Estimation
por: Pauls, Jan, et al.
Publicado: (2025)
por: Pauls, Jan, et al.
Publicado: (2025)
Ejemplares similares
-
On the Byzantine-Resilience of Distillation-Based Federated Learning
por: Roux, Christophe, et al.
Publicado: (2024) -
The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning
por: Zimmer, Max, et al.
Publicado: (2026) -
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
por: Wagner, Moritz, et al.
Publicado: (2025) -
SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
por: Zimmer, Max, et al.
Publicado: (2025) -
PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs
por: Zimmer, Max, et al.
Publicado: (2023)