Saved in:
| Main Authors: | Musil, Tomáš, Mareček, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.19813 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Interpretability of Independent Components of Word Embeddings with Automated Word Intruder Test
by: Musil, Tomáš, et al.
Published: (2022)
by: Musil, Tomáš, et al.
Published: (2022)
Dual Debiasing: Remove Stereotypes and Keep Factual Gender for Fair Language Modeling and Translation
by: Limisiewicz, Tomasz, et al.
Published: (2025)
by: Limisiewicz, Tomasz, et al.
Published: (2025)
Debiasing Algorithm through Model Adaptation
by: Limisiewicz, Tomasz, et al.
Published: (2023)
by: Limisiewicz, Tomasz, et al.
Published: (2023)
Suppressing Final Layer Hidden State Jumps in Transformer Pretraining
by: Shibata, Keigo, et al.
Published: (2026)
by: Shibata, Keigo, et al.
Published: (2026)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
by: Chen, Junhao, et al.
Published: (2024)
by: Chen, Junhao, et al.
Published: (2024)
Teaching LLMs at Charles University: Assignments and Activities
by: Helcl, Jindřich, et al.
Published: (2024)
by: Helcl, Jindřich, et al.
Published: (2024)
Mixture of Hidden-Dimensions Transformer
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
The Hidden Space of Transformer Language Adapters
by: Alabi, Jesujoba O., et al.
Published: (2024)
by: Alabi, Jesujoba O., et al.
Published: (2024)
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
by: Schiekiera, Louis, et al.
Published: (2026)
by: Schiekiera, Louis, et al.
Published: (2026)
LLM-based Embeddings: Attention Values Encode Sentence Semantics Better Than Hidden States
by: Zhang, Yeqin, et al.
Published: (2026)
by: Zhang, Yeqin, et al.
Published: (2026)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
by: Chanin, David, et al.
Published: (2025)
by: Chanin, David, et al.
Published: (2025)
HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
by: Jiang, Yilei, et al.
Published: (2025)
by: Jiang, Yilei, et al.
Published: (2025)
AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue
by: Park, Jihyung, et al.
Published: (2026)
by: Park, Jihyung, et al.
Published: (2026)
Mechanistic Interpretability of Binary and Ternary Transformers
by: Li, Jason
Published: (2024)
by: Li, Jason
Published: (2024)
Future Lens: Anticipating Subsequent Tokens from a Single Hidden State
by: Pal, Koyena, et al.
Published: (2023)
by: Pal, Koyena, et al.
Published: (2023)
What Am I Missing? Question-Answering as Hidden State Probing
by: Luo, Chu Fei, et al.
Published: (2026)
by: Luo, Chu Fei, et al.
Published: (2026)
Skill-RAG: Failure-State-Aware Retrieval Augmentation via Hidden-State Probing and Skill Routing
by: Wei, Kai, et al.
Published: (2026)
by: Wei, Kai, et al.
Published: (2026)
Do LLMs Know about Hallucination? An Empirical Investigation of LLM's Hidden States
by: Duan, Hanyu, et al.
Published: (2024)
by: Duan, Hanyu, et al.
Published: (2024)
CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
by: Liang, Zhenwen, et al.
Published: (2025)
by: Liang, Zhenwen, et al.
Published: (2025)
Evolutionary Feature-wise Thresholding for Binary Representation of NLP Embeddings
by: Sinha, Soumen, et al.
Published: (2025)
by: Sinha, Soumen, et al.
Published: (2025)
Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models
by: Bajaj, Anooshka, et al.
Published: (2025)
by: Bajaj, Anooshka, et al.
Published: (2025)
Improving Interpretability of Lexical Semantic Change with Neurobiological Features
by: Oda, Kohei, et al.
Published: (2026)
by: Oda, Kohei, et al.
Published: (2026)
LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers
by: Razzhigaev, Anton, et al.
Published: (2025)
by: Razzhigaev, Anton, et al.
Published: (2025)
Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
CLASP: Defending Hybrid Large Language Models Against Hidden State Poisoning Attacks
by: Mercier, Alexandre Le, et al.
Published: (2026)
by: Mercier, Alexandre Le, et al.
Published: (2026)
Semformer: Transformer Language Models with Semantic Planning
by: Yin, Yongjing, et al.
Published: (2024)
by: Yin, Yongjing, et al.
Published: (2024)
TRACE for Tracking the Emergence of Semantic Representations in Transformers
by: Aljaafari, Nura, et al.
Published: (2025)
by: Aljaafari, Nura, et al.
Published: (2025)
Explicit Grammar Semantic Feature Fusion for Robust Text Classification
by: Sultana, Azrin, et al.
Published: (2026)
by: Sultana, Azrin, et al.
Published: (2026)
Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
by: Yang, Rui, et al.
Published: (2024)
by: Yang, Rui, et al.
Published: (2024)
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
by: Chanin, David, et al.
Published: (2024)
by: Chanin, David, et al.
Published: (2024)
Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs
by: Sternfeld, Alexander, et al.
Published: (2025)
by: Sternfeld, Alexander, et al.
Published: (2025)
Inside-Out: Hidden Factual Knowledge in LLMs
by: Gekhman, Zorik, et al.
Published: (2025)
by: Gekhman, Zorik, et al.
Published: (2025)
LLM Hallucination Detection: A Fast Fourier Transform Method Based on Hidden Layer Temporal Signals
by: Li, Jinxin, et al.
Published: (2025)
by: Li, Jinxin, et al.
Published: (2025)
Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach
by: Zhang, Chong, et al.
Published: (2025)
by: Zhang, Chong, et al.
Published: (2025)
When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding?
by: Liu, Tianyu, et al.
Published: (2026)
by: Liu, Tianyu, et al.
Published: (2026)
Comateformer: Combined Attention Transformer for Semantic Sentence Matching
by: Li, Bo, et al.
Published: (2024)
by: Li, Bo, et al.
Published: (2024)
Improving Multilingual Neural Machine Translation by Utilizing Semantic and Linguistic Features
by: Bu, Mengyu, et al.
Published: (2024)
by: Bu, Mengyu, et al.
Published: (2024)
CSF: Contrastive Semantic Features for Direct Multilingual Sign Language Generation
by: Bao, Tran Sy
Published: (2026)
by: Bao, Tran Sy
Published: (2026)
Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
by: Boggust, Angie, et al.
Published: (2025)
by: Boggust, Angie, et al.
Published: (2025)
Transformers are Multi-State RNNs
by: Oren, Matanel, et al.
Published: (2024)
by: Oren, Matanel, et al.
Published: (2024)
Similar Items
-
Exploring Interpretability of Independent Components of Word Embeddings with Automated Word Intruder Test
by: Musil, Tomáš, et al.
Published: (2022) -
Dual Debiasing: Remove Stereotypes and Keep Factual Gender for Fair Language Modeling and Translation
by: Limisiewicz, Tomasz, et al.
Published: (2025) -
Debiasing Algorithm through Model Adaptation
by: Limisiewicz, Tomasz, et al.
Published: (2023) -
Suppressing Final Layer Hidden State Jumps in Transformer Pretraining
by: Shibata, Keigo, et al.
Published: (2026) -
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
by: Chen, Junhao, et al.
Published: (2024)