Transforming Hidden States into Binary Semantic Features
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Musil, Tomáš, Mareček, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring Interpretability of Independent Components of Word Embeddings with Automated Word Intruder Test
von: Musil, Tomáš, et al.
Veröffentlicht: (2022)
von: Musil, Tomáš, et al.
Veröffentlicht: (2022)
Dual Debiasing: Remove Stereotypes and Keep Factual Gender for Fair Language Modeling and Translation
von: Limisiewicz, Tomasz, et al.
Veröffentlicht: (2025)
von: Limisiewicz, Tomasz, et al.
Veröffentlicht: (2025)
Debiasing Algorithm through Model Adaptation
von: Limisiewicz, Tomasz, et al.
Veröffentlicht: (2023)
von: Limisiewicz, Tomasz, et al.
Veröffentlicht: (2023)
Suppressing Final Layer Hidden State Jumps in Transformer Pretraining
von: Shibata, Keigo, et al.
Veröffentlicht: (2026)
von: Shibata, Keigo, et al.
Veröffentlicht: (2026)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
von: Chen, Junhao, et al.
Veröffentlicht: (2024)
von: Chen, Junhao, et al.
Veröffentlicht: (2024)
Mixture of Hidden-Dimensions Transformer
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
The Hidden Space of Transformer Language Adapters
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
Teaching LLMs at Charles University: Assignments and Activities
von: Helcl, Jindřich, et al.
Veröffentlicht: (2024)
von: Helcl, Jindřich, et al.
Veröffentlicht: (2024)
LLM-based Embeddings: Attention Values Encode Sentence Semantics Better Than Hidden States
von: Zhang, Yeqin, et al.
Veröffentlicht: (2026)
von: Zhang, Yeqin, et al.
Veröffentlicht: (2026)
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
von: Schiekiera, Louis, et al.
Veröffentlicht: (2026)
von: Schiekiera, Louis, et al.
Veröffentlicht: (2026)
HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
von: Jiang, Yilei, et al.
Veröffentlicht: (2025)
von: Jiang, Yilei, et al.
Veröffentlicht: (2025)
AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue
von: Park, Jihyung, et al.
Veröffentlicht: (2026)
von: Park, Jihyung, et al.
Veröffentlicht: (2026)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025)
von: Chanin, David, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Binary and Ternary Transformers
von: Li, Jason
Veröffentlicht: (2024)
von: Li, Jason
Veröffentlicht: (2024)
What Am I Missing? Question-Answering as Hidden State Probing
von: Luo, Chu Fei, et al.
Veröffentlicht: (2026)
von: Luo, Chu Fei, et al.
Veröffentlicht: (2026)
Future Lens: Anticipating Subsequent Tokens from a Single Hidden State
von: Pal, Koyena, et al.
Veröffentlicht: (2023)
von: Pal, Koyena, et al.
Veröffentlicht: (2023)
Skill-RAG: Failure-State-Aware Retrieval Augmentation via Hidden-State Probing and Skill Routing
von: Wei, Kai, et al.
Veröffentlicht: (2026)
von: Wei, Kai, et al.
Veröffentlicht: (2026)
Do LLMs Know about Hallucination? An Empirical Investigation of LLM's Hidden States
von: Duan, Hanyu, et al.
Veröffentlicht: (2024)
von: Duan, Hanyu, et al.
Veröffentlicht: (2024)
CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
Improving Interpretability of Lexical Semantic Change with Neurobiological Features
von: Oda, Kohei, et al.
Veröffentlicht: (2026)
von: Oda, Kohei, et al.
Veröffentlicht: (2026)
Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models
von: Bajaj, Anooshka, et al.
Veröffentlicht: (2025)
von: Bajaj, Anooshka, et al.
Veröffentlicht: (2025)
Evolutionary Feature-wise Thresholding for Binary Representation of NLP Embeddings
von: Sinha, Soumen, et al.
Veröffentlicht: (2025)
von: Sinha, Soumen, et al.
Veröffentlicht: (2025)
Semformer: Transformer Language Models with Semantic Planning
von: Yin, Yongjing, et al.
Veröffentlicht: (2024)
von: Yin, Yongjing, et al.
Veröffentlicht: (2024)
TRACE for Tracking the Emergence of Semantic Representations in Transformers
von: Aljaafari, Nura, et al.
Veröffentlicht: (2025)
von: Aljaafari, Nura, et al.
Veröffentlicht: (2025)
Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
CLASP: Defending Hybrid Large Language Models Against Hidden State Poisoning Attacks
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
Explicit Grammar Semantic Feature Fusion for Robust Text Classification
von: Sultana, Azrin, et al.
Veröffentlicht: (2026)
von: Sultana, Azrin, et al.
Veröffentlicht: (2026)
LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers
von: Razzhigaev, Anton, et al.
Veröffentlicht: (2025)
von: Razzhigaev, Anton, et al.
Veröffentlicht: (2025)
LLM Hallucination Detection: A Fast Fourier Transform Method Based on Hidden Layer Temporal Signals
von: Li, Jinxin, et al.
Veröffentlicht: (2025)
von: Li, Jinxin, et al.
Veröffentlicht: (2025)
Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2025)
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2025)
When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding?
von: Liu, Tianyu, et al.
Veröffentlicht: (2026)
von: Liu, Tianyu, et al.
Veröffentlicht: (2026)
Comateformer: Combined Attention Transformer for Semantic Sentence Matching
von: Li, Bo, et al.
Veröffentlicht: (2024)
von: Li, Bo, et al.
Veröffentlicht: (2024)
Inside-Out: Hidden Factual Knowledge in LLMs
von: Gekhman, Zorik, et al.
Veröffentlicht: (2025)
von: Gekhman, Zorik, et al.
Veröffentlicht: (2025)
Improving Multilingual Neural Machine Translation by Utilizing Semantic and Linguistic Features
von: Bu, Mengyu, et al.
Veröffentlicht: (2024)
von: Bu, Mengyu, et al.
Veröffentlicht: (2024)
CSF: Contrastive Semantic Features for Direct Multilingual Sign Language Generation
von: Bao, Tran Sy
Veröffentlicht: (2026)
von: Bao, Tran Sy
Veröffentlicht: (2026)
Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
von: Boggust, Angie, et al.
Veröffentlicht: (2025)
von: Boggust, Angie, et al.
Veröffentlicht: (2025)
Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
von: Yang, Rui, et al.
Veröffentlicht: (2024)
von: Yang, Rui, et al.
Veröffentlicht: (2024)
Transformers are Multi-State RNNs
von: Oren, Matanel, et al.
Veröffentlicht: (2024)
von: Oren, Matanel, et al.
Veröffentlicht: (2024)
Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach
von: Zhang, Chong, et al.
Veröffentlicht: (2025)
von: Zhang, Chong, et al.
Veröffentlicht: (2025)
Semantics of Multiword Expressions in Transformer-Based Models: A Survey
von: Miletić, Filip, et al.
Veröffentlicht: (2024)
von: Miletić, Filip, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Exploring Interpretability of Independent Components of Word Embeddings with Automated Word Intruder Test
von: Musil, Tomáš, et al.
Veröffentlicht: (2022) -
Dual Debiasing: Remove Stereotypes and Keep Factual Gender for Fair Language Modeling and Translation
von: Limisiewicz, Tomasz, et al.
Veröffentlicht: (2025) -
Debiasing Algorithm through Model Adaptation
von: Limisiewicz, Tomasz, et al.
Veröffentlicht: (2023) -
Suppressing Final Layer Hidden State Jumps in Transformer Pretraining
von: Shibata, Keigo, et al.
Veröffentlicht: (2026) -
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
von: Chen, Junhao, et al.
Veröffentlicht: (2024)