Salvato in:
| Autori principali: | Musil, Tomáš, Mareček, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2409.19813 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Exploring Interpretability of Independent Components of Word Embeddings with Automated Word Intruder Test
di: Musil, Tomáš, et al.
Pubblicazione: (2022)
di: Musil, Tomáš, et al.
Pubblicazione: (2022)
Dual Debiasing: Remove Stereotypes and Keep Factual Gender for Fair Language Modeling and Translation
di: Limisiewicz, Tomasz, et al.
Pubblicazione: (2025)
di: Limisiewicz, Tomasz, et al.
Pubblicazione: (2025)
Debiasing Algorithm through Model Adaptation
di: Limisiewicz, Tomasz, et al.
Pubblicazione: (2023)
di: Limisiewicz, Tomasz, et al.
Pubblicazione: (2023)
Suppressing Final Layer Hidden State Jumps in Transformer Pretraining
di: Shibata, Keigo, et al.
Pubblicazione: (2026)
di: Shibata, Keigo, et al.
Pubblicazione: (2026)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
di: Chen, Junhao, et al.
Pubblicazione: (2024)
di: Chen, Junhao, et al.
Pubblicazione: (2024)
Teaching LLMs at Charles University: Assignments and Activities
di: Helcl, Jindřich, et al.
Pubblicazione: (2024)
di: Helcl, Jindřich, et al.
Pubblicazione: (2024)
Mixture of Hidden-Dimensions Transformer
di: Chen, Yilong, et al.
Pubblicazione: (2024)
di: Chen, Yilong, et al.
Pubblicazione: (2024)
The Hidden Space of Transformer Language Adapters
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
di: Schiekiera, Louis, et al.
Pubblicazione: (2026)
di: Schiekiera, Louis, et al.
Pubblicazione: (2026)
LLM-based Embeddings: Attention Values Encode Sentence Semantics Better Than Hidden States
di: Zhang, Yeqin, et al.
Pubblicazione: (2026)
di: Zhang, Yeqin, et al.
Pubblicazione: (2026)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
di: Jiang, Yilei, et al.
Pubblicazione: (2025)
di: Jiang, Yilei, et al.
Pubblicazione: (2025)
AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue
di: Park, Jihyung, et al.
Pubblicazione: (2026)
di: Park, Jihyung, et al.
Pubblicazione: (2026)
Mechanistic Interpretability of Binary and Ternary Transformers
di: Li, Jason
Pubblicazione: (2024)
di: Li, Jason
Pubblicazione: (2024)
Future Lens: Anticipating Subsequent Tokens from a Single Hidden State
di: Pal, Koyena, et al.
Pubblicazione: (2023)
di: Pal, Koyena, et al.
Pubblicazione: (2023)
What Am I Missing? Question-Answering as Hidden State Probing
di: Luo, Chu Fei, et al.
Pubblicazione: (2026)
di: Luo, Chu Fei, et al.
Pubblicazione: (2026)
Skill-RAG: Failure-State-Aware Retrieval Augmentation via Hidden-State Probing and Skill Routing
di: Wei, Kai, et al.
Pubblicazione: (2026)
di: Wei, Kai, et al.
Pubblicazione: (2026)
Do LLMs Know about Hallucination? An Empirical Investigation of LLM's Hidden States
di: Duan, Hanyu, et al.
Pubblicazione: (2024)
di: Duan, Hanyu, et al.
Pubblicazione: (2024)
CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
di: Liang, Zhenwen, et al.
Pubblicazione: (2025)
di: Liang, Zhenwen, et al.
Pubblicazione: (2025)
Evolutionary Feature-wise Thresholding for Binary Representation of NLP Embeddings
di: Sinha, Soumen, et al.
Pubblicazione: (2025)
di: Sinha, Soumen, et al.
Pubblicazione: (2025)
Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models
di: Bajaj, Anooshka, et al.
Pubblicazione: (2025)
di: Bajaj, Anooshka, et al.
Pubblicazione: (2025)
Improving Interpretability of Lexical Semantic Change with Neurobiological Features
di: Oda, Kohei, et al.
Pubblicazione: (2026)
di: Oda, Kohei, et al.
Pubblicazione: (2026)
LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers
di: Razzhigaev, Anton, et al.
Pubblicazione: (2025)
di: Razzhigaev, Anton, et al.
Pubblicazione: (2025)
Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
di: Yang, Haolin, et al.
Pubblicazione: (2025)
di: Yang, Haolin, et al.
Pubblicazione: (2025)
CLASP: Defending Hybrid Large Language Models Against Hidden State Poisoning Attacks
di: Mercier, Alexandre Le, et al.
Pubblicazione: (2026)
di: Mercier, Alexandre Le, et al.
Pubblicazione: (2026)
Semformer: Transformer Language Models with Semantic Planning
di: Yin, Yongjing, et al.
Pubblicazione: (2024)
di: Yin, Yongjing, et al.
Pubblicazione: (2024)
TRACE for Tracking the Emergence of Semantic Representations in Transformers
di: Aljaafari, Nura, et al.
Pubblicazione: (2025)
di: Aljaafari, Nura, et al.
Pubblicazione: (2025)
Explicit Grammar Semantic Feature Fusion for Robust Text Classification
di: Sultana, Azrin, et al.
Pubblicazione: (2026)
di: Sultana, Azrin, et al.
Pubblicazione: (2026)
Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
di: Yang, Rui, et al.
Pubblicazione: (2024)
di: Yang, Rui, et al.
Pubblicazione: (2024)
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2024)
di: Chanin, David, et al.
Pubblicazione: (2024)
Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs
di: Sternfeld, Alexander, et al.
Pubblicazione: (2025)
di: Sternfeld, Alexander, et al.
Pubblicazione: (2025)
Inside-Out: Hidden Factual Knowledge in LLMs
di: Gekhman, Zorik, et al.
Pubblicazione: (2025)
di: Gekhman, Zorik, et al.
Pubblicazione: (2025)
LLM Hallucination Detection: A Fast Fourier Transform Method Based on Hidden Layer Temporal Signals
di: Li, Jinxin, et al.
Pubblicazione: (2025)
di: Li, Jinxin, et al.
Pubblicazione: (2025)
Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach
di: Zhang, Chong, et al.
Pubblicazione: (2025)
di: Zhang, Chong, et al.
Pubblicazione: (2025)
When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding?
di: Liu, Tianyu, et al.
Pubblicazione: (2026)
di: Liu, Tianyu, et al.
Pubblicazione: (2026)
Comateformer: Combined Attention Transformer for Semantic Sentence Matching
di: Li, Bo, et al.
Pubblicazione: (2024)
di: Li, Bo, et al.
Pubblicazione: (2024)
Improving Multilingual Neural Machine Translation by Utilizing Semantic and Linguistic Features
di: Bu, Mengyu, et al.
Pubblicazione: (2024)
di: Bu, Mengyu, et al.
Pubblicazione: (2024)
CSF: Contrastive Semantic Features for Direct Multilingual Sign Language Generation
di: Bao, Tran Sy
Pubblicazione: (2026)
di: Bao, Tran Sy
Pubblicazione: (2026)
Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
di: Boggust, Angie, et al.
Pubblicazione: (2025)
di: Boggust, Angie, et al.
Pubblicazione: (2025)
Transformers are Multi-State RNNs
di: Oren, Matanel, et al.
Pubblicazione: (2024)
di: Oren, Matanel, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Exploring Interpretability of Independent Components of Word Embeddings with Automated Word Intruder Test
di: Musil, Tomáš, et al.
Pubblicazione: (2022) -
Dual Debiasing: Remove Stereotypes and Keep Factual Gender for Fair Language Modeling and Translation
di: Limisiewicz, Tomasz, et al.
Pubblicazione: (2025) -
Debiasing Algorithm through Model Adaptation
di: Limisiewicz, Tomasz, et al.
Pubblicazione: (2023) -
Suppressing Final Layer Hidden State Jumps in Transformer Pretraining
di: Shibata, Keigo, et al.
Pubblicazione: (2026) -
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
di: Chen, Junhao, et al.
Pubblicazione: (2024)