Lost in State Space: Probing Frozen Mamba Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wagh, Bhagyashree, Singh, Akash |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MambaByte: Token-free Selective State Space Model
von: Wang, Junxiong, et al.
Veröffentlicht: (2024)
von: Wang, Junxiong, et al.
Veröffentlicht: (2024)
AtteSTNet -- An attention and subword tokenization based approach for code-switched text hate speech detection
von: Shingi, Geet, et al.
Veröffentlicht: (2021)
von: Shingi, Geet, et al.
Veröffentlicht: (2021)
MatMamba: A Matryoshka State Space Model
von: Shukla, Abhinav, et al.
Veröffentlicht: (2024)
von: Shukla, Abhinav, et al.
Veröffentlicht: (2024)
MemMamba: Rethinking Memory Patterns in State Space Model
von: Wang, Youjin, et al.
Veröffentlicht: (2025)
von: Wang, Youjin, et al.
Veröffentlicht: (2025)
DenseMamba: State Space Models with Dense Hidden Connection for Efficient Large Language Models
von: He, Wei, et al.
Veröffentlicht: (2024)
von: He, Wei, et al.
Veröffentlicht: (2024)
Characterizing the Behavior of Training Mamba-based State Space Models on GPUs
von: Baruah, Trinayan, et al.
Veröffentlicht: (2025)
von: Baruah, Trinayan, et al.
Veröffentlicht: (2025)
Detection vs. Execution: Single-Bucket Probes Miss Half the Mamba-2 State Sink
von: Jiang, Yuhang
Veröffentlicht: (2026)
von: Jiang, Yuhang
Veröffentlicht: (2026)
The UNDO Flip-Flop: A Controlled Probe for Reversible Semantic State Management in State Space Model
von: Zhou, Hongxu
Veröffentlicht: (2026)
von: Zhou, Hongxu
Veröffentlicht: (2026)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
von: Pióro, Maciej, et al.
Veröffentlicht: (2024)
von: Pióro, Maciej, et al.
Veröffentlicht: (2024)
BlackMamba: Mixture of Experts for State-Space Models
von: Anthony, Quentin, et al.
Veröffentlicht: (2024)
von: Anthony, Quentin, et al.
Veröffentlicht: (2024)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
von: Chen, Yifang, et al.
Veröffentlicht: (2024)
von: Chen, Yifang, et al.
Veröffentlicht: (2024)
A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
von: Hoang, Nhat M., et al.
Veröffentlicht: (2025)
von: Hoang, Nhat M., et al.
Veröffentlicht: (2025)
MoFE: Mixture of Frozen Experts Architecture
von: Seo, Jean, et al.
Veröffentlicht: (2025)
von: Seo, Jean, et al.
Veröffentlicht: (2025)
An Empirical Study of Mamba-based Language Models
von: Waleffe, Roger, et al.
Veröffentlicht: (2024)
von: Waleffe, Roger, et al.
Veröffentlicht: (2024)
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
von: Simoncini, Walter, et al.
Veröffentlicht: (2024)
von: Simoncini, Walter, et al.
Veröffentlicht: (2024)
Comparative Study of Pre-Trained BERT and Large Language Models for Code-Mixed Named Entity Recognition
von: Shirke, Mayur, et al.
Veröffentlicht: (2025)
von: Shirke, Mayur, et al.
Veröffentlicht: (2025)
On Importance of Layer Pruning for Smaller BERT Models and Low Resource Languages
von: Shirke, Mayur, et al.
Veröffentlicht: (2025)
von: Shirke, Mayur, et al.
Veröffentlicht: (2025)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
von: Liang, Weixin, et al.
Veröffentlicht: (2025)
von: Liang, Weixin, et al.
Veröffentlicht: (2025)
Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate
von: Bochkov, A.
Veröffentlicht: (2025)
von: Bochkov, A.
Veröffentlicht: (2025)
On Pruning State-Space LLMs
von: Ghattas, Tamer, et al.
Veröffentlicht: (2025)
von: Ghattas, Tamer, et al.
Veröffentlicht: (2025)
Style Classification of Rabbinic Literature for Detection of Lost Midrash Tanhuma Material
von: Tannor, Shlomo, et al.
Veröffentlicht: (2022)
von: Tannor, Shlomo, et al.
Veröffentlicht: (2022)
Monitoring Latent World States in Language Models with Propositional Probes
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
R2Gen-Mamba: A Selective State Space Model for Radiology Report Generation
von: Sun, Yongheng, et al.
Veröffentlicht: (2024)
von: Sun, Yongheng, et al.
Veröffentlicht: (2024)
ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings
von: Hao, Shibo, et al.
Veröffentlicht: (2023)
von: Hao, Shibo, et al.
Veröffentlicht: (2023)
Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models
von: Muñoz, J. Pablo, et al.
Veröffentlicht: (2025)
von: Muñoz, J. Pablo, et al.
Veröffentlicht: (2025)
Hidden State Poisoning Attacks against Mamba-based Language Models
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
Detection Without Correction: A Robust Asymmetry in Activation-Based Hallucination Probing
von: Roy, Dip, et al.
Veröffentlicht: (2026)
von: Roy, Dip, et al.
Veröffentlicht: (2026)
Mamba Knockout for Unraveling Factual Information Flow
von: Endy, Nir, et al.
Veröffentlicht: (2025)
von: Endy, Nir, et al.
Veröffentlicht: (2025)
Rethinking Token Reduction for State Space Models
von: Zhan, Zheng, et al.
Veröffentlicht: (2024)
von: Zhan, Zheng, et al.
Veröffentlicht: (2024)
Borrowed Geometry: Cross-Distribution Head-Importance Fingerprints of Frozen Pretrained Gemma 4 31B
von: Bektursun, Abay
Veröffentlicht: (2026)
von: Bektursun, Abay
Veröffentlicht: (2026)
Low-Dimensional Structure in the Space of Language Representations is Reflected in Brain Responses
von: Antonello, Richard, et al.
Veröffentlicht: (2021)
von: Antonello, Richard, et al.
Veröffentlicht: (2021)
Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps
von: Chua, Gabriel, et al.
Veröffentlicht: (2025)
von: Chua, Gabriel, et al.
Veröffentlicht: (2025)
The Illusion of State in State-Space Models
von: Merrill, William, et al.
Veröffentlicht: (2024)
von: Merrill, William, et al.
Veröffentlicht: (2024)
Differential Mamba
von: Schneider, Nadav, et al.
Veröffentlicht: (2025)
von: Schneider, Nadav, et al.
Veröffentlicht: (2025)
Jamba: A Hybrid Transformer-Mamba Language Model
von: Lieber, Opher, et al.
Veröffentlicht: (2024)
von: Lieber, Opher, et al.
Veröffentlicht: (2024)
Parameter-Efficient Fine-Tuning of State Space Models
von: Galim, Kevin, et al.
Veröffentlicht: (2024)
von: Galim, Kevin, et al.
Veröffentlicht: (2024)
Geometric Organization of Cognitive States in Transformer Embedding Spaces
von: Zhao, Sophie
Veröffentlicht: (2025)
von: Zhao, Sophie
Veröffentlicht: (2025)
Rhetorical Questions in LLM Representations: A Linear Probing Study
von: Yao, Louie Hong, et al.
Veröffentlicht: (2026)
von: Yao, Louie Hong, et al.
Veröffentlicht: (2026)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MambaByte: Token-free Selective State Space Model
von: Wang, Junxiong, et al.
Veröffentlicht: (2024) -
AtteSTNet -- An attention and subword tokenization based approach for code-switched text hate speech detection
von: Shingi, Geet, et al.
Veröffentlicht: (2021) -
MatMamba: A Matryoshka State Space Model
von: Shukla, Abhinav, et al.
Veröffentlicht: (2024) -
MemMamba: Rethinking Memory Patterns in State Space Model
von: Wang, Youjin, et al.
Veröffentlicht: (2025) -
DenseMamba: State Space Models with Dense Hidden Connection for Efficient Large Language Models
von: He, Wei, et al.
Veröffentlicht: (2024)