Detection vs. Execution: Single-Bucket Probes Miss Half the Mamba-2 State Sink
Fuente:
arXiv
Guardado en:
| Autor principal: | Jiang, Yuhang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On the Existence and Behavior of Secondary Attention Sinks
por: Wong, Jeffrey T. H., et al.
Publicado: (2025)
por: Wong, Jeffrey T. H., et al.
Publicado: (2025)
Sink-Aware Pruning for Diffusion Language Models
por: Myrzakhan, Aidar, et al.
Publicado: (2026)
por: Myrzakhan, Aidar, et al.
Publicado: (2026)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
por: Chen, Yingfa, et al.
Publicado: (2024)
por: Chen, Yingfa, et al.
Publicado: (2024)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
por: Xu, Zukang, et al.
Publicado: (2025)
por: Xu, Zukang, et al.
Publicado: (2025)
MemMamba: Rethinking Memory Patterns in State Space Model
por: Wang, Youjin, et al.
Publicado: (2025)
por: Wang, Youjin, et al.
Publicado: (2025)
When Attention Sink Emerges in Language Models: An Empirical View
por: Gu, Xiangming, et al.
Publicado: (2024)
por: Gu, Xiangming, et al.
Publicado: (2024)
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
por: Zhang, Stephen, et al.
Publicado: (2025)
por: Zhang, Stephen, et al.
Publicado: (2025)
Hidden State Poisoning Attacks against Mamba-based Language Models
por: Mercier, Alexandre Le, et al.
Publicado: (2026)
por: Mercier, Alexandre Le, et al.
Publicado: (2026)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
por: Pióro, Maciej, et al.
Publicado: (2024)
por: Pióro, Maciej, et al.
Publicado: (2024)
Differential Mamba
por: Schneider, Nadav, et al.
Publicado: (2025)
por: Schneider, Nadav, et al.
Publicado: (2025)
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
por: Shin, Seungjun, et al.
Publicado: (2025)
por: Shin, Seungjun, et al.
Publicado: (2025)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
por: Chen, Yifang, et al.
Publicado: (2024)
por: Chen, Yifang, et al.
Publicado: (2024)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
por: Kossen, Jannik, et al.
Publicado: (2024)
por: Kossen, Jannik, et al.
Publicado: (2024)
Exploring the Limitations of Mamba in COPY and CoT Reasoning
por: Ren, Ruifeng, et al.
Publicado: (2024)
por: Ren, Ruifeng, et al.
Publicado: (2024)
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
por: NVIDIA, et al.
Publicado: (2025)
por: NVIDIA, et al.
Publicado: (2025)
Synthetic vs. Gold: The Role of LLM Generated Labels and Data in Cyberbullying Detection
por: Kazemi, Arefeh, et al.
Publicado: (2025)
por: Kazemi, Arefeh, et al.
Publicado: (2025)
Rhetorical Questions in LLM Representations: A Linear Probing Study
por: Yao, Louie Hong, et al.
Publicado: (2026)
por: Yao, Louie Hong, et al.
Publicado: (2026)
Max It or Miss It: Benchmarking LLM On Solving Extremal Problems
por: Gao, Binxin, et al.
Publicado: (2025)
por: Gao, Binxin, et al.
Publicado: (2025)
Post-Trained MoE Can Skip Half Experts via Self-Distillation
por: Lv, Xingtai, et al.
Publicado: (2026)
por: Lv, Xingtai, et al.
Publicado: (2026)
When Numbers Tell Half the Story: Human-Metric Alignment in Topic Model Evaluation
por: Prouteau, Thibault, et al.
Publicado: (2026)
por: Prouteau, Thibault, et al.
Publicado: (2026)
Towards Execution-Grounded Automated AI Research
por: Si, Chenglei, et al.
Publicado: (2026)
por: Si, Chenglei, et al.
Publicado: (2026)
From Emergence to Control: Probing and Modulating Self-Reflection in Language Models
por: Zhu, Xudong, et al.
Publicado: (2025)
por: Zhu, Xudong, et al.
Publicado: (2025)
BlackMamba: Mixture of Experts for State-Space Models
por: Anthony, Quentin, et al.
Publicado: (2024)
por: Anthony, Quentin, et al.
Publicado: (2024)
Single-pass Detection of Jailbreaking Input in Large Language Models
por: Candogan, Leyla Naz, et al.
Publicado: (2025)
por: Candogan, Leyla Naz, et al.
Publicado: (2025)
Accelerated AI Inference via Dynamic Execution Methods
por: Barad, Haim, et al.
Publicado: (2024)
por: Barad, Haim, et al.
Publicado: (2024)
Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm
por: Yang, Kaisen, et al.
Publicado: (2025)
por: Yang, Kaisen, et al.
Publicado: (2025)
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
por: NVIDIA, et al.
Publicado: (2025)
por: NVIDIA, et al.
Publicado: (2025)
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
por: Fan, Chenrui, et al.
Publicado: (2025)
por: Fan, Chenrui, et al.
Publicado: (2025)
R2Gen-Mamba: A Selective State Space Model for Radiology Report Generation
por: Sun, Yongheng, et al.
Publicado: (2024)
por: Sun, Yongheng, et al.
Publicado: (2024)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
por: Le, Qi, et al.
Publicado: (2025)
por: Le, Qi, et al.
Publicado: (2025)
Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
por: Luo, Zheng, et al.
Publicado: (2026)
por: Luo, Zheng, et al.
Publicado: (2026)
TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Composition
por: Kulkarni, Anay, et al.
Publicado: (2026)
por: Kulkarni, Anay, et al.
Publicado: (2026)
Flaming-hot Initiation with Regular Execution Sampling for Large Language Models
por: Chen, Weizhe, et al.
Publicado: (2024)
por: Chen, Weizhe, et al.
Publicado: (2024)
ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution
por: Coca, Alexandru, et al.
Publicado: (2025)
por: Coca, Alexandru, et al.
Publicado: (2025)
Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems
por: Khan, Zaid, et al.
Publicado: (2025)
por: Khan, Zaid, et al.
Publicado: (2025)
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
por: Kuratov, Yuri, et al.
Publicado: (2024)
por: Kuratov, Yuri, et al.
Publicado: (2024)
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
por: NVIDIA, et al.
Publicado: (2026)
por: NVIDIA, et al.
Publicado: (2026)
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
por: NVIDIA, et al.
Publicado: (2025)
por: NVIDIA, et al.
Publicado: (2025)
The Information Geometry of Softmax: Probing and Steering
por: Park, Kiho, et al.
Publicado: (2026)
por: Park, Kiho, et al.
Publicado: (2026)
Building Production-Ready Probes For Gemini
por: Kramár, János, et al.
Publicado: (2026)
por: Kramár, János, et al.
Publicado: (2026)
Ejemplares similares
-
On the Existence and Behavior of Secondary Attention Sinks
por: Wong, Jeffrey T. H., et al.
Publicado: (2025) -
Sink-Aware Pruning for Diffusion Language Models
por: Myrzakhan, Aidar, et al.
Publicado: (2026) -
Stuffed Mamba: Oversized States Lead to the Inability to Forget
por: Chen, Yingfa, et al.
Publicado: (2024) -
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
por: Xu, Zukang, et al.
Publicado: (2025) -
MemMamba: Rethinking Memory Patterns in State Space Model
por: Wang, Youjin, et al.
Publicado: (2025)