Attention to Mamba: A Recipe for Cross-Architecture Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Moudgil, Abhinav, Huang, Ningyuan, Dhekane, Eeshan Gunesh, Rodríguez, Pau, Zappella, Luca, Danieli, Federico |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
by: Huang, Ningyuan, et al.
Published: (2025)
by: Huang, Ningyuan, et al.
Published: (2025)
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
by: Danieli, Federico, et al.
Published: (2025)
by: Danieli, Federico, et al.
Published: (2025)
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
by: Nogales, Miguel, et al.
Published: (2025)
by: Nogales, Miguel, et al.
Published: (2025)
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
by: Liu, Linyu, et al.
Published: (2024)
by: Liu, Linyu, et al.
Published: (2024)
Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
by: de Seyssel, Maureen, et al.
Published: (2025)
by: de Seyssel, Maureen, et al.
Published: (2025)
Controlling Language and Diffusion Models by Transporting Activations
by: Rodriguez, Pau, et al.
Published: (2024)
by: Rodriguez, Pau, et al.
Published: (2024)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
by: Kirchhof, Michael, et al.
Published: (2025)
by: Kirchhof, Michael, et al.
Published: (2025)
A Language Model-Driven Semi-Supervised Ensemble Framework for Illicit Market Detection Across Deep/Dark Web and Social Platforms
by: Yazdanjue, Navid, et al.
Published: (2025)
by: Yazdanjue, Navid, et al.
Published: (2025)
Clustering in pure-attention hardmax transformers and its role in sentiment analysis
by: Alcalde, Albert, et al.
Published: (2024)
by: Alcalde, Albert, et al.
Published: (2024)
A Generalization Bound for a Family of Implicit Networks
by: Fung, Samy Wu, et al.
Published: (2024)
by: Fung, Samy Wu, et al.
Published: (2024)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
by: Wang, Youkang, et al.
Published: (2025)
by: Wang, Youkang, et al.
Published: (2025)
Enhanced QKNorm normalization for neural transformers with the Lp norm
by: Lopez-Rubio, Ezequiel, et al.
Published: (2026)
by: Lopez-Rubio, Ezequiel, et al.
Published: (2026)
Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics
by: Dahlem, Dominik, et al.
Published: (2026)
by: Dahlem, Dominik, et al.
Published: (2026)
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization
by: Pavlov, Gorgi
Published: (2026)
by: Pavlov, Gorgi
Published: (2026)
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
by: Brännvall, Rickard, et al.
Published: (2023)
by: Brännvall, Rickard, et al.
Published: (2023)
H-Model: Dynamic Neural Architectures for Adaptive Processing
by: Hospodarchuk, Dmytro
Published: (2025)
by: Hospodarchuk, Dmytro
Published: (2025)
Mamba for Scalable and Efficient Personalized Recommendations
by: Starnes, Andrew, et al.
Published: (2024)
by: Starnes, Andrew, et al.
Published: (2024)
BrainDistill: Implantable Motor Decoding with Task-Specific Knowledge Distillation
by: Xie, Yuhan, et al.
Published: (2026)
by: Xie, Yuhan, et al.
Published: (2026)
Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
by: Ilin, Ivan, et al.
Published: (2025)
by: Ilin, Ivan, et al.
Published: (2025)
Parameter-Efficient Transformer Embeddings
by: Ndubuaku, Henry, et al.
Published: (2025)
by: Ndubuaku, Henry, et al.
Published: (2025)
Deep Learning and Transfer Learning Architectures for English Premier League Player Performance Forecasting
by: Frees, Daniel, et al.
Published: (2024)
by: Frees, Daniel, et al.
Published: (2024)
Graph-Conditional Flow Matching for Relational Data Generation
by: Scassola, Davide, et al.
Published: (2025)
by: Scassola, Davide, et al.
Published: (2025)
Wave-Attractor-Tree: A Hierarchical Binary Tree Reduction Architecture for Efficient Sequence Modeling
by: Berezkin, Igor
Published: (2026)
by: Berezkin, Igor
Published: (2026)
Transformers are Efficient Compilers, Provably
by: Zhai, Xiyu, et al.
Published: (2024)
by: Zhai, Xiyu, et al.
Published: (2024)
Pay Attention to What You Need
by: Gao, Yifei, et al.
Published: (2023)
by: Gao, Yifei, et al.
Published: (2023)
Research on a hybrid LSTM-CNN-Attention model for text-based web content classification
by: Kuz, Mykola, et al.
Published: (2025)
by: Kuz, Mykola, et al.
Published: (2025)
HERCULES: Hardware-Efficient, Robust, Continual Learning Neural Architecture Search
by: Gambella, Matteo, et al.
Published: (2026)
by: Gambella, Matteo, et al.
Published: (2026)
ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
by: Chen, Yihong, et al.
Published: (2022)
by: Chen, Yihong, et al.
Published: (2022)
Diffeomorphic Measure Matching with Kernels for Generative Modeling
by: Pandey, Biraj, et al.
Published: (2024)
by: Pandey, Biraj, et al.
Published: (2024)
Optimizing Basis Function Selection in Constructive Wavelet Neural Networks and Its Applications
by: Huang, Dunsheng, et al.
Published: (2025)
by: Huang, Dunsheng, et al.
Published: (2025)
Scaling Properties of Continuous Diffusion Spoken Language Models
by: Ramapuram, Jason, et al.
Published: (2026)
by: Ramapuram, Jason, et al.
Published: (2026)
Subgroups of $U(d)$ Induce Natural RNN and Transformer Architectures
by: Nunley, Joshua
Published: (2026)
by: Nunley, Joshua
Published: (2026)
CNNtention: Can CNNs do better with Attention?
by: Kapila, Nikhil, et al.
Published: (2024)
by: Kapila, Nikhil, et al.
Published: (2024)
Temporal Attention Evolutional Graph Convolutional Network for Multivariate Time Series Forecasting
by: Zhao, Xinlong, et al.
Published: (2025)
by: Zhao, Xinlong, et al.
Published: (2025)
PolyTruth: Multilingual Disinformation Detection using Transformer-Based Language Models
by: Gouliev, Zaur, et al.
Published: (2025)
by: Gouliev, Zaur, et al.
Published: (2025)
Empirical analysis of binding precedent efficiency in Brazilian Supreme Court via case classification
by: Tinarrage, Raphaël, et al.
Published: (2024)
by: Tinarrage, Raphaël, et al.
Published: (2024)
Beyond Long Context: When Semantics Matter More than Tokens
by: Chawdhury, Tarun Kumar, et al.
Published: (2025)
by: Chawdhury, Tarun Kumar, et al.
Published: (2025)
Smoothed Embeddings for Robust Language Models
by: Hase, Ryo, et al.
Published: (2025)
by: Hase, Ryo, et al.
Published: (2025)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
by: Imanov, Olaf Yunus Laitinen
Published: (2026)
by: Imanov, Olaf Yunus Laitinen
Published: (2026)
Similar Items
-
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
by: Huang, Ningyuan, et al.
Published: (2025) -
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
by: Danieli, Federico, et al.
Published: (2025) -
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
by: Nogales, Miguel, et al.
Published: (2025) -
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
by: Liu, Linyu, et al.
Published: (2024) -
Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
by: de Seyssel, Maureen, et al.
Published: (2025)