Attention to Mamba: A Recipe for Cross-Architecture Distillation
Fuente:
arXiv
Guardado en:
| Autores principales: | Moudgil, Abhinav, Huang, Ningyuan, Dhekane, Eeshan Gunesh, Rodríguez, Pau, Zappella, Luca, Danieli, Federico |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
por: Huang, Ningyuan, et al.
Publicado: (2025)
por: Huang, Ningyuan, et al.
Publicado: (2025)
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
por: Danieli, Federico, et al.
Publicado: (2025)
por: Danieli, Federico, et al.
Publicado: (2025)
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
por: Nogales, Miguel, et al.
Publicado: (2025)
por: Nogales, Miguel, et al.
Publicado: (2025)
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
por: Liu, Linyu, et al.
Publicado: (2024)
por: Liu, Linyu, et al.
Publicado: (2024)
Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
por: de Seyssel, Maureen, et al.
Publicado: (2025)
por: de Seyssel, Maureen, et al.
Publicado: (2025)
Controlling Language and Diffusion Models by Transporting Activations
por: Rodriguez, Pau, et al.
Publicado: (2024)
por: Rodriguez, Pau, et al.
Publicado: (2024)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
por: Kirchhof, Michael, et al.
Publicado: (2025)
por: Kirchhof, Michael, et al.
Publicado: (2025)
A Language Model-Driven Semi-Supervised Ensemble Framework for Illicit Market Detection Across Deep/Dark Web and Social Platforms
por: Yazdanjue, Navid, et al.
Publicado: (2025)
por: Yazdanjue, Navid, et al.
Publicado: (2025)
Clustering in pure-attention hardmax transformers and its role in sentiment analysis
por: Alcalde, Albert, et al.
Publicado: (2024)
por: Alcalde, Albert, et al.
Publicado: (2024)
A Generalization Bound for a Family of Implicit Networks
por: Fung, Samy Wu, et al.
Publicado: (2024)
por: Fung, Samy Wu, et al.
Publicado: (2024)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
por: Wang, Youkang, et al.
Publicado: (2025)
por: Wang, Youkang, et al.
Publicado: (2025)
Enhanced QKNorm normalization for neural transformers with the Lp norm
por: Lopez-Rubio, Ezequiel, et al.
Publicado: (2026)
por: Lopez-Rubio, Ezequiel, et al.
Publicado: (2026)
Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics
por: Dahlem, Dominik, et al.
Publicado: (2026)
por: Dahlem, Dominik, et al.
Publicado: (2026)
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
por: Fu, Tianyu, et al.
Publicado: (2025)
por: Fu, Tianyu, et al.
Publicado: (2025)
Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization
por: Pavlov, Gorgi
Publicado: (2026)
por: Pavlov, Gorgi
Publicado: (2026)
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
por: Brännvall, Rickard, et al.
Publicado: (2023)
por: Brännvall, Rickard, et al.
Publicado: (2023)
H-Model: Dynamic Neural Architectures for Adaptive Processing
por: Hospodarchuk, Dmytro
Publicado: (2025)
por: Hospodarchuk, Dmytro
Publicado: (2025)
Mamba for Scalable and Efficient Personalized Recommendations
por: Starnes, Andrew, et al.
Publicado: (2024)
por: Starnes, Andrew, et al.
Publicado: (2024)
BrainDistill: Implantable Motor Decoding with Task-Specific Knowledge Distillation
por: Xie, Yuhan, et al.
Publicado: (2026)
por: Xie, Yuhan, et al.
Publicado: (2026)
Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
por: Ilin, Ivan, et al.
Publicado: (2025)
por: Ilin, Ivan, et al.
Publicado: (2025)
Parameter-Efficient Transformer Embeddings
por: Ndubuaku, Henry, et al.
Publicado: (2025)
por: Ndubuaku, Henry, et al.
Publicado: (2025)
Deep Learning and Transfer Learning Architectures for English Premier League Player Performance Forecasting
por: Frees, Daniel, et al.
Publicado: (2024)
por: Frees, Daniel, et al.
Publicado: (2024)
Graph-Conditional Flow Matching for Relational Data Generation
por: Scassola, Davide, et al.
Publicado: (2025)
por: Scassola, Davide, et al.
Publicado: (2025)
Wave-Attractor-Tree: A Hierarchical Binary Tree Reduction Architecture for Efficient Sequence Modeling
por: Berezkin, Igor
Publicado: (2026)
por: Berezkin, Igor
Publicado: (2026)
Transformers are Efficient Compilers, Provably
por: Zhai, Xiyu, et al.
Publicado: (2024)
por: Zhai, Xiyu, et al.
Publicado: (2024)
Pay Attention to What You Need
por: Gao, Yifei, et al.
Publicado: (2023)
por: Gao, Yifei, et al.
Publicado: (2023)
Research on a hybrid LSTM-CNN-Attention model for text-based web content classification
por: Kuz, Mykola, et al.
Publicado: (2025)
por: Kuz, Mykola, et al.
Publicado: (2025)
HERCULES: Hardware-Efficient, Robust, Continual Learning Neural Architecture Search
por: Gambella, Matteo, et al.
Publicado: (2026)
por: Gambella, Matteo, et al.
Publicado: (2026)
ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
por: Chen, Yihong, et al.
Publicado: (2022)
por: Chen, Yihong, et al.
Publicado: (2022)
Diffeomorphic Measure Matching with Kernels for Generative Modeling
por: Pandey, Biraj, et al.
Publicado: (2024)
por: Pandey, Biraj, et al.
Publicado: (2024)
Optimizing Basis Function Selection in Constructive Wavelet Neural Networks and Its Applications
por: Huang, Dunsheng, et al.
Publicado: (2025)
por: Huang, Dunsheng, et al.
Publicado: (2025)
Scaling Properties of Continuous Diffusion Spoken Language Models
por: Ramapuram, Jason, et al.
Publicado: (2026)
por: Ramapuram, Jason, et al.
Publicado: (2026)
Subgroups of $U(d)$ Induce Natural RNN and Transformer Architectures
por: Nunley, Joshua
Publicado: (2026)
por: Nunley, Joshua
Publicado: (2026)
CNNtention: Can CNNs do better with Attention?
por: Kapila, Nikhil, et al.
Publicado: (2024)
por: Kapila, Nikhil, et al.
Publicado: (2024)
Temporal Attention Evolutional Graph Convolutional Network for Multivariate Time Series Forecasting
por: Zhao, Xinlong, et al.
Publicado: (2025)
por: Zhao, Xinlong, et al.
Publicado: (2025)
PolyTruth: Multilingual Disinformation Detection using Transformer-Based Language Models
por: Gouliev, Zaur, et al.
Publicado: (2025)
por: Gouliev, Zaur, et al.
Publicado: (2025)
Empirical analysis of binding precedent efficiency in Brazilian Supreme Court via case classification
por: Tinarrage, Raphaël, et al.
Publicado: (2024)
por: Tinarrage, Raphaël, et al.
Publicado: (2024)
Beyond Long Context: When Semantics Matter More than Tokens
por: Chawdhury, Tarun Kumar, et al.
Publicado: (2025)
por: Chawdhury, Tarun Kumar, et al.
Publicado: (2025)
Smoothed Embeddings for Robust Language Models
por: Hase, Ryo, et al.
Publicado: (2025)
por: Hase, Ryo, et al.
Publicado: (2025)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
por: Imanov, Olaf Yunus Laitinen
Publicado: (2026)
por: Imanov, Olaf Yunus Laitinen
Publicado: (2026)
Ejemplares similares
-
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
por: Huang, Ningyuan, et al.
Publicado: (2025) -
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
por: Danieli, Federico, et al.
Publicado: (2025) -
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
por: Nogales, Miguel, et al.
Publicado: (2025) -
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
por: Liu, Linyu, et al.
Publicado: (2024) -
Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
por: de Seyssel, Maureen, et al.
Publicado: (2025)