Attention to Mamba: A Recipe for Cross-Architecture Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Moudgil, Abhinav, Huang, Ningyuan, Dhekane, Eeshan Gunesh, Rodríguez, Pau, Zappella, Luca, Danieli, Federico |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
von: Huang, Ningyuan, et al.
Veröffentlicht: (2025)
von: Huang, Ningyuan, et al.
Veröffentlicht: (2025)
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
von: Danieli, Federico, et al.
Veröffentlicht: (2025)
von: Danieli, Federico, et al.
Veröffentlicht: (2025)
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
von: Nogales, Miguel, et al.
Veröffentlicht: (2025)
von: Nogales, Miguel, et al.
Veröffentlicht: (2025)
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
von: Liu, Linyu, et al.
Veröffentlicht: (2024)
von: Liu, Linyu, et al.
Veröffentlicht: (2024)
Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2025)
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2025)
Controlling Language and Diffusion Models by Transporting Activations
von: Rodriguez, Pau, et al.
Veröffentlicht: (2024)
von: Rodriguez, Pau, et al.
Veröffentlicht: (2024)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
von: Kirchhof, Michael, et al.
Veröffentlicht: (2025)
von: Kirchhof, Michael, et al.
Veröffentlicht: (2025)
A Language Model-Driven Semi-Supervised Ensemble Framework for Illicit Market Detection Across Deep/Dark Web and Social Platforms
von: Yazdanjue, Navid, et al.
Veröffentlicht: (2025)
von: Yazdanjue, Navid, et al.
Veröffentlicht: (2025)
Clustering in pure-attention hardmax transformers and its role in sentiment analysis
von: Alcalde, Albert, et al.
Veröffentlicht: (2024)
von: Alcalde, Albert, et al.
Veröffentlicht: (2024)
A Generalization Bound for a Family of Implicit Networks
von: Fung, Samy Wu, et al.
Veröffentlicht: (2024)
von: Fung, Samy Wu, et al.
Veröffentlicht: (2024)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
von: Wang, Youkang, et al.
Veröffentlicht: (2025)
von: Wang, Youkang, et al.
Veröffentlicht: (2025)
Enhanced QKNorm normalization for neural transformers with the Lp norm
von: Lopez-Rubio, Ezequiel, et al.
Veröffentlicht: (2026)
von: Lopez-Rubio, Ezequiel, et al.
Veröffentlicht: (2026)
Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics
von: Dahlem, Dominik, et al.
Veröffentlicht: (2026)
von: Dahlem, Dominik, et al.
Veröffentlicht: (2026)
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization
von: Pavlov, Gorgi
Veröffentlicht: (2026)
von: Pavlov, Gorgi
Veröffentlicht: (2026)
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
von: Brännvall, Rickard, et al.
Veröffentlicht: (2023)
von: Brännvall, Rickard, et al.
Veröffentlicht: (2023)
H-Model: Dynamic Neural Architectures for Adaptive Processing
von: Hospodarchuk, Dmytro
Veröffentlicht: (2025)
von: Hospodarchuk, Dmytro
Veröffentlicht: (2025)
Mamba for Scalable and Efficient Personalized Recommendations
von: Starnes, Andrew, et al.
Veröffentlicht: (2024)
von: Starnes, Andrew, et al.
Veröffentlicht: (2024)
BrainDistill: Implantable Motor Decoding with Task-Specific Knowledge Distillation
von: Xie, Yuhan, et al.
Veröffentlicht: (2026)
von: Xie, Yuhan, et al.
Veröffentlicht: (2026)
Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
von: Ilin, Ivan, et al.
Veröffentlicht: (2025)
von: Ilin, Ivan, et al.
Veröffentlicht: (2025)
Parameter-Efficient Transformer Embeddings
von: Ndubuaku, Henry, et al.
Veröffentlicht: (2025)
von: Ndubuaku, Henry, et al.
Veröffentlicht: (2025)
Deep Learning and Transfer Learning Architectures for English Premier League Player Performance Forecasting
von: Frees, Daniel, et al.
Veröffentlicht: (2024)
von: Frees, Daniel, et al.
Veröffentlicht: (2024)
Graph-Conditional Flow Matching for Relational Data Generation
von: Scassola, Davide, et al.
Veröffentlicht: (2025)
von: Scassola, Davide, et al.
Veröffentlicht: (2025)
Wave-Attractor-Tree: A Hierarchical Binary Tree Reduction Architecture for Efficient Sequence Modeling
von: Berezkin, Igor
Veröffentlicht: (2026)
von: Berezkin, Igor
Veröffentlicht: (2026)
Transformers are Efficient Compilers, Provably
von: Zhai, Xiyu, et al.
Veröffentlicht: (2024)
von: Zhai, Xiyu, et al.
Veröffentlicht: (2024)
Pay Attention to What You Need
von: Gao, Yifei, et al.
Veröffentlicht: (2023)
von: Gao, Yifei, et al.
Veröffentlicht: (2023)
Research on a hybrid LSTM-CNN-Attention model for text-based web content classification
von: Kuz, Mykola, et al.
Veröffentlicht: (2025)
von: Kuz, Mykola, et al.
Veröffentlicht: (2025)
HERCULES: Hardware-Efficient, Robust, Continual Learning Neural Architecture Search
von: Gambella, Matteo, et al.
Veröffentlicht: (2026)
von: Gambella, Matteo, et al.
Veröffentlicht: (2026)
ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
von: Chen, Yihong, et al.
Veröffentlicht: (2022)
von: Chen, Yihong, et al.
Veröffentlicht: (2022)
Diffeomorphic Measure Matching with Kernels for Generative Modeling
von: Pandey, Biraj, et al.
Veröffentlicht: (2024)
von: Pandey, Biraj, et al.
Veröffentlicht: (2024)
Optimizing Basis Function Selection in Constructive Wavelet Neural Networks and Its Applications
von: Huang, Dunsheng, et al.
Veröffentlicht: (2025)
von: Huang, Dunsheng, et al.
Veröffentlicht: (2025)
Scaling Properties of Continuous Diffusion Spoken Language Models
von: Ramapuram, Jason, et al.
Veröffentlicht: (2026)
von: Ramapuram, Jason, et al.
Veröffentlicht: (2026)
Subgroups of $U(d)$ Induce Natural RNN and Transformer Architectures
von: Nunley, Joshua
Veröffentlicht: (2026)
von: Nunley, Joshua
Veröffentlicht: (2026)
CNNtention: Can CNNs do better with Attention?
von: Kapila, Nikhil, et al.
Veröffentlicht: (2024)
von: Kapila, Nikhil, et al.
Veröffentlicht: (2024)
Temporal Attention Evolutional Graph Convolutional Network for Multivariate Time Series Forecasting
von: Zhao, Xinlong, et al.
Veröffentlicht: (2025)
von: Zhao, Xinlong, et al.
Veröffentlicht: (2025)
PolyTruth: Multilingual Disinformation Detection using Transformer-Based Language Models
von: Gouliev, Zaur, et al.
Veröffentlicht: (2025)
von: Gouliev, Zaur, et al.
Veröffentlicht: (2025)
Empirical analysis of binding precedent efficiency in Brazilian Supreme Court via case classification
von: Tinarrage, Raphaël, et al.
Veröffentlicht: (2024)
von: Tinarrage, Raphaël, et al.
Veröffentlicht: (2024)
Beyond Long Context: When Semantics Matter More than Tokens
von: Chawdhury, Tarun Kumar, et al.
Veröffentlicht: (2025)
von: Chawdhury, Tarun Kumar, et al.
Veröffentlicht: (2025)
Smoothed Embeddings for Robust Language Models
von: Hase, Ryo, et al.
Veröffentlicht: (2025)
von: Hase, Ryo, et al.
Veröffentlicht: (2025)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
von: Imanov, Olaf Yunus Laitinen
Veröffentlicht: (2026)
von: Imanov, Olaf Yunus Laitinen
Veröffentlicht: (2026)
Ähnliche Einträge
-
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
von: Huang, Ningyuan, et al.
Veröffentlicht: (2025) -
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
von: Danieli, Federico, et al.
Veröffentlicht: (2025) -
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
von: Nogales, Miguel, et al.
Veröffentlicht: (2025) -
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
von: Liu, Linyu, et al.
Veröffentlicht: (2024) -
Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2025)