The Mamba in the Llama: Distilling and Accelerating Hybrid Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Junxiong, Paliotta, Daniele, May, Avner, Rush, Alexander M., Dao, Tri |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
von: Wang, Junxiong, et al.
Veröffentlicht: (2025)
von: Wang, Junxiong, et al.
Veröffentlicht: (2025)
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
von: Gu, Albert, et al.
Veröffentlicht: (2023)
von: Gu, Albert, et al.
Veröffentlicht: (2023)
Speculative Speculative Decoding
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
von: Paliotta, Daniele, et al.
Veröffentlicht: (2025)
von: Paliotta, Daniele, et al.
Veröffentlicht: (2025)
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
von: Guo, Wentao, et al.
Veröffentlicht: (2025)
von: Guo, Wentao, et al.
Veröffentlicht: (2025)
HybriDNA: A Hybrid Transformer-Mamba2 Long-Range DNA Language Model
von: Ma, Mingqian, et al.
Veröffentlicht: (2025)
von: Ma, Mingqian, et al.
Veröffentlicht: (2025)
InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model
von: Wang, Youjin, et al.
Veröffentlicht: (2026)
von: Wang, Youjin, et al.
Veröffentlicht: (2026)
Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers
von: Hwang, Sukjun, et al.
Veröffentlicht: (2024)
von: Hwang, Sukjun, et al.
Veröffentlicht: (2024)
eMamba: Efficient Acceleration Framework for Mamba Models in Edge Computing
von: Kim, Jiyong, et al.
Veröffentlicht: (2025)
von: Kim, Jiyong, et al.
Veröffentlicht: (2025)
MambaByte: Token-free Selective State Space Model
von: Wang, Junxiong, et al.
Veröffentlicht: (2024)
von: Wang, Junxiong, et al.
Veröffentlicht: (2024)
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
von: Mishra, Mayank, et al.
Veröffentlicht: (2026)
von: Mishra, Mayank, et al.
Veröffentlicht: (2026)
Scaling Algorithm Distillation for Continuous Control with Mamba
von: Beaussant, Samuel, et al.
Veröffentlicht: (2025)
von: Beaussant, Samuel, et al.
Veröffentlicht: (2025)
Lugha-Llama: Adapting Large Language Models for African Languages
von: Buzaaba, Happy, et al.
Veröffentlicht: (2025)
von: Buzaaba, Happy, et al.
Veröffentlicht: (2025)
Llama-Nemotron: Efficient Reasoning Models
von: Bercovich, Akhiad, et al.
Veröffentlicht: (2025)
von: Bercovich, Akhiad, et al.
Veröffentlicht: (2025)
Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting
von: Rasul, Kashif, et al.
Veröffentlicht: (2023)
von: Rasul, Kashif, et al.
Veröffentlicht: (2023)
Marconi: Prefix Caching for the Era of Hybrid LLMs
von: Pan, Rui, et al.
Veröffentlicht: (2024)
von: Pan, Rui, et al.
Veröffentlicht: (2024)
Compute-Constrained Data Selection
von: Yin, Junjie Oscar, et al.
Veröffentlicht: (2024)
von: Yin, Junjie Oscar, et al.
Veröffentlicht: (2024)
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
von: Shah, Jay, et al.
Veröffentlicht: (2024)
von: Shah, Jay, et al.
Veröffentlicht: (2024)
Open Llama2 Model for the Lithuanian Language
von: Nakvosas, Artūras, et al.
Veröffentlicht: (2024)
von: Nakvosas, Artūras, et al.
Veröffentlicht: (2024)
NASH: Neural Architecture and Accelerator Search for Multiplication-Reduced Hybrid Models
von: Xu, Yang, et al.
Veröffentlicht: (2024)
von: Xu, Yang, et al.
Veröffentlicht: (2024)
Mamba Modulation: On the Length Generalization of Mamba
von: Lu, Peng, et al.
Veröffentlicht: (2025)
von: Lu, Peng, et al.
Veröffentlicht: (2025)
Llama-Affinity: A Predictive Antibody Antigen Binding Model Integrating Antibody Sequences with Llama3 Backbone Architecture
von: Hossain, Delower, et al.
Veröffentlicht: (2025)
von: Hossain, Delower, et al.
Veröffentlicht: (2025)
Steering Llama 2 via Contrastive Activation Addition
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
von: Ding, Ken
Veröffentlicht: (2026)
von: Ding, Ken
Veröffentlicht: (2026)
Global Convergence of Multiplicative Updates for the Matrix Mechanism: A Collaborative Proof with Gemini 3
von: Rush, Keith
Veröffentlicht: (2026)
von: Rush, Keith
Veröffentlicht: (2026)
Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction
von: Xia, Xiaojie, et al.
Veröffentlicht: (2026)
von: Xia, Xiaojie, et al.
Veröffentlicht: (2026)
From Efficient Multimodal Models to World Models: A Survey
von: Mai, Xinji, et al.
Veröffentlicht: (2024)
von: Mai, Xinji, et al.
Veröffentlicht: (2024)
CADENT: Gated Hybrid Distillation for Sample-Efficient Transfer in Reinforcement Learning
von: Alinejad, Mahyar, et al.
Veröffentlicht: (2026)
von: Alinejad, Mahyar, et al.
Veröffentlicht: (2026)
SST: Multi-Scale Hybrid Mamba-Transformer Experts for Time Series Forecasting
von: Xu, Xiongxiao, et al.
Veröffentlicht: (2024)
von: Xu, Xiongxiao, et al.
Veröffentlicht: (2024)
Retrieval-Aware Distillation for Transformer-SSM Hybrids
von: Bick, Aviv, et al.
Veröffentlicht: (2026)
von: Bick, Aviv, et al.
Veröffentlicht: (2026)
Challenges in Mechanistically Interpreting Model Representations
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
MELCOT: A Hybrid Learning Architecture with Marginal Preservation for Matrix-Valued Regression
von: Tran, Khang, et al.
Veröffentlicht: (2025)
von: Tran, Khang, et al.
Veröffentlicht: (2025)
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
SQT -- std $Q$-target
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
ODIA: Oriented Distillation for Inline Acceleration of LLM-based Function Calling
von: Zhang, Hanlong, et al.
Veröffentlicht: (2025)
von: Zhang, Hanlong, et al.
Veröffentlicht: (2025)
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
von: Zhan, Zheng, et al.
Veröffentlicht: (2025)
von: Zhan, Zheng, et al.
Veröffentlicht: (2025)
DeciMamba: Exploring the Length Extrapolation Potential of Mamba
von: Ben-Kish, Assaf, et al.
Veröffentlicht: (2024)
von: Ben-Kish, Assaf, et al.
Veröffentlicht: (2024)
OSDTW: Optimal Shared Depth and Task Weighting for Long-Tailed Recognition
von: Chu, Chang, et al.
Veröffentlicht: (2026)
von: Chu, Chang, et al.
Veröffentlicht: (2026)
Hybrid Mamba-Transformer Decoder for Error-Correcting Codes
von: Cohen, Shy-el, et al.
Veröffentlicht: (2025)
von: Cohen, Shy-el, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
von: Wang, Junxiong, et al.
Veröffentlicht: (2025) -
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
von: Gu, Albert, et al.
Veröffentlicht: (2023) -
Speculative Speculative Decoding
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026) -
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
von: Paliotta, Daniele, et al.
Veröffentlicht: (2025) -
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
von: Guo, Wentao, et al.
Veröffentlicht: (2025)