Retrieval-Aware Distillation for Transformer-SSM Hybrids
Fuente:
arXiv
Saved in:
| Main Authors: | Bick, Aviv, Xing, Eric P., Gu, Albert |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
by: Bick, Aviv, et al.
Published: (2025)
by: Bick, Aviv, et al.
Published: (2025)
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
by: Bick, Aviv, et al.
Published: (2024)
by: Bick, Aviv, et al.
Published: (2024)
Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing
by: Bick, Aviv, et al.
Published: (2025)
by: Bick, Aviv, et al.
Published: (2025)
Zamba: A Compact 7B SSM Hybrid Model
by: Glorioso, Paolo, et al.
Published: (2024)
by: Glorioso, Paolo, et al.
Published: (2024)
A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models
by: Kong, Jason, et al.
Published: (2026)
by: Kong, Jason, et al.
Published: (2026)
ECGMamba: Towards Efficient ECG Classification with BiSSM
by: Qiang, Yupeng, et al.
Published: (2024)
by: Qiang, Yupeng, et al.
Published: (2024)
Heracles: A Hybrid SSM-Transformer Model for High-Resolution Image and Time-Series Analysis
by: Patro, Badri N., et al.
Published: (2024)
by: Patro, Badri N., et al.
Published: (2024)
Hybrid Data-Driven SSM for Interpretable and Label-Free mmWave Channel Prediction
by: Sun, Yiyong, et al.
Published: (2024)
by: Sun, Yiyong, et al.
Published: (2024)
A Comparison of Methods for Neural Network Aggregation
by: Pomerat, John, et al.
Published: (2023)
by: Pomerat, John, et al.
Published: (2023)
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
by: Terzić, Aleksandar, et al.
Published: (2026)
by: Terzić, Aleksandar, et al.
Published: (2026)
Time-SSM: Simplifying and Unifying State Space Models for Time Series Forecasting
by: Hu, Jiaxi, et al.
Published: (2024)
by: Hu, Jiaxi, et al.
Published: (2024)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
by: Ding, Ken
Published: (2026)
by: Ding, Ken
Published: (2026)
Decision MetaMamba: Enhancing Selective SSM in Offline RL with Heterogeneous Sequence Mixing
by: Kim, Wall, et al.
Published: (2026)
by: Kim, Wall, et al.
Published: (2026)
Decision MetaMamba: Enhancing Selective SSM in Offline RL with Heterogeneous Sequence Mixing
by: Kim, Wall, et al.
Published: (2024)
by: Kim, Wall, et al.
Published: (2024)
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
by: Segal-Feldman, Yael, et al.
Published: (2024)
by: Segal-Feldman, Yael, et al.
Published: (2024)
The Mamba in the Llama: Distilling and Accelerating Hybrid Models
by: Wang, Junxiong, et al.
Published: (2024)
by: Wang, Junxiong, et al.
Published: (2024)
How Many Heads Make an SSM? A Unified Framework for Attention and State Space Models
by: Ghodsi, Ali
Published: (2025)
by: Ghodsi, Ali
Published: (2025)
GradMetaNet: An Equivariant Architecture for Learning on Gradients
by: Gelberg, Yoav, et al.
Published: (2025)
by: Gelberg, Yoav, et al.
Published: (2025)
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
by: Gu, Albert, et al.
Published: (2023)
by: Gu, Albert, et al.
Published: (2023)
Meta Reinforcement Learning with Finite Training Tasks -- a Density Estimation Approach
by: Rimon, Zohar, et al.
Published: (2022)
by: Rimon, Zohar, et al.
Published: (2022)
MoGU: Mixture-of-Gaussians with Uncertainty-based Gating for Time Series Forecasting
by: Aviv, Gilad, et al.
Published: (2025)
by: Aviv, Gilad, et al.
Published: (2025)
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
Improved Generalization of Weight Space Networks via Augmentations
by: Shamsian, Aviv, et al.
Published: (2024)
by: Shamsian, Aviv, et al.
Published: (2024)
Generalized Policy Gradient with History-Aware Decision Transformer for Reliable Routing over Graph Signals
by: Wei, Xing, et al.
Published: (2025)
by: Wei, Xing, et al.
Published: (2025)
CADENT: Gated Hybrid Distillation for Sample-Efficient Transfer in Reinforcement Learning
by: Alinejad, Mahyar, et al.
Published: (2026)
by: Alinejad, Mahyar, et al.
Published: (2026)
Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction
by: Xia, Xiaojie, et al.
Published: (2026)
by: Xia, Xiaojie, et al.
Published: (2026)
Dataset Distillation-based Hybrid Federated Learning on Non-IID Data
by: Shi, Xiufang, et al.
Published: (2024)
by: Shi, Xiufang, et al.
Published: (2024)
HybriDNA: A Hybrid Transformer-Mamba2 Long-Range DNA Language Model
by: Ma, Mingqian, et al.
Published: (2025)
by: Ma, Mingqian, et al.
Published: (2025)
EEG-SSM: Leveraging State-Space Model for Dementia Detection
by: Tran, Xuan-The, et al.
Published: (2024)
by: Tran, Xuan-The, et al.
Published: (2024)
CoRPO: Adding a Correctness Bias to GRPO Improves Generalization
by: Garg, Anisha, et al.
Published: (2025)
by: Garg, Anisha, et al.
Published: (2025)
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation
by: Liu, Junzhang, et al.
Published: (2024)
by: Liu, Junzhang, et al.
Published: (2024)
Diversity-Aware Reverse Kullback-Leibler Divergence for Large Language Model Distillation
by: Luong, Hoang-Chau, et al.
Published: (2026)
by: Luong, Hoang-Chau, et al.
Published: (2026)
Distilling and Retrieving Generalizable Knowledge for Robot Manipulation via Language Corrections
by: Zha, Lihan, et al.
Published: (2023)
by: Zha, Lihan, et al.
Published: (2023)
Spatially-Aware Transformer for Embodied Agents
by: Cho, Junmo, et al.
Published: (2024)
by: Cho, Junmo, et al.
Published: (2024)
Aligning Dense Retrievers with LLM Utility via Distillation
by: Sandhu, Rajinder, et al.
Published: (2026)
by: Sandhu, Rajinder, et al.
Published: (2026)
Enhancing Transformer with GNN Structural Knowledge via Distillation: A Novel Approach
by: Duan, Zhihua, et al.
Published: (2025)
by: Duan, Zhihua, et al.
Published: (2025)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
by: Shen, Yuhao, et al.
Published: (2026)
by: Shen, Yuhao, et al.
Published: (2026)
$\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control
by: Chen, Xianwei, et al.
Published: (2026)
by: Chen, Xianwei, et al.
Published: (2026)
Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO
by: Yu, Bowen, et al.
Published: (2026)
by: Yu, Bowen, et al.
Published: (2026)
Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
by: Duan, Xintong, et al.
Published: (2025)
by: Duan, Xintong, et al.
Published: (2025)
Similar Items
-
Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
by: Bick, Aviv, et al.
Published: (2025) -
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
by: Bick, Aviv, et al.
Published: (2024) -
Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing
by: Bick, Aviv, et al.
Published: (2025) -
Zamba: A Compact 7B SSM Hybrid Model
by: Glorioso, Paolo, et al.
Published: (2024) -
A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models
by: Kong, Jason, et al.
Published: (2026)