Gated Delta Networks: Improving Mamba2 with Delta Rule
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Songlin, Kautz, Jan, Hatamizadeh, Ali |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
por: Yang, Songlin, et al.
Publicado: (2024)
por: Yang, Songlin, et al.
Publicado: (2024)
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
por: Hatamizadeh, Ali, et al.
Publicado: (2026)
por: Hatamizadeh, Ali, et al.
Publicado: (2026)
OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention
por: Zhou, Chenyu, et al.
Publicado: (2026)
por: Zhou, Chenyu, et al.
Publicado: (2026)
MambaVision: A Hybrid Mamba-Transformer Vision Backbone
por: Hatamizadeh, Ali, et al.
Publicado: (2024)
por: Hatamizadeh, Ali, et al.
Publicado: (2024)
An Empirical Study of Mamba-based Language Models
por: Waleffe, Roger, et al.
Publicado: (2024)
por: Waleffe, Roger, et al.
Publicado: (2024)
RLP: Reinforcement as a Pretraining Objective
por: Hatamizadeh, Ali, et al.
Publicado: (2025)
por: Hatamizadeh, Ali, et al.
Publicado: (2025)
Gated Linear Attention Transformers with Hardware-Efficient Training
por: Yang, Songlin, et al.
Publicado: (2023)
por: Yang, Songlin, et al.
Publicado: (2023)
DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products
por: Siems, Julien, et al.
Publicado: (2025)
por: Siems, Julien, et al.
Publicado: (2025)
DiffiT: Diffusion Vision Transformers for Image Generation
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
Deep Delta Learning
por: Zhang, Yifan, et al.
Publicado: (2026)
por: Zhang, Yifan, et al.
Publicado: (2026)
Delta Knowledge Distillation for Large Language Models
por: Cao, Yihan, et al.
Publicado: (2025)
por: Cao, Yihan, et al.
Publicado: (2025)
BitDelta: Your Fine-Tune May Only Be Worth One Bit
por: Liu, James, et al.
Publicado: (2024)
por: Liu, James, et al.
Publicado: (2024)
A Unified View of Delta Parameter Editing in Post-Trained Large-Scale Models
por: Tang, Qiaoyu, et al.
Publicado: (2024)
por: Tang, Qiaoyu, et al.
Publicado: (2024)
DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models
por: Deng, Wenlong, et al.
Publicado: (2024)
por: Deng, Wenlong, et al.
Publicado: (2024)
Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
por: Xiong, Boya, et al.
Publicado: (2025)
por: Xiong, Boya, et al.
Publicado: (2025)
ViR: Towards Efficient Vision Retention Backbones
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
Delta Activations: A Representation for Finetuned Large Language Models
por: Xu, Zhiqiu, et al.
Publicado: (2025)
por: Xu, Zhiqiu, et al.
Publicado: (2025)
RuleR: Improving LLM Controllability by Rule-based Data Recycling
por: Li, Ming, et al.
Publicado: (2024)
por: Li, Ming, et al.
Publicado: (2024)
EfficientXpert: Efficient Domain Adaptation for Large Language Models via Propagation-Aware Pruning
por: Zhao, Songlin, et al.
Publicado: (2025)
por: Zhao, Songlin, et al.
Publicado: (2025)
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
por: Wang, Tevin, et al.
Publicado: (2025)
por: Wang, Tevin, et al.
Publicado: (2025)
FasterViT: Fast Vision Transformers with Hierarchical Attention
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
por: Xu, Zukang, et al.
Publicado: (2025)
por: Xu, Zukang, et al.
Publicado: (2025)
Flextron: Many-in-One Flexible Large Language Model
por: Cai, Ruisi, et al.
Publicado: (2024)
por: Cai, Ruisi, et al.
Publicado: (2024)
Representation Learning with Conditional Information Flow Maximization
por: Hu, Dou, et al.
Publicado: (2024)
por: Hu, Dou, et al.
Publicado: (2024)
Mamba Knockout for Unraveling Factual Information Flow
por: Endy, Nir, et al.
Publicado: (2025)
por: Endy, Nir, et al.
Publicado: (2025)
Differential Mamba
por: Schneider, Nadav, et al.
Publicado: (2025)
por: Schneider, Nadav, et al.
Publicado: (2025)
Jamba: A Hybrid Transformer-Mamba Language Model
por: Lieber, Opher, et al.
Publicado: (2024)
por: Lieber, Opher, et al.
Publicado: (2024)
Lost in State Space: Probing Frozen Mamba Representations
por: Wagh, Bhagyashree, et al.
Publicado: (2026)
por: Wagh, Bhagyashree, et al.
Publicado: (2026)
Masked Gated Linear Unit
por: Tajima, Yukito, et al.
Publicado: (2025)
por: Tajima, Yukito, et al.
Publicado: (2025)
RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling
por: Liu, Yang, et al.
Publicado: (2025)
por: Liu, Yang, et al.
Publicado: (2025)
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
por: Yang, Xikang, et al.
Publicado: (2024)
por: Yang, Xikang, et al.
Publicado: (2024)
DenseMamba: State Space Models with Dense Hidden Connection for Efficient Large Language Models
por: He, Wei, et al.
Publicado: (2024)
por: He, Wei, et al.
Publicado: (2024)
Rule2Text: Natural Language Explanation of Logical Rules in Knowledge Graphs
por: Shirvani-Mahdavi, Nasim, et al.
Publicado: (2025)
por: Shirvani-Mahdavi, Nasim, et al.
Publicado: (2025)
R2T: Rule-Encoded Loss Functions for Low-Resource Sequence Tagging
por: Keita, Mamadou K., et al.
Publicado: (2025)
por: Keita, Mamadou K., et al.
Publicado: (2025)
MambaByte: Token-free Selective State Space Model
por: Wang, Junxiong, et al.
Publicado: (2024)
por: Wang, Junxiong, et al.
Publicado: (2024)
Jamba-1.5: Hybrid Transformer-Mamba Models at Scale
por: Jamba Team, et al.
Publicado: (2024)
por: Jamba Team, et al.
Publicado: (2024)
Structured Probabilistic Coding
por: Hu, Dou, et al.
Publicado: (2023)
por: Hu, Dou, et al.
Publicado: (2023)
BroRL: Scaling Reinforcement Learning via Broadened Exploration
por: Hu, Jian, et al.
Publicado: (2025)
por: Hu, Jian, et al.
Publicado: (2025)
CoGate-LSTM: Prototype-Guided Feature-Space Gating for Mitigating Gradient Dilution in Imbalanced Toxic Comment Classification
por: Mohammad, Noor Islam S.
Publicado: (2025)
por: Mohammad, Noor Islam S.
Publicado: (2025)
Leveraging Logical Rules in Knowledge Editing: A Cherry on the Top
por: Cheng, Keyuan, et al.
Publicado: (2024)
por: Cheng, Keyuan, et al.
Publicado: (2024)
Ejemplares similares
-
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
por: Yang, Songlin, et al.
Publicado: (2024) -
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
por: Hatamizadeh, Ali, et al.
Publicado: (2026) -
OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention
por: Zhou, Chenyu, et al.
Publicado: (2026) -
MambaVision: A Hybrid Mamba-Transformer Vision Backbone
por: Hatamizadeh, Ali, et al.
Publicado: (2024) -
An Empirical Study of Mamba-based Language Models
por: Waleffe, Roger, et al.
Publicado: (2024)