Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
Fuente:
arXiv
Salvato in:
| Autori principali: | Benfeghoul, Martin, Delgado, Teresa, Oomerjee, Adnan, Ammar, Haitham Bou, Wang, Jun, Fountas, Zafeirios |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SuRe: Surprise-Driven Prioritised Replay for Continual LLM Learning
di: Hazard, Hugo, et al.
Pubblicazione: (2025)
di: Hazard, Hugo, et al.
Pubblicazione: (2025)
Human-inspired Episodic Memory for Infinite Context LLMs
di: Fountas, Zafeirios, et al.
Pubblicazione: (2024)
di: Fountas, Zafeirios, et al.
Pubblicazione: (2024)
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
di: Oomerjee, Adnan, et al.
Pubblicazione: (2025)
di: Oomerjee, Adnan, et al.
Pubblicazione: (2025)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
di: Wieser, Frederico, et al.
Pubblicazione: (2025)
di: Wieser, Frederico, et al.
Pubblicazione: (2025)
Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
di: Fountas, Zafeirios, et al.
Pubblicazione: (2026)
di: Fountas, Zafeirios, et al.
Pubblicazione: (2026)
When in Doubt, Think Slow: Iterative Reasoning with Latent Imagination
di: Benfeghoul, Martin, et al.
Pubblicazione: (2024)
di: Benfeghoul, Martin, et al.
Pubblicazione: (2024)
Mixture of Attentions For Speculative Decoding
di: Zimmer, Matthieu, et al.
Pubblicazione: (2024)
di: Zimmer, Matthieu, et al.
Pubblicazione: (2024)
Emergent Bayesian Behaviour and Optimal Cue Combination in LLMs
di: Ma, Julian, et al.
Pubblicazione: (2025)
di: Ma, Julian, et al.
Pubblicazione: (2025)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
di: Christopoulou, Fenia, et al.
Pubblicazione: (2024)
di: Christopoulou, Fenia, et al.
Pubblicazione: (2024)
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
di: Ji, Xiaotong, et al.
Pubblicazione: (2026)
di: Ji, Xiaotong, et al.
Pubblicazione: (2026)
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
di: Ji, Xiaotong, et al.
Pubblicazione: (2026)
di: Ji, Xiaotong, et al.
Pubblicazione: (2026)
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
di: Zimmer, Matthieu, et al.
Pubblicazione: (2025)
di: Zimmer, Matthieu, et al.
Pubblicazione: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2026)
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2026)
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
di: Zimmer, Matthieu, et al.
Pubblicazione: (2025)
di: Zimmer, Matthieu, et al.
Pubblicazione: (2025)
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
di: Nguyen, Tu, et al.
Pubblicazione: (2026)
di: Nguyen, Tu, et al.
Pubblicazione: (2026)
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
di: Roy, Amartya, et al.
Pubblicazione: (2026)
di: Roy, Amartya, et al.
Pubblicazione: (2026)
Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information
di: Tutnov, Rasul, et al.
Pubblicazione: (2025)
di: Tutnov, Rasul, et al.
Pubblicazione: (2025)
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
di: Sun, Weigao, et al.
Pubblicazione: (2025)
di: Sun, Weigao, et al.
Pubblicazione: (2025)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
di: MiniCPM Team, et al.
Pubblicazione: (2026)
di: MiniCPM Team, et al.
Pubblicazione: (2026)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
di: Chen, Yingfa, et al.
Pubblicazione: (2026)
di: Chen, Yingfa, et al.
Pubblicazione: (2026)
Sample as You Infer: Predictive Coding With Langevin Dynamics
di: Zahid, Umais, et al.
Pubblicazione: (2023)
di: Zahid, Umais, et al.
Pubblicazione: (2023)
A Brain-like Synergistic Core in LLMs Drives Behaviour and Learning
di: Urbina-Rodriguez, Pedro, et al.
Pubblicazione: (2026)
di: Urbina-Rodriguez, Pedro, et al.
Pubblicazione: (2026)
Higher-order Linear Attention
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
di: Zhussip, Magauiya, et al.
Pubblicazione: (2025)
di: Zhussip, Magauiya, et al.
Pubblicazione: (2025)
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
di: Diwan, Anish, et al.
Pubblicazione: (2026)
di: Diwan, Anish, et al.
Pubblicazione: (2026)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
di: Su, Jingtong, et al.
Pubblicazione: (2025)
di: Su, Jingtong, et al.
Pubblicazione: (2025)
Parallax: Parameterized Local Linear Attention for Language Modeling
di: Zuo, Yifei, et al.
Pubblicazione: (2026)
di: Zuo, Yifei, et al.
Pubblicazione: (2026)
Native Hybrid Attention for Efficient Sequence Modeling
di: Du, Jusen, et al.
Pubblicazione: (2025)
di: Du, Jusen, et al.
Pubblicazione: (2025)
TED: Turn Emphasis with Dialogue Feature Attention for Emotion Recognition in Conversation
di: Ono, Junya, et al.
Pubblicazione: (2025)
di: Ono, Junya, et al.
Pubblicazione: (2025)
ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
di: Shopkhoev, Dmitriy, et al.
Pubblicazione: (2025)
di: Shopkhoev, Dmitriy, et al.
Pubblicazione: (2025)
Conversational Tree Search: A New Hybrid Dialog Task
di: Väth, Dirk, et al.
Pubblicazione: (2023)
di: Väth, Dirk, et al.
Pubblicazione: (2023)
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
di: Dentamaro, Vincenzo
Pubblicazione: (2025)
di: Dentamaro, Vincenzo
Pubblicazione: (2025)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
di: Ling Team, et al.
Pubblicazione: (2025)
di: Ling Team, et al.
Pubblicazione: (2025)
Functional Component Ablation Reveals Specialization Patterns in Hybrid Language Model Architectures
di: Borobia, Hector, et al.
Pubblicazione: (2026)
di: Borobia, Hector, et al.
Pubblicazione: (2026)
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
di: You, Haoran, et al.
Pubblicazione: (2024)
di: You, Haoran, et al.
Pubblicazione: (2024)
Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
Training Tensor Attention Efficiently: From Cubic to Almost Linear Time
di: Cao, Yang, et al.
Pubblicazione: (2024)
di: Cao, Yang, et al.
Pubblicazione: (2024)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2025)
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2025)
DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU
di: Chen, Weizhe, et al.
Pubblicazione: (2026)
di: Chen, Weizhe, et al.
Pubblicazione: (2026)
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
di: Ji, Xiaotong, et al.
Pubblicazione: (2025)
di: Ji, Xiaotong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SuRe: Surprise-Driven Prioritised Replay for Continual LLM Learning
di: Hazard, Hugo, et al.
Pubblicazione: (2025) -
Human-inspired Episodic Memory for Infinite Context LLMs
di: Fountas, Zafeirios, et al.
Pubblicazione: (2024) -
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
di: Oomerjee, Adnan, et al.
Pubblicazione: (2025) -
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
di: Wieser, Frederico, et al.
Pubblicazione: (2025) -
Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
di: Fountas, Zafeirios, et al.
Pubblicazione: (2026)