When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | You, Haoran, Fu, Yichao, Wang, Zheng, Yazdanbakhsh, Amir, Lin, Yingyan Celine |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
di: You, Haoran, et al.
Pubblicazione: (2024)
di: You, Haoran, et al.
Pubblicazione: (2024)
ShiftAddNAS: Hardware-Inspired Search for More Accurate and Efficient Neural Networks
di: You, Haoran, et al.
Pubblicazione: (2022)
di: You, Haoran, et al.
Pubblicazione: (2022)
Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed
di: Fu, Yonggan, et al.
Pubblicazione: (2025)
di: Fu, Yonggan, et al.
Pubblicazione: (2025)
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
di: Shi, Dachuan, et al.
Pubblicazione: (2025)
di: Shi, Dachuan, et al.
Pubblicazione: (2025)
ShiftAddViT: Mixture of Multiplication Primitives Towards Efficient Vision Transformer
di: You, Haoran, et al.
Pubblicazione: (2023)
di: You, Haoran, et al.
Pubblicazione: (2023)
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
di: Yu, Zhongzhi, et al.
Pubblicazione: (2024)
di: Yu, Zhongzhi, et al.
Pubblicazione: (2024)
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
di: Haller, Patrick, et al.
Pubblicazione: (2025)
di: Haller, Patrick, et al.
Pubblicazione: (2025)
The Structure of Relation Decoding Linear Operators in Large Language Models
di: Christ, Miranda Anna, et al.
Pubblicazione: (2025)
di: Christ, Miranda Anna, et al.
Pubblicazione: (2025)
Linearly Decoding Refused Knowledge in Aligned Language Models
di: Shrivastava, Aryan, et al.
Pubblicazione: (2025)
di: Shrivastava, Aryan, et al.
Pubblicazione: (2025)
When Compression Meets Model Compression: Memory-Efficient Double Compression for Large Language Models
di: Wang, Weilan, et al.
Pubblicazione: (2025)
di: Wang, Weilan, et al.
Pubblicazione: (2025)
Contextual Attention Modulation: Towards Efficient Multi-Task Adaptation in Large Language Models
di: Pan, Dayan, et al.
Pubblicazione: (2025)
di: Pan, Dayan, et al.
Pubblicazione: (2025)
EffGen: Enabling Small Language Models as Capable Autonomous Agents
di: Srivastava, Gaurav, et al.
Pubblicazione: (2026)
di: Srivastava, Gaurav, et al.
Pubblicazione: (2026)
Beyond Moore's Law: Harnessing the Redshift of Generative AI with Effective Hardware-Software Co-Design
di: Yazdanbakhsh, Amir
Pubblicazione: (2025)
di: Yazdanbakhsh, Amir
Pubblicazione: (2025)
RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
di: Goldstein, Daniel, et al.
Pubblicazione: (2025)
di: Goldstein, Daniel, et al.
Pubblicazione: (2025)
When Large Language Models Meet Personalization: Perspectives of Challenges and Opportunities
di: Chen, Jin, et al.
Pubblicazione: (2023)
di: Chen, Jin, et al.
Pubblicazione: (2023)
Parallax: Parameterized Local Linear Attention for Language Modeling
di: Zuo, Yifei, et al.
Pubblicazione: (2026)
di: Zuo, Yifei, et al.
Pubblicazione: (2026)
Efficient and Effective Vocabulary Expansion Towards Multilingual Large Language Models
di: Kim, Seungduk, et al.
Pubblicazione: (2024)
di: Kim, Seungduk, et al.
Pubblicazione: (2024)
Plato: Plan to Efficiently Decode for Large Language Model Inference
di: Jin, Shuowei, et al.
Pubblicazione: (2024)
di: Jin, Shuowei, et al.
Pubblicazione: (2024)
ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
di: Li, Jia-Nan, et al.
Pubblicazione: (2025)
di: Li, Jia-Nan, et al.
Pubblicazione: (2025)
Early-Bird GCNs: Graph-Network Co-Optimization Towards More Efficient GCN Training and Inference via Drawing Early-Bird Lottery Tickets
di: You, Haoran, et al.
Pubblicazione: (2021)
di: You, Haoran, et al.
Pubblicazione: (2021)
Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model
di: Ding, Bowen, et al.
Pubblicazione: (2025)
di: Ding, Bowen, et al.
Pubblicazione: (2025)
Sliding Window Attention Training for Efficient Large Language Models
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
di: Fu, Tairan, et al.
Pubblicazione: (2025)
di: Fu, Tairan, et al.
Pubblicazione: (2025)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
di: Chen, Yingfa, et al.
Pubblicazione: (2026)
di: Chen, Yingfa, et al.
Pubblicazione: (2026)
Identifying Linear Relational Concepts in Large Language Models
di: Chanin, David, et al.
Pubblicazione: (2023)
di: Chanin, David, et al.
Pubblicazione: (2023)
KGLens: Towards Efficient and Effective Knowledge Probing of Large Language Models with Knowledge Graphs
di: Zheng, Shangshang, et al.
Pubblicazione: (2023)
di: Zheng, Shangshang, et al.
Pubblicazione: (2023)
On Linearizing Structured Data in Encoder-Decoder Language Models: Insights from Text-to-SQL
di: Shao, Yutong, et al.
Pubblicazione: (2024)
di: Shao, Yutong, et al.
Pubblicazione: (2024)
BlockBatch: Multi-Scale Consensus Decoding for Efficient Diffusion Language Model Inference
di: Wu, Xiaoyou, et al.
Pubblicazione: (2026)
di: Wu, Xiaoyou, et al.
Pubblicazione: (2026)
Equivalent Linear Mappings of Large Language Models
di: Golden, James R.
Pubblicazione: (2025)
di: Golden, James R.
Pubblicazione: (2025)
Hierarchical Skip Decoding for Efficient Autoregressive Text Generation
di: Zhu, Yunqi, et al.
Pubblicazione: (2024)
di: Zhu, Yunqi, et al.
Pubblicazione: (2024)
FinLLM-B: When Large Language Models Meet Financial Breakout Trading
di: Zhang, Kang, et al.
Pubblicazione: (2024)
di: Zhang, Kang, et al.
Pubblicazione: (2024)
HICD: Hallucination-Inducing via Attention Dispersion for Contrastive Decoding to Mitigate Hallucinations in Large Language Models
di: Jiang, Xinyan, et al.
Pubblicazione: (2025)
di: Jiang, Xinyan, et al.
Pubblicazione: (2025)
Hitting "Probe"rty with Non-Linearity, and More
di: Pal, Avik, et al.
Pubblicazione: (2024)
di: Pal, Avik, et al.
Pubblicazione: (2024)
Generation Meets Verification: Accelerating Large Language Model Inference with Smart Parallel Auto-Correct Decoding
di: Yi, Hanling, et al.
Pubblicazione: (2024)
di: Yi, Hanling, et al.
Pubblicazione: (2024)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
di: MiniCPM Team, et al.
Pubblicazione: (2026)
di: MiniCPM Team, et al.
Pubblicazione: (2026)
Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference
di: You, Haoran, et al.
Pubblicazione: (2022)
di: You, Haoran, et al.
Pubblicazione: (2022)
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
di: Yang, Ning, et al.
Pubblicazione: (2026)
di: Yang, Ning, et al.
Pubblicazione: (2026)
When Large Language Models Meet Vector Databases: A Survey
di: Jing, Zhi, et al.
Pubblicazione: (2024)
di: Jing, Zhi, et al.
Pubblicazione: (2024)
CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
di: Sun, Weigao, et al.
Pubblicazione: (2025)
di: Sun, Weigao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
di: You, Haoran, et al.
Pubblicazione: (2024) -
ShiftAddNAS: Hardware-Inspired Search for More Accurate and Efficient Neural Networks
di: You, Haoran, et al.
Pubblicazione: (2022) -
Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed
di: Fu, Yonggan, et al.
Pubblicazione: (2025) -
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
di: Shi, Dachuan, et al.
Pubblicazione: (2025) -
ShiftAddViT: Mixture of Multiplication Primitives Towards Efficient Vision Transformer
di: You, Haoran, et al.
Pubblicazione: (2023)