Hierarchical Skip Decoding for Efficient Autoregressive Text Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Yunqi, Yang, Xuebing, Wu, Yuanyuan, Zhang, Wensheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Potential of LLMs in Medical Education: Generating Questions and Answers for Qualification Exams
von: Zhu, Yunqi, et al.
Veröffentlicht: (2024)
von: Zhu, Yunqi, et al.
Veröffentlicht: (2024)
Pipelined Decoder for Efficient Context-Aware Text Generation
von: Huang, Zixian, et al.
Veröffentlicht: (2025)
von: Huang, Zixian, et al.
Veröffentlicht: (2025)
Differences in Text Generated by Diffusion and Autoregressive Language Models
von: Zhang, Zeyang, et al.
Veröffentlicht: (2026)
von: Zhang, Zeyang, et al.
Veröffentlicht: (2026)
Breaking the Autoregressive Chain: Hyper-Parallel Decoding for Efficient LLM-Based Attribute Value Extraction
von: Glavas, Theodore, et al.
Veröffentlicht: (2026)
von: Glavas, Theodore, et al.
Veröffentlicht: (2026)
HiChunk: Evaluating and Enhancing Retrieval-Augmented Generation with Hierarchical Chunking
von: Lu, Wensheng, et al.
Veröffentlicht: (2025)
von: Lu, Wensheng, et al.
Veröffentlicht: (2025)
AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
von: He, Zhuomin, et al.
Veröffentlicht: (2025)
von: He, Zhuomin, et al.
Veröffentlicht: (2025)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
Trustful LLMs: Customizing and Grounding Text Generation with Knowledge Bases and Dual Decoders
von: Zhu, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Zhu, Xiaofeng, et al.
Veröffentlicht: (2024)
WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models
von: Lee, Hanna, et al.
Veröffentlicht: (2026)
von: Lee, Hanna, et al.
Veröffentlicht: (2026)
Generating Diverse and High-Quality Texts by Minimum Bayes Risk Decoding
von: Jinnai, Yuu, et al.
Veröffentlicht: (2024)
von: Jinnai, Yuu, et al.
Veröffentlicht: (2024)
PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block Skipping
von: Lu, Yu-Chen, et al.
Veröffentlicht: (2025)
von: Lu, Yu-Chen, et al.
Veröffentlicht: (2025)
Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
von: Bergner, Benjamin, et al.
Veröffentlicht: (2024)
von: Bergner, Benjamin, et al.
Veröffentlicht: (2024)
Model-Based Minimum Bayes Risk Decoding for Text Generation
von: Jinnai, Yuu, et al.
Veröffentlicht: (2023)
von: Jinnai, Yuu, et al.
Veröffentlicht: (2023)
A Comparative Study of Decoding Strategies in Medical Text Generation
von: Presacan, Oriana, et al.
Veröffentlicht: (2025)
von: Presacan, Oriana, et al.
Veröffentlicht: (2025)
DiffScore: Text Evaluation Beyond Autoregressive Likelihood
von: Lai, Wen, et al.
Veröffentlicht: (2026)
von: Lai, Wen, et al.
Veröffentlicht: (2026)
Why Diffusion Language Models Struggle with Truly Parallel (Non-Autoregressive) Decoding?
von: Li, Pengxiang, et al.
Veröffentlicht: (2026)
von: Li, Pengxiang, et al.
Veröffentlicht: (2026)
Autoregressive Models Rival Diffusion Models at ANY-ORDER Generation
von: Du, Tianqi, et al.
Veröffentlicht: (2026)
von: Du, Tianqi, et al.
Veröffentlicht: (2026)
Avoiding Knowledge Edit Skipping in Multi-hop Question Answering with Guided Decomposition
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
Adaptive GoGI-Skip: Coupling Goal-Gradient Importance with Dynamic Uncertainty for Efficient Reasoning
von: Zhuang, Ren
Veröffentlicht: (2025)
von: Zhuang, Ren
Veröffentlicht: (2025)
Enhancing Retrieval Augmented Generation with Hierarchical Text Segmentation Chunking
von: Nguyen, Hai Toan, et al.
Veröffentlicht: (2025)
von: Nguyen, Hai Toan, et al.
Veröffentlicht: (2025)
VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping
von: Dong, Haotian, et al.
Veröffentlicht: (2025)
von: Dong, Haotian, et al.
Veröffentlicht: (2025)
Federated Learning with Layer Skipping: Efficient Training of Large Language Models for Healthcare NLP
von: Zhang, Lihong, et al.
Veröffentlicht: (2025)
von: Zhang, Lihong, et al.
Veröffentlicht: (2025)
ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
von: Li, Jia-Nan, et al.
Veröffentlicht: (2025)
von: Li, Jia-Nan, et al.
Veröffentlicht: (2025)
Reflection-Window Decoding: Text Generation with Selective Refinement
von: Tang, Zeyu, et al.
Veröffentlicht: (2025)
von: Tang, Zeyu, et al.
Veröffentlicht: (2025)
Exploiting Text Semantics for Few and Zero Shot Node Classification on Text-attributed Graph
von: Wang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wang, Yuxiang, et al.
Veröffentlicht: (2025)
Enhancing Text Authenticity: A Novel Hybrid Approach for AI-Generated Text Detection
von: Zhang, Ye, et al.
Veröffentlicht: (2024)
von: Zhang, Ye, et al.
Veröffentlicht: (2024)
Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
von: Zhang, Yong, et al.
Veröffentlicht: (2025)
von: Zhang, Yong, et al.
Veröffentlicht: (2025)
TokenSkip: Controllable Chain-of-Thought Compression in LLMs
von: Xia, Heming, et al.
Veröffentlicht: (2025)
von: Xia, Heming, et al.
Veröffentlicht: (2025)
Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport
von: Jinnai, Yuu
Veröffentlicht: (2025)
von: Jinnai, Yuu
Veröffentlicht: (2025)
Unlocking Anticipatory Text Generation: A Constrained Approach for Large Language Models Decoding
von: Tu, Lifu, et al.
Veröffentlicht: (2023)
von: Tu, Lifu, et al.
Veröffentlicht: (2023)
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
von: You, Haoran, et al.
Veröffentlicht: (2024)
von: You, Haoran, et al.
Veröffentlicht: (2024)
Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese
von: Xu, Yunqi, et al.
Veröffentlicht: (2024)
von: Xu, Yunqi, et al.
Veröffentlicht: (2024)
Hierarchical Memory Organization for Wikipedia Generation
von: Yu, Eugene J., et al.
Veröffentlicht: (2025)
von: Yu, Eugene J., et al.
Veröffentlicht: (2025)
Plato: Plan to Efficiently Decode for Large Language Model Inference
von: Jin, Shuowei, et al.
Veröffentlicht: (2024)
von: Jin, Shuowei, et al.
Veröffentlicht: (2024)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
von: Zhang, Zhuoyang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhuoyang, et al.
Veröffentlicht: (2025)
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
von: Zhao, Haozhe, et al.
Veröffentlicht: (2025)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Potential of LLMs in Medical Education: Generating Questions and Answers for Qualification Exams
von: Zhu, Yunqi, et al.
Veröffentlicht: (2024) -
Pipelined Decoder for Efficient Context-Aware Text Generation
von: Huang, Zixian, et al.
Veröffentlicht: (2025) -
Differences in Text Generated by Diffusion and Autoregressive Language Models
von: Zhang, Zeyang, et al.
Veröffentlicht: (2026) -
Breaking the Autoregressive Chain: Hyper-Parallel Decoding for Efficient LLM-Based Attribute Value Extraction
von: Glavas, Theodore, et al.
Veröffentlicht: (2026) -
HiChunk: Evaluating and Enhancing Retrieval-Augmented Generation with Hierarchical Chunking
von: Lu, Wensheng, et al.
Veröffentlicht: (2025)