TokenSwift: Lossless Acceleration of Ultra Long Sequence Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Tong, Shen, Junzhe, Jia, Zixia, Wang, Yuxuan, Zheng, Zilong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
von: Lou, Chao, et al.
Veröffentlicht: (2024)
von: Lou, Chao, et al.
Veröffentlicht: (2024)
LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model
von: Shao, Wei, et al.
Veröffentlicht: (2025)
von: Shao, Wei, et al.
Veröffentlicht: (2025)
Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
von: Bai, Jun, et al.
Veröffentlicht: (2025)
von: Bai, Jun, et al.
Veröffentlicht: (2025)
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
LangSuitE: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments
von: Jia, Zixia, et al.
Veröffentlicht: (2024)
von: Jia, Zixia, et al.
Veröffentlicht: (2024)
Lossless Token Sequence Compression via Meta-Tokens
von: Harvill, John, et al.
Veröffentlicht: (2025)
von: Harvill, John, et al.
Veröffentlicht: (2025)
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
von: Wang, Xiaobo, et al.
Veröffentlicht: (2025)
von: Wang, Xiaobo, et al.
Veröffentlicht: (2025)
An Efficient Recipe for Long Context Extension via Middle-Focused Positional Encoding
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
Towards Lossless Token Pruning in Late-Interaction Retrieval Models
von: Zong, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zong, Yuxuan, et al.
Veröffentlicht: (2025)
Combining Supervised Learning and Reinforcement Learning for Multi-Label Classification Tasks with Partial Labels
von: Jia, Zixia, et al.
Veröffentlicht: (2024)
von: Jia, Zixia, et al.
Veröffentlicht: (2024)
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
von: Wu, Tong, et al.
Veröffentlicht: (2025)
von: Wu, Tong, et al.
Veröffentlicht: (2025)
RAM: Towards an Ever-Improving Memory System by Learning from Communications
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
BEDA: Belief Estimation as Probabilistic Constraints for Performing Strategic Dialogue Acts
von: Li, Hengli, et al.
Veröffentlicht: (2025)
von: Li, Hengli, et al.
Veröffentlicht: (2025)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Chimera: A Lossless Decoding Method for Accelerating Large Language Models Inference by Fusing all Tokens
von: Zeng, Ziqian, et al.
Veröffentlicht: (2024)
von: Zeng, Ziqian, et al.
Veröffentlicht: (2024)
Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space
von: Li, Hengli, et al.
Veröffentlicht: (2025)
von: Li, Hengli, et al.
Veröffentlicht: (2025)
Make an Offer They Can't Refuse: Grounding Bayesian Persuasion in Real-World Dialogues without Pre-Commitment
von: He, Buwei, et al.
Veröffentlicht: (2025)
von: He, Buwei, et al.
Veröffentlicht: (2025)
Parallel Decoding via Hidden Transfer for Lossless Large Language Model Acceleration
von: Wu, Pengfei, et al.
Veröffentlicht: (2024)
von: Wu, Pengfei, et al.
Veröffentlicht: (2024)
Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding
von: Wang, Pei-Shuo, et al.
Veröffentlicht: (2025)
von: Wang, Pei-Shuo, et al.
Veröffentlicht: (2025)
SwiftEmbed: Ultra-Fast Text Embeddings via Static Token Lookup for Real-Time Applications
von: Lansiaux, Edouard, et al.
Veröffentlicht: (2025)
von: Lansiaux, Edouard, et al.
Veröffentlicht: (2025)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
von: Mao, Yu, et al.
Veröffentlicht: (2025)
von: Mao, Yu, et al.
Veröffentlicht: (2025)
Team Ryu's Submission to SIGMORPHON 2024 Shared Task on Subword Tokenization
von: Li, Zilong
Veröffentlicht: (2024)
von: Li, Zilong
Veröffentlicht: (2024)
Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models
von: Hu, Xiang, et al.
Veröffentlicht: (2025)
von: Hu, Xiang, et al.
Veröffentlicht: (2025)
LooGLE: Can Long-Context Language Models Understand Long Contexts?
von: Li, Jiaqi, et al.
Veröffentlicht: (2023)
von: Li, Jiaqi, et al.
Veröffentlicht: (2023)
Behavior-Equivalent Token: Single-Token Replacement for Long Prompts in LLMs
von: Dong, Jiancheng, et al.
Veröffentlicht: (2025)
von: Dong, Jiancheng, et al.
Veröffentlicht: (2025)
Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs
von: Zhang, Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Xuan, et al.
Veröffentlicht: (2025)
Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
von: Zhang, Jun, et al.
Veröffentlicht: (2023)
von: Zhang, Jun, et al.
Veröffentlicht: (2023)
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
von: Yan, Shaotian, et al.
Veröffentlicht: (2026)
von: Yan, Shaotian, et al.
Veröffentlicht: (2026)
LongRoPE2: Near-Lossless LLM Context Window Scaling
von: Shang, Ning, et al.
Veröffentlicht: (2025)
von: Shang, Ning, et al.
Veröffentlicht: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
From Token to Line: Enhancing Code Generation with a Long-Term Perspective
von: Lu, Tingwei, et al.
Veröffentlicht: (2025)
von: Lu, Tingwei, et al.
Veröffentlicht: (2025)
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement
von: Wang, Xiaobo, et al.
Veröffentlicht: (2026)
von: Wang, Xiaobo, et al.
Veröffentlicht: (2026)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
Hybrid OCR-LLM Framework for Enterprise-Scale Document Information Extraction Under Copy-heavy Task
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning
von: Long, Lingkun, et al.
Veröffentlicht: (2025)
von: Long, Lingkun, et al.
Veröffentlicht: (2025)
Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding
von: Cho, Sukmin, et al.
Veröffentlicht: (2025)
von: Cho, Sukmin, et al.
Veröffentlicht: (2025)
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
ResFormer: All-Time Reservoir Memory for Long Sequence Classification
von: Liu, Hongbo, et al.
Veröffentlicht: (2025)
von: Liu, Hongbo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
von: Lou, Chao, et al.
Veröffentlicht: (2024) -
LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model
von: Shao, Wei, et al.
Veröffentlicht: (2025) -
Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
von: Bai, Jun, et al.
Veröffentlicht: (2025) -
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
von: Sun, Hanshi, et al.
Veröffentlicht: (2024) -
LangSuitE: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments
von: Jia, Zixia, et al.
Veröffentlicht: (2024)