Beyond the Target: From Imitation to Collaboration in Speculative Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jinze, Xu, Yixing, Li, Guanchen, Xu, Jinfeng, Yang, Shuo, Zhang, Yang, Yin, Xuanwu, Li, Dong, Ngai, Edith C. H., Barsoum, Emad |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
by: Li, Jinze, et al.
Published: (2025)
by: Li, Jinze, et al.
Published: (2025)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
by: Li, Jinze, et al.
Published: (2025)
by: Li, Jinze, et al.
Published: (2025)
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Learnable Permutation for Structured Sparsity on Transformer Models
by: Li, Zekai, et al.
Published: (2026)
by: Li, Zekai, et al.
Published: (2026)
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
by: Liao, Huanxuan, et al.
Published: (2025)
by: Liao, Huanxuan, et al.
Published: (2025)
Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates
by: Xu, Yixing, et al.
Published: (2025)
by: Xu, Yixing, et al.
Published: (2025)
Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution Optimization
by: Li, Guanchen, et al.
Published: (2025)
by: Li, Guanchen, et al.
Published: (2025)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
by: An, Zihao, et al.
Published: (2026)
by: An, Zihao, et al.
Published: (2026)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
MSWA: Refining Local Attention with Multi-ScaleWindow Attention
by: Xu, Yixing, et al.
Published: (2025)
by: Xu, Yixing, et al.
Published: (2025)
OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory
by: Li, Jinze, et al.
Published: (2026)
by: Li, Jinze, et al.
Published: (2026)
LATTE: Forecasting Peer Anchored Preference Trajectories for Personalized LLM Generation
by: Li, Jinze, et al.
Published: (2026)
by: Li, Jinze, et al.
Published: (2026)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference
by: Li, Zeping, et al.
Published: (2024)
by: Li, Zeping, et al.
Published: (2024)
RAMA: Retrieval-Augmented Multi-Agent Framework for Misinformation Detection in Multimodal Fact-Checking
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
The Best is Yet to Come: Graph Convolution in the Testing Phase for Multimodal Recommendation
by: Xu, Jinfeng, et al.
Published: (2025)
by: Xu, Jinfeng, et al.
Published: (2025)
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
Enhancing Graph Collaborative Filtering with FourierKAN Feature Transformation
by: Xu, Jinfeng, et al.
Published: (2024)
by: Xu, Jinfeng, et al.
Published: (2024)
AlignGroup: Learning and Aligning Group Consensus with Member Preferences for Group Recommendation
by: Xu, Jinfeng, et al.
Published: (2024)
by: Xu, Jinfeng, et al.
Published: (2024)
MENTOR: Multi-level Self-supervised Learning for Multimodal Recommendation
by: Xu, Jinfeng, et al.
Published: (2024)
by: Xu, Jinfeng, et al.
Published: (2024)
Learning and Editing Universal Graph Prompt Tuning via Reinforcement Learning
by: Xu, Jinfeng, et al.
Published: (2025)
by: Xu, Jinfeng, et al.
Published: (2025)
Enhancing One-shot Pruned Pre-trained Language Models through Sparse-Dense-Sparse Mechanism
by: Li, Guanchen, et al.
Published: (2024)
by: Li, Guanchen, et al.
Published: (2024)
Generative AI for Vulnerability Detection in 6G Wireless Networks: Advances, Case Study, and Future Directions
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
Fast Large Language Model Collaborative Decoding via Speculation
by: Fu, Jiale, et al.
Published: (2025)
by: Fu, Jiale, et al.
Published: (2025)
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
by: Li, Guihong, et al.
Published: (2025)
by: Li, Guihong, et al.
Published: (2025)
Speculative Decoding and Beyond: An In-Depth Survey of Techniques
by: Hu, Yunhai, et al.
Published: (2025)
by: Hu, Yunhai, et al.
Published: (2025)
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
by: Joshi, Vinay, et al.
Published: (2025)
by: Joshi, Vinay, et al.
Published: (2025)
Zebra-Llama: Towards Extremely Efficient Hybrid Models
by: Yang, Mingyu, et al.
Published: (2025)
by: Yang, Mingyu, et al.
Published: (2025)
Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding
by: Wang, Ziyao, et al.
Published: (2025)
by: Wang, Ziyao, et al.
Published: (2025)
NLGCL: Naturally Existing Neighbor Layers Graph Contrastive Learning for Recommendation
by: Xu, Jinfeng, et al.
Published: (2025)
by: Xu, Jinfeng, et al.
Published: (2025)
A Survey on Multimodal Recommender Systems: Recent Advances and Future Directions
by: Xu, Jinfeng, et al.
Published: (2025)
by: Xu, Jinfeng, et al.
Published: (2025)
Speculative Decoding with a Speculative Vocabulary
by: Williams, Miles, et al.
Published: (2026)
by: Williams, Miles, et al.
Published: (2026)
Large Language Models for Network Intrusion Detection Systems: Foundations, Implementations, and Future Directions
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
MDVT: Enhancing Multimodal Recommendation with Model-Agnostic Multimodal-Driven Virtual Triplets
by: Xu, Jinfeng, et al.
Published: (2025)
by: Xu, Jinfeng, et al.
Published: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding
by: Du, Cunxiao, et al.
Published: (2024)
by: Du, Cunxiao, et al.
Published: (2024)
CAMMSR: Category-Guided Attentive Mixture of Experts for Multimodal Sequential Recommendation
by: Xu, Jinfeng, et al.
Published: (2026)
by: Xu, Jinfeng, et al.
Published: (2026)
VI-MMRec: Similarity-Aware Training Cost-free Virtual User-Item Interactions for Multimodal Recommendation
by: Xu, Jinfeng, et al.
Published: (2025)
by: Xu, Jinfeng, et al.
Published: (2025)
Spec-LLaVA: Accelerating Vision-Language Models with Dynamic Tree-Based Speculative Decoding
by: Huo, Mingxiao, et al.
Published: (2025)
by: Huo, Mingxiao, et al.
Published: (2025)
Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inference
by: Zhang, Libo, et al.
Published: (2024)
by: Zhang, Libo, et al.
Published: (2024)
Similar Items
-
Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
by: Li, Jinze, et al.
Published: (2025) -
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
by: Li, Jinze, et al.
Published: (2025) -
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
by: Huang, Haiduo, et al.
Published: (2025) -
Learnable Permutation for Structured Sparsity on Transformer Models
by: Li, Zekai, et al.
Published: (2026) -
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
by: Liao, Huanxuan, et al.
Published: (2025)