Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jinze, Xu, Yixing, Li, Guanchen, Yang, Shuo, Xu, Jinfeng, Yin, Xuanwu, Li, Dong, Ngai, Edith C. H., Barsoum, Emad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond the Target: From Imitation to Collaboration in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2026)
von: Li, Jinze, et al.
Veröffentlicht: (2026)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2025)
von: Li, Jinze, et al.
Veröffentlicht: (2025)
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution Optimization
von: Li, Guanchen, et al.
Veröffentlicht: (2025)
von: Li, Guanchen, et al.
Veröffentlicht: (2025)
Learnable Permutation for Structured Sparsity on Transformer Models
von: Li, Zekai, et al.
Veröffentlicht: (2026)
von: Li, Zekai, et al.
Veröffentlicht: (2026)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
von: An, Zihao, et al.
Veröffentlicht: (2026)
von: An, Zihao, et al.
Veröffentlicht: (2026)
Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates
von: Xu, Yixing, et al.
Veröffentlicht: (2025)
von: Xu, Yixing, et al.
Veröffentlicht: (2025)
Well Begun is Half Done: Training-Free and Model-Agnostic Semantically Guaranteed User Representation Initialization for Multimodal Recommendation
von: Xu, Jinfeng, et al.
Veröffentlicht: (2026)
von: Xu, Jinfeng, et al.
Veröffentlicht: (2026)
The Best is Yet to Come: Graph Convolution in the Testing Phase for Multimodal Recommendation
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
Enhancing Graph Collaborative Filtering with FourierKAN Feature Transformation
von: Xu, Jinfeng, et al.
Veröffentlicht: (2024)
von: Xu, Jinfeng, et al.
Veröffentlicht: (2024)
AlignGroup: Learning and Aligning Group Consensus with Member Preferences for Group Recommendation
von: Xu, Jinfeng, et al.
Veröffentlicht: (2024)
von: Xu, Jinfeng, et al.
Veröffentlicht: (2024)
MENTOR: Multi-level Self-supervised Learning for Multimodal Recommendation
von: Xu, Jinfeng, et al.
Veröffentlicht: (2024)
von: Xu, Jinfeng, et al.
Veröffentlicht: (2024)
Learning and Editing Universal Graph Prompt Tuning via Reinforcement Learning
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
MSWA: Refining Local Attention with Multi-ScaleWindow Attention
von: Xu, Yixing, et al.
Veröffentlicht: (2025)
von: Xu, Yixing, et al.
Veröffentlicht: (2025)
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
von: Hong, Fenglu, et al.
Veröffentlicht: (2025)
von: Hong, Fenglu, et al.
Veröffentlicht: (2025)
VI-MMRec: Similarity-Aware Training Cost-free Virtual User-Item Interactions for Multimodal Recommendation
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
NLGCL: Naturally Existing Neighbor Layers Graph Contrastive Learning for Recommendation
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
Generative AI for Vulnerability Detection in 6G Wireless Networks: Advances, Case Study, and Future Directions
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
A Survey on Multimodal Recommender Systems: Recent Advances and Future Directions
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
von: An, Zihao, et al.
Veröffentlicht: (2025)
von: An, Zihao, et al.
Veröffentlicht: (2025)
Make Every Draft Count: Hidden State based Speculative Decoding
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
MDVT: Enhancing Multimodal Recommendation with Model-Agnostic Multimodal-Driven Virtual Triplets
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference
von: Li, Zeping, et al.
Veröffentlicht: (2024)
von: Li, Zeping, et al.
Veröffentlicht: (2024)
Draft, Verify, and Improve: Toward Training-Aware Speculative Decoding
von: Bhansali, Shrenik, et al.
Veröffentlicht: (2025)
von: Bhansali, Shrenik, et al.
Veröffentlicht: (2025)
Theory-optimal Quantization Based on Flatness
von: Huang, Xiusheng, et al.
Veröffentlicht: (2026)
von: Huang, Xiusheng, et al.
Veröffentlicht: (2026)
OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory
von: Li, Jinze, et al.
Veröffentlicht: (2026)
von: Li, Jinze, et al.
Veröffentlicht: (2026)
Large Language Models for Network Intrusion Detection Systems: Foundations, Implementations, and Future Directions
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
LATTE: Forecasting Peer Anchored Preference Trajectories for Personalized LLM Generation
von: Li, Jinze, et al.
Veröffentlicht: (2026)
von: Li, Jinze, et al.
Veröffentlicht: (2026)
PEARL: Parallel Speculative Decoding with Adaptive Draft Length
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMs
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
CAMMSR: Category-Guided Attentive Mixture of Experts for Multimodal Sequential Recommendation
von: Xu, Jinfeng, et al.
Veröffentlicht: (2026)
von: Xu, Jinfeng, et al.
Veröffentlicht: (2026)
SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting
von: Shi, Weijie, et al.
Veröffentlicht: (2026)
von: Shi, Weijie, et al.
Veröffentlicht: (2026)
Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
von: Zhao, Weilin, et al.
Veröffentlicht: (2024)
von: Zhao, Weilin, et al.
Veröffentlicht: (2024)
OPT-Tree: Speculative Decoding with Adaptive Draft Tree Structure
von: Wang, Jikai, et al.
Veröffentlicht: (2024)
von: Wang, Jikai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond the Target: From Imitation to Collaboration in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2026) -
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2025) -
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
von: Huang, Haiduo, et al.
Veröffentlicht: (2025) -
Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution Optimization
von: Li, Guanchen, et al.
Veröffentlicht: (2025) -
Learnable Permutation for Structured Sparsity on Transformer Models
von: Li, Zekai, et al.
Veröffentlicht: (2026)