Gespeichert in:
| Hauptverfasser: | Yao, Yuncheng, Xia, Yuxuan, Wang, Shengjie, Zhuo, Danyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2605.04263 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training
von: Wang, Xi, et al.
Veröffentlicht: (2026)
von: Wang, Xi, et al.
Veröffentlicht: (2026)
TAPS: Target-Aware Prefix Tree Selection for Diffusion-Drafted Speculative Decoding
von: Wang, Zhuoyu, et al.
Veröffentlicht: (2026)
von: Wang, Zhuoyu, et al.
Veröffentlicht: (2026)
DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution
von: Hu, Yunhai, et al.
Veröffentlicht: (2026)
von: Hu, Yunhai, et al.
Veröffentlicht: (2026)
HilbertA: Hilbert Attention for Image Generation with Diffusion Models
von: Zheng, Shaoyi, et al.
Veröffentlicht: (2025)
von: Zheng, Shaoyi, et al.
Veröffentlicht: (2025)
ToMA: Token Merge with Attention for Diffusion Models
von: Lu, Wenbo, et al.
Veröffentlicht: (2025)
von: Lu, Wenbo, et al.
Veröffentlicht: (2025)
VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping
von: Dong, Haotian, et al.
Veröffentlicht: (2025)
von: Dong, Haotian, et al.
Veröffentlicht: (2025)
Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs
von: Athiwaratkun, Ben, et al.
Veröffentlicht: (2024)
von: Athiwaratkun, Ben, et al.
Veröffentlicht: (2024)
PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
von: Wu, Tianyu, et al.
Veröffentlicht: (2026)
von: Wu, Tianyu, et al.
Veröffentlicht: (2026)
Traversal Verification for Speculative Tree Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
PACER: Blockwise Pre-verification for Speculative Decoding with Adaptive Length
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
Hydra: Efficient, Correct Code Generation via Checkpoint-and-Rollback Support
von: Du, Alexander, et al.
Veröffentlicht: (2026)
von: Du, Alexander, et al.
Veröffentlicht: (2026)
PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
von: Ding, Ruogu, et al.
Veröffentlicht: (2025)
von: Ding, Ruogu, et al.
Veröffentlicht: (2025)
PrefixLLM: LLM-aided Prefix Circuit Design
von: Xiao, Weihua, et al.
Veröffentlicht: (2024)
von: Xiao, Weihua, et al.
Veröffentlicht: (2024)
WISV: Wireless-Informed Semantic Verification for Distributed Speculative Decoding in Device-Edge LLM Inference
von: Liu, Zixuan, et al.
Veröffentlicht: (2026)
von: Liu, Zixuan, et al.
Veröffentlicht: (2026)
Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding
von: Su, Xin, et al.
Veröffentlicht: (2026)
von: Su, Xin, et al.
Veröffentlicht: (2026)
Overcoming Joint Intractability with Lossless Hierarchical Speculative Decoding
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2026)
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2026)
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
First Ask Then Answer: A Framework Design for AI Dialogue Based on Supplementary Questioning with Large Language Models
von: Fu, Chuanruo, et al.
Veröffentlicht: (2025)
von: Fu, Chuanruo, et al.
Veröffentlicht: (2025)
Hypothesize-Then-Verify: Speculative Root Cause Analysis for Microservices with Pathwise Parallelism
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2026)
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2026)
PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2025)
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2025)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
ConfSpec: Efficient Step-Level Speculative Reasoning via Confidence-Gated Verification
von: Liu, Siran, et al.
Veröffentlicht: (2026)
von: Liu, Siran, et al.
Veröffentlicht: (2026)
Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference
von: Timor, Nadav, et al.
Veröffentlicht: (2024)
von: Timor, Nadav, et al.
Veröffentlicht: (2024)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
von: An, Zihao, et al.
Veröffentlicht: (2026)
von: An, Zihao, et al.
Veröffentlicht: (2026)
DIVERSED: Relaxed Speculative Decoding via Dynamic Ensemble Verification
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
Enhancing High-Quality Code Generation in Large Language Models with Comparative Prefix-Tuning
von: Jiang, Yuan, et al.
Veröffentlicht: (2025)
von: Jiang, Yuan, et al.
Veröffentlicht: (2025)
SAM Decoding: Speculative Decoding via Suffix Automaton
von: Hu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2024)
Human-Guided Image Generation for Expanding Small-Scale Training Image Datasets
von: Chen, Changjian, et al.
Veröffentlicht: (2024)
von: Chen, Changjian, et al.
Veröffentlicht: (2024)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
von: Li, Ruanjun, et al.
Veröffentlicht: (2025)
von: Li, Ruanjun, et al.
Veröffentlicht: (2025)
Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
von: Wan, Yuxuan, et al.
Veröffentlicht: (2026)
von: Wan, Yuxuan, et al.
Veröffentlicht: (2026)
Layered LA-MAPF: a decomposition of large agent MAPF instance to accelerate solving without compromising solvability
von: Yao, Zhuo
Veröffentlicht: (2024)
von: Yao, Zhuo
Veröffentlicht: (2024)
Generating Visual Stories with Grounded and Coreferent Characters
von: Liu, Danyang, et al.
Veröffentlicht: (2024)
von: Liu, Danyang, et al.
Veröffentlicht: (2024)
Accelerating Large Language Model Reasoning via Speculative Search
von: Wang, Zhihai, et al.
Veröffentlicht: (2025)
von: Wang, Zhihai, et al.
Veröffentlicht: (2025)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training
von: Wang, Xi, et al.
Veröffentlicht: (2026) -
TAPS: Target-Aware Prefix Tree Selection for Diffusion-Drafted Speculative Decoding
von: Wang, Zhuoyu, et al.
Veröffentlicht: (2026) -
DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution
von: Hu, Yunhai, et al.
Veröffentlicht: (2026) -
HilbertA: Hilbert Attention for Image Generation with Diffusion Models
von: Zheng, Shaoyi, et al.
Veröffentlicht: (2025) -
ToMA: Token Merge with Attention for Diffusion Models
von: Lu, Wenbo, et al.
Veröffentlicht: (2025)