Salvato in:
| Autori principali: | Yao, Yuncheng, Xia, Yuxuan, Wang, Shengjie, Zhuo, Danyang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2605.04263 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training
di: Wang, Xi, et al.
Pubblicazione: (2026)
di: Wang, Xi, et al.
Pubblicazione: (2026)
TAPS: Target-Aware Prefix Tree Selection for Diffusion-Drafted Speculative Decoding
di: Wang, Zhuoyu, et al.
Pubblicazione: (2026)
di: Wang, Zhuoyu, et al.
Pubblicazione: (2026)
DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution
di: Hu, Yunhai, et al.
Pubblicazione: (2026)
di: Hu, Yunhai, et al.
Pubblicazione: (2026)
HilbertA: Hilbert Attention for Image Generation with Diffusion Models
di: Zheng, Shaoyi, et al.
Pubblicazione: (2025)
di: Zheng, Shaoyi, et al.
Pubblicazione: (2025)
ToMA: Token Merge with Attention for Diffusion Models
di: Lu, Wenbo, et al.
Pubblicazione: (2025)
di: Lu, Wenbo, et al.
Pubblicazione: (2025)
VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping
di: Dong, Haotian, et al.
Pubblicazione: (2025)
di: Dong, Haotian, et al.
Pubblicazione: (2025)
Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs
di: Athiwaratkun, Ben, et al.
Pubblicazione: (2024)
di: Athiwaratkun, Ben, et al.
Pubblicazione: (2024)
PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention
di: Wang, Haonan, et al.
Pubblicazione: (2025)
di: Wang, Haonan, et al.
Pubblicazione: (2025)
Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
di: Liu, Yuhan, et al.
Pubblicazione: (2025)
di: Liu, Yuhan, et al.
Pubblicazione: (2025)
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
di: Wu, Tianyu, et al.
Pubblicazione: (2026)
di: Wu, Tianyu, et al.
Pubblicazione: (2026)
Traversal Verification for Speculative Tree Decoding
di: Weng, Yepeng, et al.
Pubblicazione: (2025)
di: Weng, Yepeng, et al.
Pubblicazione: (2025)
PACER: Blockwise Pre-verification for Speculative Decoding with Adaptive Length
di: Zhang, Situo, et al.
Pubblicazione: (2026)
di: Zhang, Situo, et al.
Pubblicazione: (2026)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
di: Shen, Yuhao, et al.
Pubblicazione: (2025)
di: Shen, Yuhao, et al.
Pubblicazione: (2025)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
di: Zhang, Ziyin, et al.
Pubblicazione: (2024)
di: Zhang, Ziyin, et al.
Pubblicazione: (2024)
Hydra: Efficient, Correct Code Generation via Checkpoint-and-Rollback Support
di: Du, Alexander, et al.
Pubblicazione: (2026)
di: Du, Alexander, et al.
Pubblicazione: (2026)
PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
di: Ding, Ruogu, et al.
Pubblicazione: (2025)
di: Ding, Ruogu, et al.
Pubblicazione: (2025)
PrefixLLM: LLM-aided Prefix Circuit Design
di: Xiao, Weihua, et al.
Pubblicazione: (2024)
di: Xiao, Weihua, et al.
Pubblicazione: (2024)
WISV: Wireless-Informed Semantic Verification for Distributed Speculative Decoding in Device-Edge LLM Inference
di: Liu, Zixuan, et al.
Pubblicazione: (2026)
di: Liu, Zixuan, et al.
Pubblicazione: (2026)
Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding
di: Su, Xin, et al.
Pubblicazione: (2026)
di: Su, Xin, et al.
Pubblicazione: (2026)
Overcoming Joint Intractability with Lossless Hierarchical Speculative Decoding
di: Zhou, Yuxuan, et al.
Pubblicazione: (2026)
di: Zhou, Yuxuan, et al.
Pubblicazione: (2026)
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
di: Liu, Zikang, et al.
Pubblicazione: (2025)
di: Liu, Zikang, et al.
Pubblicazione: (2025)
First Ask Then Answer: A Framework Design for AI Dialogue Based on Supplementary Questioning with Large Language Models
di: Fu, Chuanruo, et al.
Pubblicazione: (2025)
di: Fu, Chuanruo, et al.
Pubblicazione: (2025)
Hypothesize-Then-Verify: Speculative Root Cause Analysis for Microservices with Pathwise Parallelism
di: Zhang, Lingzhe, et al.
Pubblicazione: (2026)
di: Zhang, Lingzhe, et al.
Pubblicazione: (2026)
PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization
di: Zuo, Dongsheng, et al.
Pubblicazione: (2025)
di: Zuo, Dongsheng, et al.
Pubblicazione: (2025)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
di: Yang, Penghui, et al.
Pubblicazione: (2025)
di: Yang, Penghui, et al.
Pubblicazione: (2025)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
di: Yoon, Kanghoon, et al.
Pubblicazione: (2025)
di: Yoon, Kanghoon, et al.
Pubblicazione: (2025)
ConfSpec: Efficient Step-Level Speculative Reasoning via Confidence-Gated Verification
di: Liu, Siran, et al.
Pubblicazione: (2026)
di: Liu, Siran, et al.
Pubblicazione: (2026)
Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference
di: Timor, Nadav, et al.
Pubblicazione: (2024)
di: Timor, Nadav, et al.
Pubblicazione: (2024)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
di: An, Zihao, et al.
Pubblicazione: (2026)
di: An, Zihao, et al.
Pubblicazione: (2026)
DIVERSED: Relaxed Speculative Decoding via Dynamic Ensemble Verification
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
Enhancing High-Quality Code Generation in Large Language Models with Comparative Prefix-Tuning
di: Jiang, Yuan, et al.
Pubblicazione: (2025)
di: Jiang, Yuan, et al.
Pubblicazione: (2025)
SAM Decoding: Speculative Decoding via Suffix Automaton
di: Hu, Yuxuan, et al.
Pubblicazione: (2024)
di: Hu, Yuxuan, et al.
Pubblicazione: (2024)
Human-Guided Image Generation for Expanding Small-Scale Training Image Datasets
di: Chen, Changjian, et al.
Pubblicazione: (2024)
di: Chen, Changjian, et al.
Pubblicazione: (2024)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
di: Li, Ruanjun, et al.
Pubblicazione: (2025)
di: Li, Ruanjun, et al.
Pubblicazione: (2025)
Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
di: Wan, Yuxuan, et al.
Pubblicazione: (2026)
di: Wan, Yuxuan, et al.
Pubblicazione: (2026)
Layered LA-MAPF: a decomposition of large agent MAPF instance to accelerate solving without compromising solvability
di: Yao, Zhuo
Pubblicazione: (2024)
di: Yao, Zhuo
Pubblicazione: (2024)
Generating Visual Stories with Grounded and Coreferent Characters
di: Liu, Danyang, et al.
Pubblicazione: (2024)
di: Liu, Danyang, et al.
Pubblicazione: (2024)
Accelerating Large Language Model Reasoning via Speculative Search
di: Wang, Zhihai, et al.
Pubblicazione: (2025)
di: Wang, Zhihai, et al.
Pubblicazione: (2025)
Speculative Decoding for Multi-Sample Inference
di: Li, Yiwei, et al.
Pubblicazione: (2025)
di: Li, Yiwei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training
di: Wang, Xi, et al.
Pubblicazione: (2026) -
TAPS: Target-Aware Prefix Tree Selection for Diffusion-Drafted Speculative Decoding
di: Wang, Zhuoyu, et al.
Pubblicazione: (2026) -
DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution
di: Hu, Yunhai, et al.
Pubblicazione: (2026) -
HilbertA: Hilbert Attention for Image Generation with Diffusion Models
di: Zheng, Shaoyi, et al.
Pubblicazione: (2025) -
ToMA: Token Merge with Attention for Diffusion Models
di: Lu, Wenbo, et al.
Pubblicazione: (2025)