Accelerating Inference of Discrete Autoregressive Normalizing Flows by Selective Jacobi Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jiaru, Lu, Juanwu, Wu, Xiaoyu, Wang, Ziran, Zhang, Ruqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On Variance Reduction in Learning Mean Flows
von: Lu, Juanwu, et al.
Veröffentlicht: (2026)
von: Lu, Juanwu, et al.
Veröffentlicht: (2026)
Analytical Correction for Subsampling Bias in Drifting Models
von: Zhang, Jiaru, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaru, et al.
Veröffentlicht: (2026)
One-Step Diffusion Samplers via Self-Distillation and Deterministic Flow
von: Jutras-Dube, Pascal, et al.
Veröffentlicht: (2025)
von: Jutras-Dube, Pascal, et al.
Veröffentlicht: (2025)
Latent Bayesian Optimization via Autoregressive Normalizing Flows
von: Lee, Seunghun, et al.
Veröffentlicht: (2025)
von: Lee, Seunghun, et al.
Veröffentlicht: (2025)
Leveraging Model Guidance to Extract Training Data from Personalized Diffusion Models
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2024)
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
von: Ning, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Ning, Zhiyuan, et al.
Veröffentlicht: (2025)
FedMomentum: Preserving LoRA Training Momentum in Federated Fine-Tuning
von: Yan, Peishen, et al.
Veröffentlicht: (2026)
von: Yan, Peishen, et al.
Veröffentlicht: (2026)
Model Merging on Loss Landscape: A Geometry Perspective
von: Lu, Juanwu, et al.
Veröffentlicht: (2026)
von: Lu, Juanwu, et al.
Veröffentlicht: (2026)
Distilling Autoregressive Models to Obtain High-Performance Non-Autoregressive Solvers for Vehicle Routing Problems with Faster Inference Speed
von: Xiao, Yubin, et al.
Veröffentlicht: (2023)
von: Xiao, Yubin, et al.
Veröffentlicht: (2023)
DIVERSED: Relaxed Speculative Decoding via Dynamic Ensemble Verification
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
A Universal and Robust Framework for Multiple Gas Recognition Based-on Spherical Normalization-Coupled Mahalanobis Algorithm
von: Chen, Shuai, et al.
Veröffentlicht: (2025)
von: Chen, Shuai, et al.
Veröffentlicht: (2025)
Autoregressive Visual Decoding from EEG Signals
von: Dai, Sicheng, et al.
Veröffentlicht: (2026)
von: Dai, Sicheng, et al.
Veröffentlicht: (2026)
Learning Discrete Autoregressive Priors with Wasserstein Gradient Flow
von: Zheng, Bowen, et al.
Veröffentlicht: (2026)
von: Zheng, Bowen, et al.
Veröffentlicht: (2026)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
Accelerating Transformer Inference for Translation via Parallel Decoding
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
PromptIntern: Saving Inference Costs by Internalizing Recurrent Prompt during Large Language Model Fine-tuning
von: Zou, Jiaru, et al.
Veröffentlicht: (2024)
von: Zou, Jiaru, et al.
Veröffentlicht: (2024)
Latent-DARM: Bridging Discrete Diffusion And Autoregressive Models For Reasoning
von: Berrayana, Lina, et al.
Veröffentlicht: (2026)
von: Berrayana, Lina, et al.
Veröffentlicht: (2026)
A Theoretical Analysis of Discrete Flow Matching Generative Models
von: Su, Maojiang, et al.
Veröffentlicht: (2025)
von: Su, Maojiang, et al.
Veröffentlicht: (2025)
Discrete Flow Matching
von: Gat, Itai, et al.
Veröffentlicht: (2024)
von: Gat, Itai, et al.
Veröffentlicht: (2024)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
DepCap: Adaptive Block-Wise Parallel Decoding for Efficient Diffusion LM Inference
von: Xia, Xiang, et al.
Veröffentlicht: (2026)
von: Xia, Xiang, et al.
Veröffentlicht: (2026)
Linear Transformers as VAR Models: Aligning Autoregressive Attention Mechanisms with Autoregressive Forecasting
von: Lu, Jiecheng, et al.
Veröffentlicht: (2025)
von: Lu, Jiecheng, et al.
Veröffentlicht: (2025)
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2025)
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2025)
PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization
von: Wang, Xinhai, et al.
Veröffentlicht: (2025)
von: Wang, Xinhai, et al.
Veröffentlicht: (2025)
Controlled LLM Decoding via Discrete Auto-regressive Biasing
von: Pynadath, Patrick, et al.
Veröffentlicht: (2025)
von: Pynadath, Patrick, et al.
Veröffentlicht: (2025)
Attention Drift: What Autoregressive Speculative Decoding Models Learn
von: Eldenk, Doğaç, et al.
Veröffentlicht: (2026)
von: Eldenk, Doğaç, et al.
Veröffentlicht: (2026)
From Values to Tokens: An LLM-Driven Framework for Context-aware Time Series Forecasting via Symbolic Discretization
von: Tao, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Tao, Xiaoyu, et al.
Veröffentlicht: (2025)
ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
Efficient Regression-Based Training of Normalizing Flows for Boltzmann Generators
von: Rehman, Danyal, et al.
Veröffentlicht: (2025)
von: Rehman, Danyal, et al.
Veröffentlicht: (2025)
Exploring Diffusion Models' Corruption Stage in Few-Shot Fine-tuning and Mitigating with Bayesian Neural Networks
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2024)
ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
von: Li, Jia-Nan, et al.
Veröffentlicht: (2025)
von: Li, Jia-Nan, et al.
Veröffentlicht: (2025)
Amortized Sampling with Transferable Normalizing Flows
von: Tan, Charlie B., et al.
Veröffentlicht: (2025)
von: Tan, Charlie B., et al.
Veröffentlicht: (2025)
Knowledge Graph Embedding by Normalizing Flows
von: Xiao, Changyi, et al.
Veröffentlicht: (2024)
von: Xiao, Changyi, et al.
Veröffentlicht: (2024)
Personalized Federated Learning via Learning Dynamic Graphs
von: Zhou, Ziran, et al.
Veröffentlicht: (2025)
von: Zhou, Ziran, et al.
Veröffentlicht: (2025)
Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation
von: Li, Xingyao, et al.
Veröffentlicht: (2026)
von: Li, Xingyao, et al.
Veröffentlicht: (2026)
FlowPG: Action-constrained Policy Gradient with Normalizing Flows
von: Brahmanage, Janaka Chathuranga, et al.
Veröffentlicht: (2024)
von: Brahmanage, Janaka Chathuranga, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On Variance Reduction in Learning Mean Flows
von: Lu, Juanwu, et al.
Veröffentlicht: (2026) -
Analytical Correction for Subsampling Bias in Drifting Models
von: Zhang, Jiaru, et al.
Veröffentlicht: (2026) -
One-Step Diffusion Samplers via Self-Distillation and Deterministic Flow
von: Jutras-Dube, Pascal, et al.
Veröffentlicht: (2025) -
Latent Bayesian Optimization via Autoregressive Normalizing Flows
von: Lee, Seunghun, et al.
Veröffentlicht: (2025) -
Leveraging Model Guidance to Extract Training Data from Personalized Diffusion Models
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2024)