Optimized Speculative Sampling for GPU Hardware Accelerators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wagner, Dominik, Lee, Seanie, Baumann, Ilja, Seeberger, Philipp, Riedhammer, Korbinian, Bocklet, Tobias |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generalizing to Unseen Disaster Events: A Causal View
von: Seeberger, Philipp, et al.
Veröffentlicht: (2025)
von: Seeberger, Philipp, et al.
Veröffentlicht: (2025)
MMUTF: Multimodal Multimedia Event Argument Extraction with Unified Template Filling
von: Seeberger, Philipp, et al.
Veröffentlicht: (2024)
von: Seeberger, Philipp, et al.
Veröffentlicht: (2024)
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
von: Baumann, Ilja, et al.
Veröffentlicht: (2025)
von: Baumann, Ilja, et al.
Veröffentlicht: (2025)
Reading Between the Waves: Robust Topic Segmentation Using Inter-Sentence Audio Features
von: Freisinger, Steffen, et al.
Veröffentlicht: (2026)
von: Freisinger, Steffen, et al.
Veröffentlicht: (2026)
Multi-Query Focused Disaster Summarization via Instruction-Based Prompting
von: Seeberger, Philipp, et al.
Veröffentlicht: (2024)
von: Seeberger, Philipp, et al.
Veröffentlicht: (2024)
Large Language Models for Dysfluency Detection in Stuttered Speech
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation
von: Freisinger, Steffen, et al.
Veröffentlicht: (2026)
von: Freisinger, Steffen, et al.
Veröffentlicht: (2026)
On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
von: Gulzar, Kashaf, et al.
Veröffentlicht: (2025)
von: Gulzar, Kashaf, et al.
Veröffentlicht: (2025)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models
von: Simic, Christopher, et al.
Veröffentlicht: (2025)
von: Simic, Christopher, et al.
Veröffentlicht: (2025)
Shared Multi-modal Embedding Space for Face-Voice Association
von: Simic, Christopher, et al.
Veröffentlicht: (2025)
von: Simic, Christopher, et al.
Veröffentlicht: (2025)
Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0
von: Engert, Natalie, et al.
Veröffentlicht: (2026)
von: Engert, Natalie, et al.
Veröffentlicht: (2026)
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
von: Wagner, Dominik, et al.
Veröffentlicht: (2023)
von: Wagner, Dominik, et al.
Veröffentlicht: (2023)
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
von: Lee, Seanie, et al.
Veröffentlicht: (2024)
von: Lee, Seanie, et al.
Veröffentlicht: (2024)
Digital Operating Mode Classification of Real-World Amateur Radio Transmissions
von: Bundscherer, Maximilian, et al.
Veröffentlicht: (2025)
von: Bundscherer, Maximilian, et al.
Veröffentlicht: (2025)
SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models
von: Lee, Seanie, et al.
Veröffentlicht: (2025)
von: Lee, Seanie, et al.
Veröffentlicht: (2025)
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
von: Bayerl, Sebastian P., et al.
Veröffentlicht: (2022)
von: Bayerl, Sebastian P., et al.
Veröffentlicht: (2022)
ConFu: Contemplate the Future for Better Speculative Sampling
von: Qin, Zongyue, et al.
Veröffentlicht: (2026)
von: Qin, Zongyue, et al.
Veröffentlicht: (2026)
Learning Harmonized Representations for Speculative Sampling
von: Zhang, Lefan, et al.
Veröffentlicht: (2024)
von: Zhang, Lefan, et al.
Veröffentlicht: (2024)
BASS: Batched Attention-optimized Speculative Sampling
von: Qian, Haifeng, et al.
Veröffentlicht: (2024)
von: Qian, Haifeng, et al.
Veröffentlicht: (2024)
Out-of-Vocabulary Sampling Boosts Speculative Decoding
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
FedSVD: Adaptive Orthogonalization for Private Federated Learning with LoRA
von: Lee, Seanie, et al.
Veröffentlicht: (2025)
von: Lee, Seanie, et al.
Veröffentlicht: (2025)
Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
von: Sun, Shuoyang, et al.
Veröffentlicht: (2026)
von: Sun, Shuoyang, et al.
Veröffentlicht: (2026)
Infusing Acoustic Pause Context into Text-Based Dementia Assessment
von: Braun, Franziska, et al.
Veröffentlicht: (2024)
von: Braun, Franziska, et al.
Veröffentlicht: (2024)
EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
von: Li, Yuhui, et al.
Veröffentlicht: (2024)
von: Li, Yuhui, et al.
Veröffentlicht: (2024)
Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion
von: Christopher, Jacob K, et al.
Veröffentlicht: (2024)
von: Christopher, Jacob K, et al.
Veröffentlicht: (2024)
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
von: Iso, Hayate, et al.
Veröffentlicht: (2026)
von: Iso, Hayate, et al.
Veröffentlicht: (2026)
SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
Accelerating Retrieval-Augmented Language Model Serving with Speculation
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
Boosting Lossless Speculative Decoding via Feature Sampling and Partial Alignment Distillation
von: Gui, Lujun, et al.
Veröffentlicht: (2024)
von: Gui, Lujun, et al.
Veröffentlicht: (2024)
Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment
von: Bachmann, Gregor, et al.
Veröffentlicht: (2025)
von: Bachmann, Gregor, et al.
Veröffentlicht: (2025)
Reviving Any-Subset Autoregressive Models with Principled Parallel Sampling and Speculative Decoding
von: Guo, Gabe, et al.
Veröffentlicht: (2025)
von: Guo, Gabe, et al.
Veröffentlicht: (2025)
LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding
von: Samarin, Alexander, et al.
Veröffentlicht: (2026)
von: Samarin, Alexander, et al.
Veröffentlicht: (2026)
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
von: Liu, Bingshuai, et al.
Veröffentlicht: (2025)
von: Liu, Bingshuai, et al.
Veröffentlicht: (2025)
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
von: McDanel, Bradley
Veröffentlicht: (2024)
von: McDanel, Bradley
Veröffentlicht: (2024)
Ähnliche Einträge
-
Generalizing to Unseen Disaster Events: A Causal View
von: Seeberger, Philipp, et al.
Veröffentlicht: (2025) -
MMUTF: Multimodal Multimedia Event Argument Extraction with Unified Template Filling
von: Seeberger, Philipp, et al.
Veröffentlicht: (2024) -
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
von: Baumann, Ilja, et al.
Veröffentlicht: (2025) -
Reading Between the Waves: Robust Topic Segmentation Using Inter-Sentence Audio Features
von: Freisinger, Steffen, et al.
Veröffentlicht: (2026) -
Multi-Query Focused Disaster Summarization via Instruction-Based Prompting
von: Seeberger, Philipp, et al.
Veröffentlicht: (2024)