DynaMo: Accelerating Language Model Inference with Dynamic Multi-Token Sampling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tuli, Shikhar, Lin, Chi-Heng, Hsu, Yen-Chang, Jha, Niraj K., Shen, Yilin, Jin, Hongxia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MoDeGPT: Modular Decomposition for Large Language Model Compression
von: Lin, Chi-Heng, et al.
Veröffentlicht: (2024)
von: Lin, Chi-Heng, et al.
Veröffentlicht: (2024)
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing
von: Smith, James Seale, et al.
Veröffentlicht: (2025)
von: Smith, James Seale, et al.
Veröffentlicht: (2025)
DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control
von: Cui, Zichen Jeff, et al.
Veröffentlicht: (2024)
von: Cui, Zichen Jeff, et al.
Veröffentlicht: (2024)
MossNet: Mixture of State-Space Experts is a Multi-Head Attention
von: Tuli, Shikhar, et al.
Veröffentlicht: (2025)
von: Tuli, Shikhar, et al.
Veröffentlicht: (2025)
GraphMERT: Efficient and Scalable Distillation of Reliable Knowledge Graphs from Unstructured Data
von: Belova, Margarita, et al.
Veröffentlicht: (2025)
von: Belova, Margarita, et al.
Veröffentlicht: (2025)
DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
DISP-LLM: Dimension-Independent Structural Pruning for Large Language Models
von: Gao, Shangqian, et al.
Veröffentlicht: (2024)
von: Gao, Shangqian, et al.
Veröffentlicht: (2024)
Continual Diffusion with STAMINA: STack-And-Mask INcremental Adapters
von: Smith, James Seale, et al.
Veröffentlicht: (2023)
von: Smith, James Seale, et al.
Veröffentlicht: (2023)
Learning Interpretable Differentiable Logic Networks for Time-Series Classification
von: Yue, Chang, et al.
Veröffentlicht: (2025)
von: Yue, Chang, et al.
Veröffentlicht: (2025)
Learning Interpretable Differentiable Logic Networks
von: Yue, Chang, et al.
Veröffentlicht: (2024)
von: Yue, Chang, et al.
Veröffentlicht: (2024)
Learning Interpretable Differentiable Logic Networks for Tabular Regression
von: Yue, Chang, et al.
Veröffentlicht: (2025)
von: Yue, Chang, et al.
Veröffentlicht: (2025)
TAD-SIE: Sample Size Estimation for Clinical Randomized Controlled Trials using a Trend-Adaptive Design with a Synthetic-Intervention-Based Estimator
von: Lala, Sayeri, et al.
Veröffentlicht: (2024)
von: Lala, Sayeri, et al.
Veröffentlicht: (2024)
Uncertainty-Aware Transformers: Conformal Prediction for Language Models
von: Vellore, Abhiram, et al.
Veröffentlicht: (2026)
von: Vellore, Abhiram, et al.
Veröffentlicht: (2026)
Continual Diffusion: Continual Customization of Text-to-Image Diffusion with C-LoRA
von: Smith, James Seale, et al.
Veröffentlicht: (2023)
von: Smith, James Seale, et al.
Veröffentlicht: (2023)
Zero-TPrune: Zero-Shot Token Pruning through Leveraging of the Attention Graph in Pre-Trained Transformers
von: Wang, Hongjie, et al.
Veröffentlicht: (2023)
von: Wang, Hongjie, et al.
Veröffentlicht: (2023)
HEED: Density-Weighted Residual Alignment for Hybrid Vision-Language Model Distillation
von: Liang, Yihao, et al.
Veröffentlicht: (2026)
von: Liang, Yihao, et al.
Veröffentlicht: (2026)
DOCTOR: A Multi-Disease Detection Continual Learning Framework Based on Wearable Medical Sensors
von: Li, Chia-Hao, et al.
Veröffentlicht: (2023)
von: Li, Chia-Hao, et al.
Veröffentlicht: (2023)
ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
von: Gao, Shangqian, et al.
Veröffentlicht: (2025)
von: Gao, Shangqian, et al.
Veröffentlicht: (2025)
METRIK: Measurement-Efficient Randomized Controlled Trials using Transformers with Input Masking
von: Lala, Sayeri, et al.
Veröffentlicht: (2024)
von: Lala, Sayeri, et al.
Veröffentlicht: (2024)
Knowledge Graph-Driven Expert-Level Reasoning for Neuroscience
von: Stephen, Jake, et al.
Veröffentlicht: (2026)
von: Stephen, Jake, et al.
Veröffentlicht: (2026)
Neural Slot Interpreters: Grounding Object Semantics in Emergent Slot Representations
von: Dedhia, Bhishma, et al.
Veröffentlicht: (2024)
von: Dedhia, Bhishma, et al.
Veröffentlicht: (2024)
LinMU: Multimodal Understanding Made Linear
von: Wang, Hongjie, et al.
Veröffentlicht: (2026)
von: Wang, Hongjie, et al.
Veröffentlicht: (2026)
Knowledge Graphs are Implicit Reward Models: Path-Derived Signals Enable Compositional Reasoning
von: Kansal, Yuval, et al.
Veröffentlicht: (2026)
von: Kansal, Yuval, et al.
Veröffentlicht: (2026)
MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
von: Kallini, Julie, et al.
Veröffentlicht: (2024)
von: Kallini, Julie, et al.
Veröffentlicht: (2024)
COMFORT: A Continual Fine-Tuning Framework for Foundation Models Targeted at Consumer Healthcare
von: Li, Chia-Hao, et al.
Veröffentlicht: (2024)
von: Li, Chia-Hao, et al.
Veröffentlicht: (2024)
PAGE: Domain-Incremental Adaptation with Past-Agnostic Generative Replay for Smart Healthcare
von: Li, Chia-Hao, et al.
Veröffentlicht: (2024)
von: Li, Chia-Hao, et al.
Veröffentlicht: (2024)
FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
DynaMoE: Dynamic Token-Level Expert Activation with Layer-Wise Adaptive Capacity for Mixture-of-Experts Neural Networks
von: Gülmez, Gökdeniz
Veröffentlicht: (2026)
von: Gülmez, Gökdeniz
Veröffentlicht: (2026)
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
von: Chen, Liang, et al.
Veröffentlicht: (2024)
von: Chen, Liang, et al.
Veröffentlicht: (2024)
AsymVLM: Asymmetric Token Pruning for Efficient Vision-Language Model Inference
von: Feng, Yilin, et al.
Veröffentlicht: (2026)
von: Feng, Yilin, et al.
Veröffentlicht: (2026)
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
von: Lin, Wenxiang, et al.
Veröffentlicht: (2025)
von: Lin, Wenxiang, et al.
Veröffentlicht: (2025)
DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
von: Zhang, Jinbin, et al.
Veröffentlicht: (2025)
von: Zhang, Jinbin, et al.
Veröffentlicht: (2025)
Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
von: Dedhia, Bhishma, et al.
Veröffentlicht: (2025)
von: Dedhia, Bhishma, et al.
Veröffentlicht: (2025)
CONFINE: Conformal Prediction for Interpretable Neural Networks
von: Huang, Linhui, et al.
Veröffentlicht: (2024)
von: Huang, Linhui, et al.
Veröffentlicht: (2024)
NI Sampling: Accelerating Discrete Diffusion Sampling by Token Order Optimization
von: Liu, Enshu, et al.
Veröffentlicht: (2026)
von: Liu, Enshu, et al.
Veröffentlicht: (2026)
DEI: Diversity in Evolutionary Inference for Quality-Diversity Search
von: Donaghy, John, et al.
Veröffentlicht: (2026)
von: Donaghy, John, et al.
Veröffentlicht: (2026)
Retraining-Free Merging of Sparse MoE via Hierarchical Clustering
von: Chen, I-Chun, et al.
Veröffentlicht: (2024)
von: Chen, I-Chun, et al.
Veröffentlicht: (2024)
DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration
von: Zhang, Hanzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Hanzhi, et al.
Veröffentlicht: (2025)
A Frequentist Statistical Introduction to Variational Inference, Autoencoders, and Diffusion Models
von: Chen, Yen-Chi
Veröffentlicht: (2025)
von: Chen, Yen-Chi
Veröffentlicht: (2025)
On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation
von: Hsu, Chan-Jan, et al.
Veröffentlicht: (2026)
von: Hsu, Chan-Jan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MoDeGPT: Modular Decomposition for Large Language Model Compression
von: Lin, Chi-Heng, et al.
Veröffentlicht: (2024) -
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing
von: Smith, James Seale, et al.
Veröffentlicht: (2025) -
DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control
von: Cui, Zichen Jeff, et al.
Veröffentlicht: (2024) -
MossNet: Mixture of State-Space Experts is a Multi-Head Attention
von: Tuli, Shikhar, et al.
Veröffentlicht: (2025) -
GraphMERT: Efficient and Scalable Distillation of Reliable Knowledge Graphs from Unstructured Data
von: Belova, Margarita, et al.
Veröffentlicht: (2025)