Efficient Neural Compression with Inference-time Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Metz, C., Bichler, O., Dupret, A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ultra-Efficient Decoding for End-to-End Neural Compression and Reconstruction
von: Rogers, Ethan G., et al.
Veröffentlicht: (2025)
von: Rogers, Ethan G., et al.
Veröffentlicht: (2025)
NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks
von: Hao, Yongchang, et al.
Veröffentlicht: (2024)
von: Hao, Yongchang, et al.
Veröffentlicht: (2024)
Inference-friendly Graph Compression for Graph Neural Networks
von: Fan, Yangxin, et al.
Veröffentlicht: (2025)
von: Fan, Yangxin, et al.
Veröffentlicht: (2025)
Progressive Mixed-Precision Decoding for Efficient LLM Inference
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
Efficient Decoder Scaling Strategy for Neural Routing Solvers
von: Luo, Qing, et al.
Veröffentlicht: (2026)
von: Luo, Qing, et al.
Veröffentlicht: (2026)
Efficient Test-Time Inference via Deterministic Exploration of Truncated Decoding Trees
von: Li, Xueyan, et al.
Veröffentlicht: (2026)
von: Li, Xueyan, et al.
Veröffentlicht: (2026)
Efficient Model Compression for Bayesian Neural Networks
von: Saha, Diptarka, et al.
Veröffentlicht: (2024)
von: Saha, Diptarka, et al.
Veröffentlicht: (2024)
Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
von: Arriola, Marianne, et al.
Veröffentlicht: (2025)
von: Arriola, Marianne, et al.
Veröffentlicht: (2025)
INEUS: Iterative Neural Solver for High-Dimensional PIDEs
von: Dupret, Jean-Loup, et al.
Veröffentlicht: (2026)
von: Dupret, Jean-Loup, et al.
Veröffentlicht: (2026)
Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference
von: Zhou, Xuwen, et al.
Veröffentlicht: (2026)
von: Zhou, Xuwen, et al.
Veröffentlicht: (2026)
Neural Embedding Compression For Efficient Multi-Task Earth Observation Modelling
von: Gomes, Carlos, et al.
Veröffentlicht: (2024)
von: Gomes, Carlos, et al.
Veröffentlicht: (2024)
Score $\times$ Decoder: A Unified View of Unsupervised Inference-Time Scaling for Hallucination Mitigation
von: Cheng, Yun-Chen, et al.
Veröffentlicht: (2026)
von: Cheng, Yun-Chen, et al.
Veröffentlicht: (2026)
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
ResponseRank: Data-Efficient Reward Modeling through Preference Strength Learning
von: Kaufmann, Timo, et al.
Veröffentlicht: (2025)
von: Kaufmann, Timo, et al.
Veröffentlicht: (2025)
Entropy Adaptive Decoding: Dynamic Model Switching for Efficient Inference
von: Simonds, Toby
Veröffentlicht: (2025)
von: Simonds, Toby
Veröffentlicht: (2025)
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2025)
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2025)
GraphPI: Efficient Protein Inference with Graph Neural Networks
von: Ma, Zheng, et al.
Veröffentlicht: (2026)
von: Ma, Zheng, et al.
Veröffentlicht: (2026)
Early-Exit with Class Exclusion for Efficient Inference of Neural Networks
von: Wang, Jingcun, et al.
Veröffentlicht: (2023)
von: Wang, Jingcun, et al.
Veröffentlicht: (2023)
Fast Inference via Hierarchical Speculative Decoding
von: Mohri, Clara, et al.
Veröffentlicht: (2025)
von: Mohri, Clara, et al.
Veröffentlicht: (2025)
Nudging: Inference-time Alignment of LLMs via Guided Decoding
von: Fei, Yu, et al.
Veröffentlicht: (2024)
von: Fei, Yu, et al.
Veröffentlicht: (2024)
From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models
von: Welleck, Sean, et al.
Veröffentlicht: (2024)
von: Welleck, Sean, et al.
Veröffentlicht: (2024)
An Efficient Compression of Deep Neural Network Checkpoints Based on Prediction and Context Modeling
von: Kim, Yuriy, et al.
Veröffentlicht: (2025)
von: Kim, Yuriy, et al.
Veröffentlicht: (2025)
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
A Scalable, Causal, and Energy Efficient Framework for Neural Decoding with Spiking Neural Networks
von: Mentzelopoulos, Georgios, et al.
Veröffentlicht: (2025)
von: Mentzelopoulos, Georgios, et al.
Veröffentlicht: (2025)
EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices
von: Sanyal, Arnab, et al.
Veröffentlicht: (2025)
von: Sanyal, Arnab, et al.
Veröffentlicht: (2025)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
von: Park, Jihoon, et al.
Veröffentlicht: (2025)
von: Park, Jihoon, et al.
Veröffentlicht: (2025)
DepCap: Adaptive Block-Wise Parallel Decoding for Efficient Diffusion LM Inference
von: Xia, Xiang, et al.
Veröffentlicht: (2026)
von: Xia, Xiang, et al.
Veröffentlicht: (2026)
BlockBatch: Multi-Scale Consensus Decoding for Efficient Diffusion Language Model Inference
von: Wu, Xiaoyou, et al.
Veröffentlicht: (2026)
von: Wu, Xiaoyou, et al.
Veröffentlicht: (2026)
NeuraLUT-Assemble: Hardware-aware Assembling of Sub-Neural Networks for Efficient LUT Inference
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding
von: Ryu, Hyun, et al.
Veröffentlicht: (2024)
von: Ryu, Hyun, et al.
Veröffentlicht: (2024)
SPIRe: Boosting LLM Inference Throughput with Speculative Decoding
von: Neelam, Sanjit, et al.
Veröffentlicht: (2025)
von: Neelam, Sanjit, et al.
Veröffentlicht: (2025)
Neural Network Compression for Reinforcement Learning Tasks
von: Ivanov, Dmitry A., et al.
Veröffentlicht: (2024)
von: Ivanov, Dmitry A., et al.
Veröffentlicht: (2024)
List-Level Distribution Coupling with Applications to Speculative Decoding and Lossy Compression
von: Rowan, Joseph, et al.
Veröffentlicht: (2025)
von: Rowan, Joseph, et al.
Veröffentlicht: (2025)
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
von: Dong, Harry, et al.
Veröffentlicht: (2024)
von: Dong, Harry, et al.
Veröffentlicht: (2024)
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
von: Xiang, Maoyang, et al.
Veröffentlicht: (2025)
von: Xiang, Maoyang, et al.
Veröffentlicht: (2025)
Neural Normalized Compression Distance and the Disconnect Between Compression and Classification
von: Hurwitz, John, et al.
Veröffentlicht: (2024)
von: Hurwitz, John, et al.
Veröffentlicht: (2024)
Gradual Binary Search and Dimension Expansion : A general method for activation quantization in LLMs
von: Maisonnave, Lucas, et al.
Veröffentlicht: (2025)
von: Maisonnave, Lucas, et al.
Veröffentlicht: (2025)
Adaptive Error-Bounded Hierarchical Matrices for Efficient Neural Network Compression
von: Mango, John, et al.
Veröffentlicht: (2024)
von: Mango, John, et al.
Veröffentlicht: (2024)
Set Block Decoding is a Language Model Inference Accelerator
von: Gat, Itai, et al.
Veröffentlicht: (2025)
von: Gat, Itai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ultra-Efficient Decoding for End-to-End Neural Compression and Reconstruction
von: Rogers, Ethan G., et al.
Veröffentlicht: (2025) -
NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks
von: Hao, Yongchang, et al.
Veröffentlicht: (2024) -
Inference-friendly Graph Compression for Graph Neural Networks
von: Fan, Yangxin, et al.
Veröffentlicht: (2025) -
Progressive Mixed-Precision Decoding for Efficient LLM Inference
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024) -
Efficient Decoder Scaling Strategy for Neural Routing Solvers
von: Luo, Qing, et al.
Veröffentlicht: (2026)