Efficient Generative Modeling with Residual Vector Quantization-Based Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jaehyeon, Moon, Taehong, Lee, Keon, Cho, Jaewoong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
by: Park, Dongmin, et al.
Published: (2024)
by: Park, Dongmin, et al.
Published: (2024)
DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
by: Lee, Keon, et al.
Published: (2024)
by: Lee, Keon, et al.
Published: (2024)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
by: Kim, Jaehyeon, et al.
Published: (2024)
by: Kim, Jaehyeon, et al.
Published: (2024)
Instance-Aware Group Quantization for Vision Transformers
by: Moon, Jaehyeon, et al.
Published: (2024)
by: Moon, Jaehyeon, et al.
Published: (2024)
Task Vector Quantization for Memory-Efficient Model Merging
by: Kim, Youngeun, et al.
Published: (2025)
by: Kim, Youngeun, et al.
Published: (2025)
Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning Models
by: Kim, Junhyuck, et al.
Published: (2025)
by: Kim, Junhyuck, et al.
Published: (2025)
Distilling Reinforcement Learning Algorithms for In-Context Model-Based Planning
by: Son, Jaehyeon, et al.
Published: (2025)
by: Son, Jaehyeon, et al.
Published: (2025)
Concept-Centric Token Interpretation for Vector-Quantized Generative Models
by: Yang, Tianze, et al.
Published: (2025)
by: Yang, Tianze, et al.
Published: (2025)
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
by: Lee, Jungi, et al.
Published: (2025)
by: Lee, Jungi, et al.
Published: (2025)
Recasting Continual Learning as Sequence Modeling
by: Lee, Soochan, et al.
Published: (2023)
by: Lee, Soochan, et al.
Published: (2023)
Variable Bitrate Residual Vector Quantization for Audio Coding
by: Chae, Yunkee, et al.
Published: (2024)
by: Chae, Yunkee, et al.
Published: (2024)
Fast and Accurate Neural Rendering Using Semi-Gradients
by: Cho, In-Young, et al.
Published: (2024)
by: Cho, In-Young, et al.
Published: (2024)
InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management
by: Lee, Wonbeom, et al.
Published: (2024)
by: Lee, Wonbeom, et al.
Published: (2024)
When Meta-Learning Meets Online and Continual Learning: A Survey
by: Son, Jaehyeon, et al.
Published: (2023)
by: Son, Jaehyeon, et al.
Published: (2023)
Mitigating Premature Discretization with Progressive Quantization for Robust Vector Tokenization
by: Zhao, Wenhao, et al.
Published: (2026)
by: Zhao, Wenhao, et al.
Published: (2026)
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries
by: Kim, Junhyuck, et al.
Published: (2024)
by: Kim, Junhyuck, et al.
Published: (2024)
ZIP: An Efficient Zeroth-order Prompt Tuning for Black-box Vision-Language Models
by: Park, Seonghwan, et al.
Published: (2025)
by: Park, Seonghwan, et al.
Published: (2025)
SBVR: Summation of BitVector Representation for Efficient LLM Quantization
by: Bang, Wonjun, et al.
Published: (2025)
by: Bang, Wonjun, et al.
Published: (2025)
Bridging the Gap between Expert and Language Models: Concept-guided Chess Commentary Generation and Evaluation
by: Kim, Jaechang, et al.
Published: (2024)
by: Kim, Jaechang, et al.
Published: (2024)
CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compression
by: Gou, Dayin, et al.
Published: (2025)
by: Gou, Dayin, et al.
Published: (2025)
MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models
by: Kim, Taehyun, et al.
Published: (2024)
by: Kim, Taehyun, et al.
Published: (2024)
RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
by: Zhong, Zhengjia, et al.
Published: (2026)
by: Zhong, Zhengjia, et al.
Published: (2026)
EdgeCodec: Onboard Lightweight High Fidelity Neural Compressor with Residual Vector Quantization
by: Hodo, Benjamin, et al.
Published: (2025)
by: Hodo, Benjamin, et al.
Published: (2025)
Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime Requantization
by: Lee, Jungi, et al.
Published: (2024)
by: Lee, Jungi, et al.
Published: (2024)
Leech Lattice Vector Quantization for Efficient LLM Compression
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026)
Adaptive Tracking of a Single-Rigid-Body Character in Various Environments
by: Kwon, Taesoo, et al.
Published: (2023)
by: Kwon, Taesoo, et al.
Published: (2023)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
by: Kim, Junhyuck, et al.
Published: (2026)
by: Kim, Junhyuck, et al.
Published: (2026)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
by: Yang, Seongjun, et al.
Published: (2023)
by: Yang, Seongjun, et al.
Published: (2023)
UOTIP: Unbalanced Optimal Transport Map for Unpaired Inverse Problems
by: Lee, Donggyu, et al.
Published: (2026)
by: Lee, Donggyu, et al.
Published: (2026)
FoldToken: Learning Protein Language via Vector Quantization and Beyond
by: Gao, Zhangyang, et al.
Published: (2024)
by: Gao, Zhangyang, et al.
Published: (2024)
QEFT: Quantization for Efficient Fine-Tuning of LLMs
by: Lee, Changhun, et al.
Published: (2024)
by: Lee, Changhun, et al.
Published: (2024)
SAiD: Speech-driven Blendshape Facial Animation with Diffusion
by: Park, Inkyu, et al.
Published: (2023)
by: Park, Inkyu, et al.
Published: (2023)
Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech
by: Kim, Semin, et al.
Published: (2026)
by: Kim, Semin, et al.
Published: (2026)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
by: Lee, Namyoon, et al.
Published: (2026)
by: Lee, Namyoon, et al.
Published: (2026)
AL-PINN: Active Learning-Driven Physics-Informed Neural Networks for Efficient Sample Selection in Solving Partial Differential Equations
by: Park, Keon Vin
Published: (2025)
by: Park, Keon Vin
Published: (2025)
Learning to Continually Learn with the Bayesian Principle
by: Lee, Soochan, et al.
Published: (2024)
by: Lee, Soochan, et al.
Published: (2024)
Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs
by: Kim, Kibum, et al.
Published: (2026)
by: Kim, Kibum, et al.
Published: (2026)
QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs
by: Noh, Kanghyun, et al.
Published: (2026)
by: Noh, Kanghyun, et al.
Published: (2026)
Multi-Token Residual Prediction
by: Xu, Yufeng, et al.
Published: (2026)
by: Xu, Yufeng, et al.
Published: (2026)
Clip-Low Increases Entropy and Clip-High Decreases Entropy in Reinforcement Learning of Large Language Models
by: Park, Jaesung R., et al.
Published: (2025)
by: Park, Jaesung R., et al.
Published: (2025)
Similar Items
-
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
by: Park, Dongmin, et al.
Published: (2024) -
DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
by: Lee, Keon, et al.
Published: (2024) -
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
by: Kim, Jaehyeon, et al.
Published: (2024) -
Instance-Aware Group Quantization for Vision Transformers
by: Moon, Jaehyeon, et al.
Published: (2024) -
Task Vector Quantization for Memory-Efficient Model Merging
by: Kim, Youngeun, et al.
Published: (2025)