Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Jo, Dongwon, Kim, Taesu, Kim, Yulhwa, Kim, Jae-Joon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025)
by: Jo, Dongwon, et al.
Published: (2025)
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
by: Song, Jiwon, et al.
Published: (2024)
by: Song, Jiwon, et al.
Published: (2024)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
by: Jeon, Hyesung, et al.
Published: (2024)
by: Jeon, Hyesung, et al.
Published: (2024)
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
by: Song, Jiwon, et al.
Published: (2025)
by: Song, Jiwon, et al.
Published: (2025)
QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs
by: Noh, Kanghyun, et al.
Published: (2026)
by: Noh, Kanghyun, et al.
Published: (2026)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)
by: Kim, Jiyoon, et al.
Published: (2025)
QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models
by: Jeon, Hyesung, et al.
Published: (2025)
by: Jeon, Hyesung, et al.
Published: (2025)
GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
by: Jung, Yeonjoon, et al.
Published: (2025)
by: Jung, Yeonjoon, et al.
Published: (2025)
LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents
by: Jeon, Hyesung, et al.
Published: (2026)
by: Jeon, Hyesung, et al.
Published: (2026)
Personalized Federated Learning for Gradient Alignment
by: Kim, Dongwon, et al.
Published: (2026)
by: Kim, Dongwon, et al.
Published: (2026)
Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
NeuralSVCD for Efficient Swept Volume Collision Detection
by: Son, Dongwon, et al.
Published: (2025)
by: Son, Dongwon, et al.
Published: (2025)
MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models
by: Kim, Taehyun, et al.
Published: (2024)
by: Kim, Taehyun, et al.
Published: (2024)
On Integer Programming for the Binarized Neural Network Verification Problem
by: Kim, Woojin, et al.
Published: (2025)
by: Kim, Woojin, et al.
Published: (2025)
Graph Generation with Diffusion Mixture
by: Jo, Jaehyeong, et al.
Published: (2023)
by: Jo, Jaehyeong, et al.
Published: (2023)
FastAV: Efficient Token Pruning for Audio-Visual Large Language Model Inference
by: Jung, Chaeyoung, et al.
Published: (2026)
by: Jung, Chaeyoung, et al.
Published: (2026)
Bootstrapping Top-down Information for Self-modulating Slot Attention
by: Kim, Dongwon, et al.
Published: (2024)
by: Kim, Dongwon, et al.
Published: (2024)
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference
by: Kang, Beomseok, et al.
Published: (2026)
by: Kang, Beomseok, et al.
Published: (2026)
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
by: Bae, Sangmin, et al.
Published: (2025)
by: Bae, Sangmin, et al.
Published: (2025)
NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
by: Chong, Hyochan, et al.
Published: (2026)
by: Chong, Hyochan, et al.
Published: (2026)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
by: Kim, Taesu, et al.
Published: (2024)
by: Kim, Taesu, et al.
Published: (2024)
COMPASS: A Compiler Framework for Resource-Constrained Crossbar-Array Based In-Memory Deep Learning Accelerators
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning
by: Lee, Yuna, et al.
Published: (2026)
by: Lee, Yuna, et al.
Published: (2026)
RDB2G-Bench: A Comprehensive Benchmark for Automatic Graph Modeling of Relational Databases
by: Choi, Dongwon, et al.
Published: (2025)
by: Choi, Dongwon, et al.
Published: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
by: Kim, Junhyuck, et al.
Published: (2026)
by: Kim, Junhyuck, et al.
Published: (2026)
Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
by: Kim, Dahun, et al.
Published: (2025)
by: Kim, Dahun, et al.
Published: (2025)
GraphT5: Unified Molecular Graph-Language Modeling via Multi-Modal Cross-Token Attention
by: Kim, Sangyeup, et al.
Published: (2025)
by: Kim, Sangyeup, et al.
Published: (2025)
Adaptive Task Vectors for Large Language Models
by: Kang, Joonseong, et al.
Published: (2025)
by: Kang, Joonseong, et al.
Published: (2025)
AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models
by: Baek, Changwoo, et al.
Published: (2026)
by: Baek, Changwoo, et al.
Published: (2026)
A Survey on Mixture of Experts in Large Language Models
by: Cai, Weilin, et al.
Published: (2024)
by: Cai, Weilin, et al.
Published: (2024)
A New Lens on Homelessness: Daily Tent Monitoring with 311 Calls and Street Images
by: Jung, Wooyong, et al.
Published: (2025)
by: Jung, Wooyong, et al.
Published: (2025)
Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language Models
by: Son, Hyegang, et al.
Published: (2024)
by: Son, Hyegang, et al.
Published: (2024)
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens
by: Kim, Jaehyeon, et al.
Published: (2024)
by: Kim, Jaehyeon, et al.
Published: (2024)
SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling
by: Kim, Dahyun, et al.
Published: (2023)
by: Kim, Dahyun, et al.
Published: (2023)
BrepCoder: A Unified Multimodal Large Language Model for Multi-task B-rep Reasoning
by: Kim, Mingi, et al.
Published: (2026)
by: Kim, Mingi, et al.
Published: (2026)
Probing the Impact of Scale on Data-Efficient, Generalist Transformer World Models for Atari
by: Kim, Jooyeon
Published: (2026)
by: Kim, Jooyeon
Published: (2026)
Clip-Low Increases Entropy and Clip-High Decreases Entropy in Reinforcement Learning of Large Language Models
by: Park, Jaesung R., et al.
Published: (2025)
by: Park, Jaesung R., et al.
Published: (2025)
A Robust Foundation Model for Conservation Laws: Injecting Context into Flux Neural Operators via Recurrent Vision Transformers
by: Kim, Taeyoung, et al.
Published: (2026)
by: Kim, Taeyoung, et al.
Published: (2026)
Efficient Epistemic Uncertainty Estimation for Large Language Models via Knowledge Distillation
by: Park, Seonghyeon, et al.
Published: (2026)
by: Park, Seonghyeon, et al.
Published: (2026)
Similar Items
-
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025) -
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
by: Song, Jiwon, et al.
Published: (2024) -
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026) -
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
by: Jeon, Hyesung, et al.
Published: (2024) -
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
by: Song, Jiwon, et al.
Published: (2025)