Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Junhyuck, Ewer, Ethan, Moon, Taehong, Park, Jongho, Papailiopoulos, Dimitris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries
von: Kim, Junhyuck, et al.
Veröffentlicht: (2024)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2024)
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
von: Park, Jongho, et al.
Veröffentlicht: (2024)
von: Park, Jongho, et al.
Veröffentlicht: (2024)
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
von: Zeng, Thomas, et al.
Veröffentlicht: (2025)
von: Zeng, Thomas, et al.
Veröffentlicht: (2025)
Beyond RLHF: A Unified Theoretical Framework of Alignment
von: Yun, Jihun, et al.
Veröffentlicht: (2025)
von: Yun, Jihun, et al.
Veröffentlicht: (2025)
Wait, Wait, Wait... Why Do Reasoning Models Loop?
von: Pipis, Charilaos, et al.
Veröffentlicht: (2025)
von: Pipis, Charilaos, et al.
Veröffentlicht: (2025)
Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
Endless Terminals: Scaling RL Environments for Terminal Agents
von: Gandhi, Kanishk, et al.
Veröffentlicht: (2026)
von: Gandhi, Kanishk, et al.
Veröffentlicht: (2026)
Not All LLM Reasoners Are Created Equal
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
ENTP: Encoder-only Next Token Prediction
von: Ewer, Ethan, et al.
Veröffentlicht: (2024)
von: Ewer, Ethan, et al.
Veröffentlicht: (2024)
Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
Looped Transformers are Better at Learning Learning Algorithms
von: Yang, Liu, et al.
Veröffentlicht: (2023)
von: Yang, Liu, et al.
Veröffentlicht: (2023)
Task Vectors in In-Context Learning: Emergence, Formation, and Benefit
von: Yang, Liu, et al.
Veröffentlicht: (2025)
von: Yang, Liu, et al.
Veröffentlicht: (2025)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
von: Xiong, Zheyang, et al.
Veröffentlicht: (2024)
von: Xiong, Zheyang, et al.
Veröffentlicht: (2024)
Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
von: Lee, Nayoung, et al.
Veröffentlicht: (2025)
von: Lee, Nayoung, et al.
Veröffentlicht: (2025)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
von: Yang, Seongjun, et al.
Veröffentlicht: (2023)
von: Yang, Seongjun, et al.
Veröffentlicht: (2023)
Not All Code Is Equal: A Data-Centric Study of Code Complexity and LLM Reasoning
von: Twist, Lukas, et al.
Veröffentlicht: (2026)
von: Twist, Lukas, et al.
Veröffentlicht: (2026)
Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
von: Hwang, Dongyoon, et al.
Veröffentlicht: (2025)
von: Hwang, Dongyoon, et al.
Veröffentlicht: (2025)
Memory-Efficient Personalization of Text-to-Image Diffusion Models via Selective Optimization Strategies
von: Choi, Seokeon, et al.
Veröffentlicht: (2025)
von: Choi, Seokeon, et al.
Veröffentlicht: (2025)
Not All Forgetting Is Equal: Architecture-Dependent Retention Dynamics in Fine-Tuned Image Classifiers
von: Daga, Miit, et al.
Veröffentlicht: (2026)
von: Daga, Miit, et al.
Veröffentlicht: (2026)
ReJump: A Tree-Jump Representation for Analyzing and Improving LLM Reasoning
von: Zeng, Yuchen, et al.
Veröffentlicht: (2025)
von: Zeng, Yuchen, et al.
Veröffentlicht: (2025)
Not All Negative Samples Are Equal: LLMs Learn Better from Plausible Reasoning
von: Di, Zixiang, et al.
Veröffentlicht: (2026)
von: Di, Zixiang, et al.
Veröffentlicht: (2026)
MSQ: Memory-Efficient Bit Sparsification Quantization
von: Han, Seokho, et al.
Veröffentlicht: (2025)
von: Han, Seokho, et al.
Veröffentlicht: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
Selective Matching Losses -- Not All Scores Are Created Equal
von: Shamir, Gil I., et al.
Veröffentlicht: (2025)
von: Shamir, Gil I., et al.
Veröffentlicht: (2025)
How Well Can Transformers Emulate In-context Newton's Method?
von: Giannou, Angeliki, et al.
Veröffentlicht: (2024)
von: Giannou, Angeliki, et al.
Veröffentlicht: (2024)
MAGNET: Autonomous Expert Model Generation via Decentralized Autoresearch and BitNet Training
von: Kim, Yongwan, et al.
Veröffentlicht: (2026)
von: Kim, Yongwan, et al.
Veröffentlicht: (2026)
Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
von: Xu, Binxing, et al.
Veröffentlicht: (2026)
von: Xu, Binxing, et al.
Veröffentlicht: (2026)
Not All Denoising Steps Are Equal: Model Scheduling for Faster Masked Diffusion Language Models
von: Sedykh, Ivan, et al.
Veröffentlicht: (2026)
von: Sedykh, Ivan, et al.
Veröffentlicht: (2026)
Process-Aware Procurement Lead Time Prediction for Shipyard Delay Mitigation
von: Lee, Yongjae, et al.
Veröffentlicht: (2026)
von: Lee, Yongjae, et al.
Veröffentlicht: (2026)
All Nodes are created Not Equal: Node-Specific Layer Aggregation and Filtration for GNN
von: Wang, Shilong, et al.
Veröffentlicht: (2024)
von: Wang, Shilong, et al.
Veröffentlicht: (2024)
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
von: Kim, Juno, et al.
Veröffentlicht: (2026)
von: Kim, Juno, et al.
Veröffentlicht: (2026)
Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
von: Kim, Yeongmin, et al.
Veröffentlicht: (2026)
von: Kim, Yeongmin, et al.
Veröffentlicht: (2026)
Taming Gradient Oversmoothing and Expansion in Graph Neural Networks
von: Park, MoonJeong, et al.
Veröffentlicht: (2024)
von: Park, MoonJeong, et al.
Veröffentlicht: (2024)
Adaptive Tracking of a Single-Rigid-Body Character in Various Environments
von: Kwon, Taesoo, et al.
Veröffentlicht: (2023)
von: Kwon, Taesoo, et al.
Veröffentlicht: (2023)
MIDUS: Memory-Infused Depth Up-Scaling
von: Kim, Taero, et al.
Veröffentlicht: (2025)
von: Kim, Taero, et al.
Veröffentlicht: (2025)
Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
von: Jo, Dongwon, et al.
Veröffentlicht: (2024)
von: Jo, Dongwon, et al.
Veröffentlicht: (2024)
Disentangling the Spectral Properties of the Hodge Laplacian: Not All Small Eigenvalues Are Equal
von: Grande, Vincent P., et al.
Veröffentlicht: (2023)
von: Grande, Vincent P., et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries
von: Kim, Junhyuck, et al.
Veröffentlicht: (2024) -
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
von: Park, Jongho, et al.
Veröffentlicht: (2024) -
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024) -
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
von: Zeng, Thomas, et al.
Veröffentlicht: (2025) -
Beyond RLHF: A Unified Theoretical Framework of Alignment
von: Yun, Jihun, et al.
Veröffentlicht: (2025)