Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Park, Yeonhong, Hyun, Jake, Cho, SangLyul, Sim, Bonggeun, Lee, Jae W. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
di: Park, Yeonhong, et al.
Pubblicazione: (2024)
di: Park, Yeonhong, et al.
Pubblicazione: (2024)
DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
di: Kwon, Sangwoo, et al.
Pubblicazione: (2025)
di: Kwon, Sangwoo, et al.
Pubblicazione: (2025)
MAGE: All-[MASK] Block Already Knows Where to Look in Diffusion LLM
di: Kwon, Omin, et al.
Pubblicazione: (2026)
di: Kwon, Omin, et al.
Pubblicazione: (2026)
GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance
di: Kim, Jinuk, et al.
Pubblicazione: (2025)
di: Kim, Jinuk, et al.
Pubblicazione: (2025)
NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
di: Lee, Haeun, et al.
Pubblicazione: (2025)
di: Lee, Haeun, et al.
Pubblicazione: (2025)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
di: Lee, Deokjae, et al.
Pubblicazione: (2025)
di: Lee, Deokjae, et al.
Pubblicazione: (2025)
Dual Precision Deep Neural Network
di: Park, Jae Hyun, et al.
Pubblicazione: (2020)
di: Park, Jae Hyun, et al.
Pubblicazione: (2020)
AmoebaLLM: Constructing Any-Shape Large Language Models for Efficient and Instant Deployment
di: Fu, Yonggan, et al.
Pubblicazione: (2024)
di: Fu, Yonggan, et al.
Pubblicazione: (2024)
Collage: Light-Weight Low-Precision Strategy for LLM Training
di: Yu, Tao, et al.
Pubblicazione: (2024)
di: Yu, Tao, et al.
Pubblicazione: (2024)
Affordable Precision Agriculture: A Deployment-Oriented Review of Low-Cost, Low-Power Edge AI and TinyML for Resource-Constrained Farming Systems
di: Samanta, Riya, et al.
Pubblicazione: (2026)
di: Samanta, Riya, et al.
Pubblicazione: (2026)
Log-Time K-Means Clustering for 1D Data: Novel Approaches with Proof and Implementation
di: Hyun, Jake
Pubblicazione: (2024)
di: Hyun, Jake
Pubblicazione: (2024)
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
di: Lee, Jungi, et al.
Pubblicazione: (2025)
di: Lee, Jungi, et al.
Pubblicazione: (2025)
MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models
di: Kim, Taehyun, et al.
Pubblicazione: (2024)
di: Kim, Taehyun, et al.
Pubblicazione: (2024)
One Size Fits All for Semantic Shifts: Adaptive Prompt Tuning for Continual Learning
di: Kim, Doyoung, et al.
Pubblicazione: (2023)
di: Kim, Doyoung, et al.
Pubblicazione: (2023)
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
di: Xu, Zifei, et al.
Pubblicazione: (2024)
di: Xu, Zifei, et al.
Pubblicazione: (2024)
CAPER: Enhancing Career Trajectory Prediction using Temporal Knowledge Graph and Ternary Relationship
di: Lee, Yeon-Chang, et al.
Pubblicazione: (2024)
di: Lee, Yeon-Chang, et al.
Pubblicazione: (2024)
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
di: Chung, Jae-Won, et al.
Pubblicazione: (2026)
di: Chung, Jae-Won, et al.
Pubblicazione: (2026)
Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving
di: Ma, Jeff J., et al.
Pubblicazione: (2025)
di: Ma, Jeff J., et al.
Pubblicazione: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
di: Wang, Dongwei, et al.
Pubblicazione: (2026)
di: Wang, Dongwei, et al.
Pubblicazione: (2026)
NExT-GPT: Any-to-Any Multimodal LLM
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
QEFT: Quantization for Efficient Fine-Tuning of LLMs
di: Lee, Changhun, et al.
Pubblicazione: (2024)
di: Lee, Changhun, et al.
Pubblicazione: (2024)
From Bits to Chips: An LLM-based Hardware-Aware Quantization Agent for Streamlined Deployment of LLMs
di: Deng, Kaiyuan, et al.
Pubblicazione: (2026)
di: Deng, Kaiyuan, et al.
Pubblicazione: (2026)
Low-pass Personalized Subgraph Federated Recommendation
di: Sim, Wooseok, et al.
Pubblicazione: (2026)
di: Sim, Wooseok, et al.
Pubblicazione: (2026)
VPO: Leveraging the Number of Votes in Preference Optimization
di: Cho, Jae Hyeon, et al.
Pubblicazione: (2024)
di: Cho, Jae Hyeon, et al.
Pubblicazione: (2024)
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
di: Zhang, Jiawei, et al.
Pubblicazione: (2025)
di: Zhang, Jiawei, et al.
Pubblicazione: (2025)
Log-Normal Multiplicative Dynamics for Stable Low-Precision Training of Large Networks
di: Nishida, Keigo, et al.
Pubblicazione: (2025)
di: Nishida, Keigo, et al.
Pubblicazione: (2025)
Self-Supervised Curriculum Generation for Autonomous Reinforcement Learning without Task-Specific Knowledge
di: Lee, Sang-Hyun, et al.
Pubblicazione: (2023)
di: Lee, Sang-Hyun, et al.
Pubblicazione: (2023)
Power Hungry Processing: Watts Driving the Cost of AI Deployment?
di: Luccioni, Alexandra Sasha, et al.
Pubblicazione: (2023)
di: Luccioni, Alexandra Sasha, et al.
Pubblicazione: (2023)
Multimodal Crystal Flow: Any-to-Any Modality Generation for Unified Crystal Modeling
di: Seong, Kiyoung, et al.
Pubblicazione: (2026)
di: Seong, Kiyoung, et al.
Pubblicazione: (2026)
Continuous Semantic Caching for Low-Cost LLM Serving
di: Atalar, Baran, et al.
Pubblicazione: (2026)
di: Atalar, Baran, et al.
Pubblicazione: (2026)
InstantNet: Automated Generation and Deployment of Instantaneously Switchable-Precision Networks
di: Fu, Yonggan, et al.
Pubblicazione: (2021)
di: Fu, Yonggan, et al.
Pubblicazione: (2021)
SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights
di: Müller, Lorenz K., et al.
Pubblicazione: (2025)
di: Müller, Lorenz K., et al.
Pubblicazione: (2025)
Ex Uno Pluria: Insights on Ensembling in Low Precision Number Systems
di: Nam, Giung, et al.
Pubblicazione: (2024)
di: Nam, Giung, et al.
Pubblicazione: (2024)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
di: Alexandridis, Kosmas, et al.
Pubblicazione: (2025)
di: Alexandridis, Kosmas, et al.
Pubblicazione: (2025)
Any-Way Meta Learning
di: Lee, Junhoo, et al.
Pubblicazione: (2024)
di: Lee, Junhoo, et al.
Pubblicazione: (2024)
Adaptive MSD-Splitting: Enhancing C4.5 and Random Forests for Skewed Continuous Attributes
di: Lee, Jake
Pubblicazione: (2026)
di: Lee, Jake
Pubblicazione: (2026)
Choose Your Model Size: Any Compression of Large Language Models Without Re-Computation
di: Genzel, Martin, et al.
Pubblicazione: (2025)
di: Genzel, Martin, et al.
Pubblicazione: (2025)
Efficiently Deploying LLMs with Controlled Risk
di: Zellinger, Michael J., et al.
Pubblicazione: (2024)
di: Zellinger, Michael J., et al.
Pubblicazione: (2024)
A Cost-Benefit Analysis of On-Premise Large Language Model Deployment: Breaking Even with Commercial LLM Services
di: Pan, Guanzhong, et al.
Pubblicazione: (2025)
di: Pan, Guanzhong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
di: Park, Yeonhong, et al.
Pubblicazione: (2024) -
DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
di: Kwon, Sangwoo, et al.
Pubblicazione: (2025) -
MAGE: All-[MASK] Block Already Knows Where to Look in Diffusion LLM
di: Kwon, Omin, et al.
Pubblicazione: (2026) -
GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance
di: Kim, Jinuk, et al.
Pubblicazione: (2025) -
NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
di: Lee, Haeun, et al.
Pubblicazione: (2025)