OTARo: Once Tuning for All Precisions toward Robust On-Device LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Shaoyuan, Chen, Zhixuan, Yang, Dawei, Yuan, Zhihang, Wu, Qiang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
por: Zhou, Sifan, et al.
Publicado: (2025)
por: Zhou, Sifan, et al.
Publicado: (2025)
MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
por: Hu, Xing, et al.
Publicado: (2025)
por: Hu, Xing, et al.
Publicado: (2025)
RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
por: Xu, Chen, et al.
Publicado: (2025)
por: Xu, Chen, et al.
Publicado: (2025)
SOLAR: Switchable Output Layer for Accuracy and Robustness in Once-for-All Training
por: Tareen, Shaharyar Ahmed Khan, et al.
Publicado: (2025)
por: Tareen, Shaharyar Ahmed Khan, et al.
Publicado: (2025)
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
por: Zhao, Zhixiong, et al.
Publicado: (2026)
por: Zhao, Zhixiong, et al.
Publicado: (2026)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
por: Xu, Zukang, et al.
Publicado: (2025)
por: Xu, Zukang, et al.
Publicado: (2025)
Mixed-Precision Conjugate Gradient Solvers with RL-Driven Precision Tuning
por: Chen, Xinye
Publicado: (2025)
por: Chen, Xinye
Publicado: (2025)
PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling
por: Yue, Yuxuan, et al.
Publicado: (2025)
por: Yue, Yuxuan, et al.
Publicado: (2025)
KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models
por: Xu, Zukang, et al.
Publicado: (2026)
por: Xu, Zukang, et al.
Publicado: (2026)
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
por: Xiong, Zheyang, et al.
Publicado: (2024)
por: Xiong, Zheyang, et al.
Publicado: (2024)
CogFormer: Learn All Your Models Once
por: Huang, Jerry M., et al.
Publicado: (2026)
por: Huang, Jerry M., et al.
Publicado: (2026)
Make Optimization Once and for All with Fine-grained Guidance
por: Shi, Mingjia, et al.
Publicado: (2025)
por: Shi, Mingjia, et al.
Publicado: (2025)
Text-to-Model: Text-Conditioned Neural Network Diffusion for Train-Once-for-All Personalization
por: Li, Zexi, et al.
Publicado: (2024)
por: Li, Zexi, et al.
Publicado: (2024)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
por: Hu, Xing, et al.
Publicado: (2024)
por: Hu, Xing, et al.
Publicado: (2024)
Goal-Conditioned Agents that Learn Everything All at Once
por: Matthews, Michael, et al.
Publicado: (2026)
por: Matthews, Michael, et al.
Publicado: (2026)
Once-for-All Channel Mixers (HYPERTINYPW): Generative Compression for TinyML
por: Shaalan, Yassien
Publicado: (2026)
por: Shaalan, Yassien
Publicado: (2026)
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
por: Hu, Xing, et al.
Publicado: (2025)
por: Hu, Xing, et al.
Publicado: (2025)
Attacking All Tasks at Once Using Adversarial Examples in Multi-Task Learning
por: Zhang, Lijun, et al.
Publicado: (2023)
por: Zhang, Lijun, et al.
Publicado: (2023)
Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers
por: Norgren, Victor
Publicado: (2026)
por: Norgren, Victor
Publicado: (2026)
FedProxy: Federated Fine-Tuning of LLMs via Proxy SLMs and Heterogeneity-Aware Fusion
por: Fan, Tao, et al.
Publicado: (2026)
por: Fan, Tao, et al.
Publicado: (2026)
Memory-Efficient Backpropagation for Fine-Tuning LLMs on Resource-Constrained Mobile Devices
por: Song, Congzheng, et al.
Publicado: (2025)
por: Song, Congzheng, et al.
Publicado: (2025)
BEAT: Balanced Frequency Adaptive Tuning for Long-Term Time-Series Forecasting
por: Li, Zhixuan, et al.
Publicado: (2025)
por: Li, Zhixuan, et al.
Publicado: (2025)
Memory-Optimized Once-For-All Network
por: Girard, Maxime, et al.
Publicado: (2024)
por: Girard, Maxime, et al.
Publicado: (2024)
PocketLLM: Enabling On-Device Fine-Tuning for Personalized LLMs
por: Peng, Dan, et al.
Publicado: (2024)
por: Peng, Dan, et al.
Publicado: (2024)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
por: Méloux, Maxime, et al.
Publicado: (2025)
por: Méloux, Maxime, et al.
Publicado: (2025)
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
por: Li, Ang, et al.
Publicado: (2025)
por: Li, Ang, et al.
Publicado: (2025)
BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations
por: Monninger, Thomas, et al.
Publicado: (2026)
por: Monninger, Thomas, et al.
Publicado: (2026)
Detecting the Future: All-at-Once Event Sequence Forecasting with Horizon Matching
por: Karpukhin, Ivan, et al.
Publicado: (2024)
por: Karpukhin, Ivan, et al.
Publicado: (2024)
GreenAuto: An Automated Platform for Sustainable AI Model Design on Edge Devices
por: Tu, Xiaolong, et al.
Publicado: (2025)
por: Tu, Xiaolong, et al.
Publicado: (2025)
From LLMs to Edge: Parameter-Efficient Fine-Tuning on Edge Devices
por: Slamanig, Georg, et al.
Publicado: (2025)
por: Slamanig, Georg, et al.
Publicado: (2025)
Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
por: Wang, Jingyao, et al.
Publicado: (2025)
por: Wang, Jingyao, et al.
Publicado: (2025)
Emergent Response Planning in LLMs
por: Dong, Zhichen, et al.
Publicado: (2025)
por: Dong, Zhichen, et al.
Publicado: (2025)
All in One and One for All: A Simple yet Effective Method towards Cross-domain Graph Pretraining
por: Zhao, Haihong, et al.
Publicado: (2024)
por: Zhao, Haihong, et al.
Publicado: (2024)
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
por: Lee, Jung Hyun, et al.
Publicado: (2025)
por: Lee, Jung Hyun, et al.
Publicado: (2025)
Uncertainty-Aware Gradient Signal-to-Noise Data Selection for Instruction Tuning
por: Yuan, Zhihang, et al.
Publicado: (2026)
por: Yuan, Zhihang, et al.
Publicado: (2026)
MobiLLM: Enabling LLM Fine-Tuning on the Mobile Device via Server Assisted Side Tuning
por: Li, Liang, et al.
Publicado: (2025)
por: Li, Liang, et al.
Publicado: (2025)
DεpS: Delayed ε-Shrinking for Faster Once-For-All Training
por: Annavajjala, Aditya, et al.
Publicado: (2024)
por: Annavajjala, Aditya, et al.
Publicado: (2024)
Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-Training
por: Fang, Wenzhi, et al.
Publicado: (2025)
por: Fang, Wenzhi, et al.
Publicado: (2025)
Equivariance Everywhere All At Once: A Recipe for Graph Foundation Models
por: Finkelshtein, Ben, et al.
Publicado: (2025)
por: Finkelshtein, Ben, et al.
Publicado: (2025)
Generalized Wasserstein Flow Matching: Transport Plans, Everywhere, All at Once
por: Piening, Moritz, et al.
Publicado: (2026)
por: Piening, Moritz, et al.
Publicado: (2026)
Ejemplares similares
-
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
por: Zhou, Sifan, et al.
Publicado: (2025) -
MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
por: Hu, Xing, et al.
Publicado: (2025) -
RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
por: Xu, Chen, et al.
Publicado: (2025) -
SOLAR: Switchable Output Layer for Accuracy and Robustness in Once-for-All Training
por: Tareen, Shaharyar Ahmed Khan, et al.
Publicado: (2025) -
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
por: Zhao, Zhixiong, et al.
Publicado: (2026)