Guardado en:
| Autores principales: | Tseng, Albert, Yu, Tao, Park, Youngsuk |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2502.20586 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Stochastic Rounding for LLM Training: Theory and Practice
por: Ozkara, Kaan, et al.
Publicado: (2025)
por: Ozkara, Kaan, et al.
Publicado: (2025)
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
por: Bian, Song, et al.
Publicado: (2025)
por: Bian, Song, et al.
Publicado: (2025)
Oscillation-Reduced MXFP4 Training for Vision Transformers
por: Chen, Yuxiang, et al.
Publicado: (2025)
por: Chen, Yuxiang, et al.
Publicado: (2025)
TritonRL: Training LLMs to Think and Code Triton Without Cheating
por: Woo, Jiin, et al.
Publicado: (2025)
por: Woo, Jiin, et al.
Publicado: (2025)
Recipes for Pre-training LLMs with MXFP8
por: Mishra, Asit, et al.
Publicado: (2025)
por: Mishra, Asit, et al.
Publicado: (2025)
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
por: Liu, Hongyi, et al.
Publicado: (2025)
por: Liu, Hongyi, et al.
Publicado: (2025)
MuonBP: Faster Muon via Block-Periodic Orthogonalization
por: Khaled, Ahmed, et al.
Publicado: (2025)
por: Khaled, Ahmed, et al.
Publicado: (2025)
Block Rotation is All You Need for MXFP4 Quantization
por: Shao, Yuantian, et al.
Publicado: (2025)
por: Shao, Yuantian, et al.
Publicado: (2025)
TORQ: Two-Level Orthogonal Rotation for MXFP4 Quantization
por: Xu, Zukang, et al.
Publicado: (2026)
por: Xu, Zukang, et al.
Publicado: (2026)
Pretraining large language models with MXFP4 on Native FP4 Hardware
por: Cim, Musa, et al.
Publicado: (2026)
por: Cim, Musa, et al.
Publicado: (2026)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
por: Chhugani, Jatin, et al.
Publicado: (2026)
por: Chhugani, Jatin, et al.
Publicado: (2026)
Collage: Light-Weight Low-Precision Strategy for LLM Training
por: Yu, Tao, et al.
Publicado: (2024)
por: Yu, Tao, et al.
Publicado: (2024)
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
por: Liu, Hongyi, et al.
Publicado: (2025)
por: Liu, Hongyi, et al.
Publicado: (2025)
Online Posterior Sampling with a Diffusion Prior
por: Kveton, Branislav, et al.
Publicado: (2024)
por: Kveton, Branislav, et al.
Publicado: (2024)
MXNorm: Reusing MXFP block scales for efficient tensor normalisation
por: McLean, Callum, et al.
Publicado: (2026)
por: McLean, Callum, et al.
Publicado: (2026)
Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference
por: Ding, Yifu, et al.
Publicado: (2026)
por: Ding, Yifu, et al.
Publicado: (2026)
Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor
por: Li, Xiaocan, et al.
Publicado: (2026)
por: Li, Xiaocan, et al.
Publicado: (2026)
Shadow Cones: A Generalized Framework for Partial Order Embeddings
por: Yu, Tao, et al.
Publicado: (2023)
por: Yu, Tao, et al.
Publicado: (2023)
Theoretical Guarantees of Learning Ensembling Strategies with Applications to Time Series Forecasting
por: Hasson, Hilaf, et al.
Publicado: (2023)
por: Hasson, Hilaf, et al.
Publicado: (2023)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
por: Gautam, Tanmay, et al.
Publicado: (2024)
por: Gautam, Tanmay, et al.
Publicado: (2024)
L$^3$: Large Lookup Layers
por: Tseng, Albert, et al.
Publicado: (2026)
por: Tseng, Albert, et al.
Publicado: (2026)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
por: Deng, Wenlong, et al.
Publicado: (2026)
por: Deng, Wenlong, et al.
Publicado: (2026)
RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models
por: Wei, Quan, et al.
Publicado: (2025)
por: Wei, Quan, et al.
Publicado: (2025)
Verifier-free Test-Time Sampling for Vision Language Action Models
por: Jang, Suhyeok, et al.
Publicado: (2025)
por: Jang, Suhyeok, et al.
Publicado: (2025)
Model-Preserving Adaptive Rounding
por: Tseng, Albert, et al.
Publicado: (2025)
por: Tseng, Albert, et al.
Publicado: (2025)
Metis: Training LLMs with FP4 Quantization
por: Cao, Hengjie, et al.
Publicado: (2025)
por: Cao, Hengjie, et al.
Publicado: (2025)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
por: Ouyang, Xu, et al.
Publicado: (2024)
por: Ouyang, Xu, et al.
Publicado: (2024)
QTIP: Quantization with Trellises and Incoherence Processing
por: Tseng, Albert, et al.
Publicado: (2024)
por: Tseng, Albert, et al.
Publicado: (2024)
MuCon: Clipped Muon Updates for LLM Training
por: Yi, Albert
Publicado: (2026)
por: Yi, Albert
Publicado: (2026)
Training Dynamics Impact Post-Training Quantization Robustness
por: Catalan-Tatjer, Albert, et al.
Publicado: (2025)
por: Catalan-Tatjer, Albert, et al.
Publicado: (2025)
Inference Optimization of Foundation Models on AI Accelerators
por: Park, Youngsuk, et al.
Publicado: (2024)
por: Park, Youngsuk, et al.
Publicado: (2024)
MedM2T: A MultiModal Framework for Time-Aware Modeling with Electronic Health Record and Electrocardiogram Data
por: Kuo, Yu-Chen, et al.
Publicado: (2025)
por: Kuo, Yu-Chen, et al.
Publicado: (2025)
Learning-Based WiFi Fingerprint Inpainting via Generative Adversarial Networks
por: Chan, Yu, et al.
Publicado: (2024)
por: Chan, Yu, et al.
Publicado: (2024)
StreetMath: Study of LLMs' Approximation Behaviors
por: Tseng, Chiung-Yi, et al.
Publicado: (2025)
por: Tseng, Chiung-Yi, et al.
Publicado: (2025)
Laplace Approximation For Tensor Train Kernel Machines In System Identification
por: Saiapin, Albert, et al.
Publicado: (2025)
por: Saiapin, Albert, et al.
Publicado: (2025)
Physics-Informed Neural Network for Predicting Out-of-Training-Range TCAD Solution with Minimized Domain Expertise
por: Lu, Albert, et al.
Publicado: (2024)
por: Lu, Albert, et al.
Publicado: (2024)
FP4 All the Way: Fully Quantized Training of LLMs
por: Chmiel, Brian, et al.
Publicado: (2025)
por: Chmiel, Brian, et al.
Publicado: (2025)
CAAP: Class-Dependent Automatic Data Augmentation Based On Adaptive Policies For Time Series
por: Chang, Tien-Yu, et al.
Publicado: (2024)
por: Chang, Tien-Yu, et al.
Publicado: (2024)
Test-Time Training on Graphs with Large Language Models (LLMs)
por: Zhang, Jiaxin, et al.
Publicado: (2024)
por: Zhang, Jiaxin, et al.
Publicado: (2024)
LLM4TS: Aligning Pre-Trained LLMs as Data-Efficient Time-Series Forecasters
por: Chang, Ching, et al.
Publicado: (2023)
por: Chang, Ching, et al.
Publicado: (2023)
Ejemplares similares
-
Stochastic Rounding for LLM Training: Theory and Practice
por: Ozkara, Kaan, et al.
Publicado: (2025) -
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
por: Bian, Song, et al.
Publicado: (2025) -
Oscillation-Reduced MXFP4 Training for Vision Transformers
por: Chen, Yuxiang, et al.
Publicado: (2025) -
TritonRL: Training LLMs to Think and Code Triton Without Cheating
por: Woo, Jiin, et al.
Publicado: (2025) -
Recipes for Pre-training LLMs with MXFP8
por: Mishra, Asit, et al.
Publicado: (2025)