Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Jiecheng, Tang, Ding, Fu, Rong, Hu, Boni, Xu, Haoran, Wang, Yi, Pei, Zhilin, Su, Zhongling, Liu, Liang, Zhang, Xingcheng, Zhang, Weiming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PackMamba: Efficient Processing of Variable-Length Sequences in Mamba training
von: Xu, Haoran, et al.
Veröffentlicht: (2024)
von: Xu, Haoran, et al.
Veröffentlicht: (2024)
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
von: Su, Zhongling, et al.
Veröffentlicht: (2025)
von: Su, Zhongling, et al.
Veröffentlicht: (2025)
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
von: Tang, Ding, et al.
Veröffentlicht: (2025)
von: Tang, Ding, et al.
Veröffentlicht: (2025)
Towards Effective and Efficient Continual Pre-training of Large Language Models
von: Chen, Jie, et al.
Veröffentlicht: (2024)
von: Chen, Jie, et al.
Veröffentlicht: (2024)
Noise Calibration: Plug-and-play Content-Preserving Video Enhancement using Pre-trained Video Diffusion Models
von: Yang, Qinyu, et al.
Veröffentlicht: (2024)
von: Yang, Qinyu, et al.
Veröffentlicht: (2024)
TC-GS: A Faster Gaussian Splatting Module Utilizing Tensor Cores
von: Liao, Zimu, et al.
Veröffentlicht: (2025)
von: Liao, Zimu, et al.
Veröffentlicht: (2025)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
von: Tang, Ding, et al.
Veröffentlicht: (2024)
von: Tang, Ding, et al.
Veröffentlicht: (2024)
HySparK: Hybrid Sparse Masking for Large Scale Medical Image Pre-Training
von: Tang, Fenghe, et al.
Veröffentlicht: (2024)
von: Tang, Fenghe, et al.
Veröffentlicht: (2024)
TraXion: Rethinking Pre-training Frameworks for Mobility and Beyond
von: Hsu, Shang-Ling, et al.
Veröffentlicht: (2026)
von: Hsu, Shang-Ling, et al.
Veröffentlicht: (2026)
FlowSteer: Towards Agents Designing Agentic Workflows via Reinforced Progressive Canvas Editing
von: Zhang, Mingda, et al.
Veröffentlicht: (2026)
von: Zhang, Mingda, et al.
Veröffentlicht: (2026)
On Initializing Transformers with Pre-trained Embeddings
von: Kim, Ha Young, et al.
Veröffentlicht: (2024)
von: Kim, Ha Young, et al.
Veröffentlicht: (2024)
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
von: Wang, Zhilin, et al.
Veröffentlicht: (2026)
von: Wang, Zhilin, et al.
Veröffentlicht: (2026)
Pre-trained Language Model with Prompts for Temporal Knowledge Graph Completion
von: Xu, Wenjie, et al.
Veröffentlicht: (2023)
von: Xu, Wenjie, et al.
Veröffentlicht: (2023)
Exploring LLM-based Verilog Code Generation with Data-Efficient Fine-Tuning and Testbench Automation
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2026)
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2026)
Asset Pricing in Pre-trained Transformer
von: Lai, Shanyan
Veröffentlicht: (2025)
von: Lai, Shanyan
Veröffentlicht: (2025)
Unveiling Attractor Cycles in Large Language Models: A Dynamical Systems View of Successive Paraphrasing
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
von: He, Yanjin, et al.
Veröffentlicht: (2025)
von: He, Yanjin, et al.
Veröffentlicht: (2025)
Zero-Shot Spam Email Classification Using Pre-trained Large Language Models
von: Rojas-Galeano, Sergio
Veröffentlicht: (2024)
von: Rojas-Galeano, Sergio
Veröffentlicht: (2024)
The Native Spiking Microarchitecture: From Iontronic Primitives to Bit-Exact FP8 Arithmetic
von: Tang, Zhengzheng
Veröffentlicht: (2025)
von: Tang, Zhengzheng
Veröffentlicht: (2025)
Efficient Attention: Attention with Linear Complexities
von: Shen, Zhuoran, et al.
Veröffentlicht: (2018)
von: Shen, Zhuoran, et al.
Veröffentlicht: (2018)
Leveraging Personalized PageRank and Higher-Order Topological Structures for Heterophily Mitigation in Graph Neural Networks
von: Wang, Yumeng, et al.
Veröffentlicht: (2025)
von: Wang, Yumeng, et al.
Veröffentlicht: (2025)
Universal Adversarial Perturbations for Vision-Language Pre-trained Models
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2024)
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2024)
Fisheye-GS: Lightweight and Extensible Gaussian Splatting Module for Fisheye Cameras
von: Liao, Zimu, et al.
Veröffentlicht: (2024)
von: Liao, Zimu, et al.
Veröffentlicht: (2024)
FlashGS: Efficient 3D Gaussian Splatting for Large-scale and High-resolution Rendering
von: Feng, Guofeng, et al.
Veröffentlicht: (2024)
von: Feng, Guofeng, et al.
Veröffentlicht: (2024)
FedAttr: Towards Privacy-preserving Client-Level Attribution in Federated LLM Fine-tuning
von: Zhang, Su, et al.
Veröffentlicht: (2026)
von: Zhang, Su, et al.
Veröffentlicht: (2026)
4D-PreNet: A Unified Preprocessing Framework for 4D-STEM Data Analysis
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
CoMA: Complementary Masking and Hierarchical Dynamic Multi-Window Self-Attention in a Unified Pre-training Framework
von: Li, Jiaxuan, et al.
Veröffentlicht: (2025)
von: Li, Jiaxuan, et al.
Veröffentlicht: (2025)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
DV-Matcher: Deformation-based Non-Rigid Point Cloud Matching Guided by Pre-trained Visual Features
von: Chen, Zhangquan, et al.
Veröffentlicht: (2024)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2024)
DatUS^2: Data-driven Unsupervised Semantic Segmentation with Pre-trained Self-supervised Vision Transformer
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
Hierarchical Pre-Training of Vision Encoders with Large Language Models
von: Lee, Eugene, et al.
Veröffentlicht: (2026)
von: Lee, Eugene, et al.
Veröffentlicht: (2026)
Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models
von: Wen, Zhihao, et al.
Veröffentlicht: (2025)
von: Wen, Zhihao, et al.
Veröffentlicht: (2025)
Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
von: Eymaël, Alexandre, et al.
Veröffentlicht: (2024)
von: Eymaël, Alexandre, et al.
Veröffentlicht: (2024)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2025)
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2025)
On the Role of Pre-trained Embeddings in Binary Code Analysis
von: Maier, Alwin, et al.
Veröffentlicht: (2025)
von: Maier, Alwin, et al.
Veröffentlicht: (2025)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
von: Zhang, Xue
Veröffentlicht: (2025)
von: Zhang, Xue
Veröffentlicht: (2025)
VLMShield: Efficient and Robust Defense of Vision-Language Models against Malicious Prompts
von: Qi, Peigui, et al.
Veröffentlicht: (2026)
von: Qi, Peigui, et al.
Veröffentlicht: (2026)
Pre-trained LLMs Meet Sequential Recommenders: Efficient User-Centric Knowledge Distillation
von: Severin, Nikita, et al.
Veröffentlicht: (2026)
von: Severin, Nikita, et al.
Veröffentlicht: (2026)
Out-of-Sight Embodied Agents: Multimodal Tracking, Sensor Fusion, and Trajectory Forecasting
von: Zhang, Haichao, et al.
Veröffentlicht: (2025)
von: Zhang, Haichao, et al.
Veröffentlicht: (2025)
A Multi-Pass Large Language Model Framework for Precise and Efficient Radiology Report Error Detection
von: Kim, Songsoo, et al.
Veröffentlicht: (2025)
von: Kim, Songsoo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PackMamba: Efficient Processing of Variable-Length Sequences in Mamba training
von: Xu, Haoran, et al.
Veröffentlicht: (2024) -
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
von: Su, Zhongling, et al.
Veröffentlicht: (2025) -
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
von: Tang, Ding, et al.
Veröffentlicht: (2025) -
Towards Effective and Efficient Continual Pre-training of Large Language Models
von: Chen, Jie, et al.
Veröffentlicht: (2024) -
Noise Calibration: Plug-and-play Content-Preserving Video Enhancement using Pre-trained Video Diffusion Models
von: Yang, Qinyu, et al.
Veröffentlicht: (2024)