Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Ziyang, Ding, Xinheng, Yuan, Jiayi, Liu, Rixin, Mao, Huizi, Xing, Jiarong, Liu, Zirui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Defeating the Training-Inference Mismatch via FP16
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
Diagnosing Training Inference Mismatch in LLM Reinforcement Learning
von: Zhong, Tianle, et al.
Veröffentlicht: (2026)
von: Zhong, Tianle, et al.
Veröffentlicht: (2026)
Communication Compression for Tensor Parallel LLM Inference
von: Hansen-Palmus, Jan, et al.
Veröffentlicht: (2024)
von: Hansen-Palmus, Jan, et al.
Veröffentlicht: (2024)
RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
Make Some Noise: Unlocking Language Model Parallel Inference Capability through Noisy Training
von: Wang, Yixuan, et al.
Veröffentlicht: (2024)
von: Wang, Yixuan, et al.
Veröffentlicht: (2024)
LLMExplainer: Large Language Model based Bayesian Inference for Graph Explanation Generation
von: Zhang, Jiaxing, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxing, et al.
Veröffentlicht: (2024)
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
von: Rodionov, Gleb, et al.
Veröffentlicht: (2025)
von: Rodionov, Gleb, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference
von: Yuan, Jiayi, et al.
Veröffentlicht: (2025)
von: Yuan, Jiayi, et al.
Veröffentlicht: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
von: Ye, Jiancai, et al.
Veröffentlicht: (2026)
von: Ye, Jiancai, et al.
Veröffentlicht: (2026)
Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It
von: Zhang, Yaxiang, et al.
Veröffentlicht: (2026)
von: Zhang, Yaxiang, et al.
Veröffentlicht: (2026)
Accelerating Transformer Inference for Translation via Parallel Decoding
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
Evolving LLMs' Self-Refinement Capability via Synergistic Training-Inference Optimization
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2025)
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2025)
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
von: Gu, Zhuohan, et al.
Veröffentlicht: (2024)
von: Gu, Zhuohan, et al.
Veröffentlicht: (2024)
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2025)
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2025)
FlashDecoding++: Faster Large Language Model Inference on GPUs
von: Hong, Ke, et al.
Veröffentlicht: (2023)
von: Hong, Ke, et al.
Veröffentlicht: (2023)
Characterization-Guided GPU Fault Resilience in NVIDIA MPS
von: Liu, Rixin, et al.
Veröffentlicht: (2026)
von: Liu, Rixin, et al.
Veröffentlicht: (2026)
Tokens for Learning, Tokens for Unlearning: Mitigating Membership Inference Attacks in Large Language Models via Dual-Purpose Training
von: Tran, Toan, et al.
Veröffentlicht: (2025)
von: Tran, Toan, et al.
Veröffentlicht: (2025)
Training-Inference Consistent Segmented Execution for Long-Context LLMs
von: Shang, Xianpeng, et al.
Veröffentlicht: (2026)
von: Shang, Xianpeng, et al.
Veröffentlicht: (2026)
VTC: DNN Compilation with Virtual Tensors for Data Movement Elimination
von: Hu, Muyan, et al.
Veröffentlicht: (2026)
von: Hu, Muyan, et al.
Veröffentlicht: (2026)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
von: He, Shwai, et al.
Veröffentlicht: (2025)
von: He, Shwai, et al.
Veröffentlicht: (2025)
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
Transparent Screening for LLM Inference and Training Impacts
von: Pachot, Arnault, et al.
Veröffentlicht: (2026)
von: Pachot, Arnault, et al.
Veröffentlicht: (2026)
MF-QAT: Multi-Format Quantization-Aware Training for Elastic Inference
von: Xu, Zifei, et al.
Veröffentlicht: (2026)
von: Xu, Zifei, et al.
Veröffentlicht: (2026)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
von: Liu, Di, et al.
Veröffentlicht: (2024)
von: Liu, Di, et al.
Veröffentlicht: (2024)
Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference
von: Liu, Yi
Veröffentlicht: (2026)
von: Liu, Yi
Veröffentlicht: (2026)
Online Cascade Learning for Efficient Inference over Streams
von: Nie, Lunyiu, et al.
Veröffentlicht: (2024)
von: Nie, Lunyiu, et al.
Veröffentlicht: (2024)
ComplexFormer: Disruptively Advancing Transformer Inference Ability via Head-Specific Complex Vector Attention
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
CoLLMLight: Cooperative Large Language Model Agents for Network-Wide Traffic Signal Control
von: Yuan, Zirui, et al.
Veröffentlicht: (2025)
von: Yuan, Zirui, et al.
Veröffentlicht: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
von: Li, Xing, et al.
Veröffentlicht: (2025)
von: Li, Xing, et al.
Veröffentlicht: (2025)
R$^2$PO: Decoupling Training Trajectories from Inference Responses for LLM Reasoning
von: Wang, Jingchu, et al.
Veröffentlicht: (2026)
von: Wang, Jingchu, et al.
Veröffentlicht: (2026)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
von: Xiao, Youshao, et al.
Veröffentlicht: (2023)
von: Xiao, Youshao, et al.
Veröffentlicht: (2023)
A Graph-Based Classical and Quantum Approach to Deterministic L-System Inference
von: Lotfi, Ali, et al.
Veröffentlicht: (2024)
von: Lotfi, Ali, et al.
Veröffentlicht: (2024)
Neuro-Symbolic Contrastive Learning for Cross-domain Inference
von: Liu, Mingyue, et al.
Veröffentlicht: (2025)
von: Liu, Mingyue, et al.
Veröffentlicht: (2025)
Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
von: Sohn, Gina, et al.
Veröffentlicht: (2025)
von: Sohn, Gina, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Defeating the Training-Inference Mismatch via FP16
von: Qi, Penghui, et al.
Veröffentlicht: (2025) -
Diagnosing Training Inference Mismatch in LLM Reinforcement Learning
von: Zhong, Tianle, et al.
Veröffentlicht: (2026) -
Communication Compression for Tensor Parallel LLM Inference
von: Hansen-Palmus, Jan, et al.
Veröffentlicht: (2024) -
RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers
von: Lu, Yifan, et al.
Veröffentlicht: (2025) -
Make Some Noise: Unlocking Language Model Parallel Inference Capability through Noisy Training
von: Wang, Yixuan, et al.
Veröffentlicht: (2024)