A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
Fuente:
arXiv
Salvato in:
| Autori principali: | Xi, Shaoke, Lao, ChonLam, Jia, Boyi, Gao, Jiaqi, Zhang, Zhipeng, Cao, Jiamin, Sutioso, Brian, Xu, Erci, Yu, Minlan, Ren, Kui, Li, Yong, Qian, Zhengping, Zhai, Ennan, Zhou, Jingren |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EdgeSight: Enabling Modeless and Cost-Efficient Inference at the Edge
di: Lao, ChonLam, et al.
Pubblicazione: (2024)
di: Lao, ChonLam, et al.
Pubblicazione: (2024)
TrainMover: An Interruption-Resilient Runtime for ML Training
di: Lao, ChonLam, et al.
Pubblicazione: (2024)
di: Lao, ChonLam, et al.
Pubblicazione: (2024)
Towards Easy and Realistic Network Infrastructure Testing for Large-scale Machine Learning
di: Yoo, Jinsun, et al.
Pubblicazione: (2025)
di: Yoo, Jinsun, et al.
Pubblicazione: (2025)
Cora: Accelerating Stateful Network Applications with SmartNICs
di: Xi, Shaoke, et al.
Pubblicazione: (2024)
di: Xi, Shaoke, et al.
Pubblicazione: (2024)
THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression
di: Li, Minghao, et al.
Pubblicazione: (2023)
di: Li, Minghao, et al.
Pubblicazione: (2023)
An LLM-based Agentic Framework for Accessible Network Control
di: Lin, Samuel, et al.
Pubblicazione: (2025)
di: Lin, Samuel, et al.
Pubblicazione: (2025)
Estudo fitossociológico de um trecho da floresta estacional semidecidual em Diamante do Norte, Estado do Paraná, Brasil
di: Erci Marcos Del Quiqui
Pubblicazione: (2007)
di: Erci Marcos Del Quiqui
Pubblicazione: (2007)
An Extensible Software Transport Layer for GPU Networking
di: Zhou, Yang, et al.
Pubblicazione: (2025)
di: Zhou, Yang, et al.
Pubblicazione: (2025)
Accelerating Compound LLM Training Workloads with Maestro
di: Yuan, Xiulong, et al.
Pubblicazione: (2026)
di: Yuan, Xiulong, et al.
Pubblicazione: (2026)
Orla: A Library for Serving LLM-Based Multi-Agent Systems
di: Shahout, Rana, et al.
Pubblicazione: (2026)
di: Shahout, Rana, et al.
Pubblicazione: (2026)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
di: Hankendi, Can, et al.
Pubblicazione: (2026)
di: Hankendi, Can, et al.
Pubblicazione: (2026)
Intra-request branch orchestration for efficient LLM reasoning
di: Jiang, Weifan, et al.
Pubblicazione: (2025)
di: Jiang, Weifan, et al.
Pubblicazione: (2025)
An Empirical Study of LLM Serving in Confidential GPUs
di: Park, Eunseong, et al.
Pubblicazione: (2026)
di: Park, Eunseong, et al.
Pubblicazione: (2026)
A Systematic Characterization of LLM Inference on GPUs
di: Wang, Haonan, et al.
Pubblicazione: (2025)
di: Wang, Haonan, et al.
Pubblicazione: (2025)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
di: Yu, Shan, et al.
Pubblicazione: (2025)
di: Yu, Shan, et al.
Pubblicazione: (2025)
Are LLM Decisions Faithful to Verbal Confidence?
di: Wang, Jiawei, et al.
Pubblicazione: (2026)
di: Wang, Jiawei, et al.
Pubblicazione: (2026)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
di: Jiang, Xuanlin, et al.
Pubblicazione: (2024)
di: Jiang, Xuanlin, et al.
Pubblicazione: (2024)
Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches
di: Fang, Shaoke, et al.
Pubblicazione: (2026)
di: Fang, Shaoke, et al.
Pubblicazione: (2026)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
di: Kim, Heehoon, et al.
Pubblicazione: (2026)
di: Kim, Heehoon, et al.
Pubblicazione: (2026)
Measurement of LLM's Philosophies of Human Nature
di: Ni, Minheng, et al.
Pubblicazione: (2025)
di: Ni, Minheng, et al.
Pubblicazione: (2025)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
di: Tamber, Manveer Singh, et al.
Pubblicazione: (2025)
di: Tamber, Manveer Singh, et al.
Pubblicazione: (2025)
LatencyPrism: Online Non-intrusive Latency Sculpting for SLO-Guaranteed LLM Inference
di: Du, Yin, et al.
Pubblicazione: (2026)
di: Du, Yin, et al.
Pubblicazione: (2026)
FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression
di: Tang, Zhenheng, et al.
Pubblicazione: (2024)
di: Tang, Zhenheng, et al.
Pubblicazione: (2024)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
di: Georganas, Evangelos, et al.
Pubblicazione: (2025)
di: Georganas, Evangelos, et al.
Pubblicazione: (2025)
EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
di: Kubwimana, Benjamin, et al.
Pubblicazione: (2025)
di: Kubwimana, Benjamin, et al.
Pubblicazione: (2025)
Scaled Block Vecchia Approximation for High-Dimensional Gaussian Process Emulation on GPUs
di: Pan, Qilong, et al.
Pubblicazione: (2025)
di: Pan, Qilong, et al.
Pubblicazione: (2025)
LLM-Slice: Dedicated Wireless Network Slicing for Large Language Models
di: Liu, Boyi, et al.
Pubblicazione: (2024)
di: Liu, Boyi, et al.
Pubblicazione: (2024)
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
di: Bhan, Milan, et al.
Pubblicazione: (2025)
di: Bhan, Milan, et al.
Pubblicazione: (2025)
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
di: Da, Wei, et al.
Pubblicazione: (2026)
di: Da, Wei, et al.
Pubblicazione: (2026)
Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation
di: Feng, Weiqi, et al.
Pubblicazione: (2024)
di: Feng, Weiqi, et al.
Pubblicazione: (2024)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2024)
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2024)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
di: Alon, Bar, et al.
Pubblicazione: (2026)
di: Alon, Bar, et al.
Pubblicazione: (2026)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
di: Mittal, Avni, et al.
Pubblicazione: (2026)
di: Mittal, Avni, et al.
Pubblicazione: (2026)
LLM-grounded Video Diffusion Models
di: Lian, Long, et al.
Pubblicazione: (2023)
di: Lian, Long, et al.
Pubblicazione: (2023)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
di: Karinshak, Elise, et al.
Pubblicazione: (2024)
di: Karinshak, Elise, et al.
Pubblicazione: (2024)
HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
NLP-AKG: Few-Shot Construction of NLP Academic Knowledge Graph Based on LLM
di: Lan, Jiayin, et al.
Pubblicazione: (2025)
di: Lan, Jiayin, et al.
Pubblicazione: (2025)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
di: Quan, Xin, et al.
Pubblicazione: (2025)
di: Quan, Xin, et al.
Pubblicazione: (2025)
Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and Method
di: Zhao, Tianzhe, et al.
Pubblicazione: (2026)
di: Zhao, Tianzhe, et al.
Pubblicazione: (2026)
Documenti analoghi
-
EdgeSight: Enabling Modeless and Cost-Efficient Inference at the Edge
di: Lao, ChonLam, et al.
Pubblicazione: (2024) -
TrainMover: An Interruption-Resilient Runtime for ML Training
di: Lao, ChonLam, et al.
Pubblicazione: (2024) -
Towards Easy and Realistic Network Infrastructure Testing for Large-scale Machine Learning
di: Yoo, Jinsun, et al.
Pubblicazione: (2025) -
Cora: Accelerating Stateful Network Applications with SmartNICs
di: Xi, Shaoke, et al.
Pubblicazione: (2024) -
THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression
di: Li, Minghao, et al.
Pubblicazione: (2023)