Distributed Order Recording Techniques for Efficient Record-and-Replay of Multi-threaded Programs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fu, Xiang, Meng, Shiman, Zhang, Weiping, Guo, Luanzheng, Sato, Kento, Ahn, Dong H., Laguna, Ignacio, Lee, Gregory L., Schulz, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scrutinizing Variables for Checkpoint Using Automatic Differentiation
von: Huang, Xin, et al.
Veröffentlicht: (2026)
von: Huang, Xin, et al.
Veröffentlicht: (2026)
LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
von: Sun, Minqiu, et al.
Veröffentlicht: (2026)
von: Sun, Minqiu, et al.
Veröffentlicht: (2026)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
Ira: Efficient Transaction Replay for Distributed Systems
von: Bhat, Adithya, et al.
Veröffentlicht: (2026)
von: Bhat, Adithya, et al.
Veröffentlicht: (2026)
Tracing Distributed Algorithms Using Replay Clocks
von: Lagwankar, Ishaan
Veröffentlicht: (2024)
von: Lagwankar, Ishaan
Veröffentlicht: (2024)
On The Reproducibility Limitations of RAG Systems
von: Wang, Baiqiang, et al.
Veröffentlicht: (2025)
von: Wang, Baiqiang, et al.
Veröffentlicht: (2025)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
von: Proaño, Andrès Rubio, et al.
Veröffentlicht: (2024)
von: Proaño, Andrès Rubio, et al.
Veröffentlicht: (2024)
Distributed Record Linkage in Healthcare Data with Apache Spark
von: Heydari, Mohammad, et al.
Veröffentlicht: (2024)
von: Heydari, Mohammad, et al.
Veröffentlicht: (2024)
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
von: Nichols, Daniel, et al.
Veröffentlicht: (2026)
von: Nichols, Daniel, et al.
Veröffentlicht: (2026)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
Towards Efficient Replay in Federated Incremental Learning
von: Li, Yichen, et al.
Veröffentlicht: (2024)
von: Li, Yichen, et al.
Veröffentlicht: (2024)
MPI Malleability Validation under Replayed Real-World HPC Conditions
von: Iserte, S., et al.
Veröffentlicht: (2026)
von: Iserte, S., et al.
Veröffentlicht: (2026)
Towards Efficient and Scalable Distributed Vector Search with RDMA
von: Zhi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhi, Xiangyu, et al.
Veröffentlicht: (2025)
Trace Replay Simulation of MIT SuperCloud for Studying Optimal Sustainability Policies
von: Brewer, Wesley, et al.
Veröffentlicht: (2025)
von: Brewer, Wesley, et al.
Veröffentlicht: (2025)
NLP-Guided Synthesis: Transitioning from Sequential Programs to Distributed Programs
von: Sanjel, Arun, et al.
Veröffentlicht: (2024)
von: Sanjel, Arun, et al.
Veröffentlicht: (2024)
MoLink: Distributed and Efficient Serving Framework for Large Models
von: Jin, Lewei, et al.
Veröffentlicht: (2025)
von: Jin, Lewei, et al.
Veröffentlicht: (2025)
Design Principles of Dynamic Resource Management for High-Performance Parallel Programming Models
von: Huber, Dominik, et al.
Veröffentlicht: (2024)
von: Huber, Dominik, et al.
Veröffentlicht: (2024)
An Explorative Study on Distributed Computing Techniques in Training and Inference of Large Language Models
von: Hakim, Sheikh Azizul, et al.
Veröffentlicht: (2025)
von: Hakim, Sheikh Azizul, et al.
Veröffentlicht: (2025)
Communication and Energy Efficient Federated Learning using Zero-Order Optimization Technique
von: Mhanna, Elissa, et al.
Veröffentlicht: (2024)
von: Mhanna, Elissa, et al.
Veröffentlicht: (2024)
HiCR, an Abstract Model for Distributed Heterogeneous Programming
von: Martin, Sergio Miguel, et al.
Veröffentlicht: (2025)
von: Martin, Sergio Miguel, et al.
Veröffentlicht: (2025)
Performance Trade-offs of High Order Meshless Approximation on Distributed Memory Systems
von: Vehovar, Jon, et al.
Veröffentlicht: (2025)
von: Vehovar, Jon, et al.
Veröffentlicht: (2025)
CXL Shared Memory Programming: Barely Distributed and Almost Persistent
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
Efficient Distributed MLLM Training with Cornstarch
von: Jang, Insu, et al.
Veröffentlicht: (2025)
von: Jang, Insu, et al.
Veröffentlicht: (2025)
MTGenRec: An Efficient Distributed Training System for Generative Recommendation Models in Meituan
von: Wang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wang, Yuxiang, et al.
Veröffentlicht: (2025)
Vortex: Efficient Sample-Free Dynamic Tensor Program Optimization via Hardware-aware Strategy Space Hierarchization
von: Zhou, Yangjie, et al.
Veröffentlicht: (2024)
von: Zhou, Yangjie, et al.
Veröffentlicht: (2024)
UniFaaS: Programming across Distributed Cyberinfrastructure with Federated Function Serving
von: Li, Yifei, et al.
Veröffentlicht: (2024)
von: Li, Yifei, et al.
Veröffentlicht: (2024)
ParaLog: Consistent Host-side Logging for Parallel Checkpoints
von: Chien, Steven W. D., et al.
Veröffentlicht: (2024)
von: Chien, Steven W. D., et al.
Veröffentlicht: (2024)
DeFT: Mitigating Data Dependencies for Flexible Communication Scheduling in Distributed Training
von: Meng, Lin, et al.
Veröffentlicht: (2025)
von: Meng, Lin, et al.
Veröffentlicht: (2025)
Locality, Not Spectral Mixing, Governs Direct Propagation in Distributed Offline Dynamic Programming
von: Shihab, Ibne Farabi
Veröffentlicht: (2026)
von: Shihab, Ibne Farabi
Veröffentlicht: (2026)
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
von: Zheng, Size, et al.
Veröffentlicht: (2025)
von: Zheng, Size, et al.
Veröffentlicht: (2025)
Toward Heterogeneous, Distributed, and Energy-Efficient Computing with SYCL
von: Cosenza, Biagio, et al.
Veröffentlicht: (2025)
von: Cosenza, Biagio, et al.
Veröffentlicht: (2025)
Adjusted Objects: An Efficient and Principled Approach to Scalable Programming (Extended Version)
von: Kane, Boubacar, et al.
Veröffentlicht: (2025)
von: Kane, Boubacar, et al.
Veröffentlicht: (2025)
VersaSlot: Efficient Fine-grained FPGA Sharing with Big.Little Slots and Live Migration in FPGA Cluster
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
An Integrated (Crop Model, Cloud and Big Data Analytic) Framework to support Agriculture Activity Monitoring System
von: Akhter, Shamim, et al.
Veröffentlicht: (2024)
von: Akhter, Shamim, et al.
Veröffentlicht: (2024)
Object Proxy Patterns for Accelerating Distributed Applications
von: Pauloski, J. Gregory, et al.
Veröffentlicht: (2024)
von: Pauloski, J. Gregory, et al.
Veröffentlicht: (2024)
TF-DDRL: A Transformer-enhanced Distributed DRL Technique for Scheduling IoT Applications in Edge and Cloud Computing Environments
von: Wang, Zhiyu, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyu, et al.
Veröffentlicht: (2024)
Using Diffusion Models as Generative Replay in Continual Federated Learning -- What will Happen?
von: Mei, Yongsheng, et al.
Veröffentlicht: (2024)
von: Mei, Yongsheng, et al.
Veröffentlicht: (2024)
Efficient Distributed Algorithms for Shape Reduction via Reconfigurable Circuits
von: Almalki, Nada, et al.
Veröffentlicht: (2025)
von: Almalki, Nada, et al.
Veröffentlicht: (2025)
Efficient and Portable Support for Overdecomposition on Distributed Memory GPGPU Platforms
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scrutinizing Variables for Checkpoint Using Automatic Differentiation
von: Huang, Xin, et al.
Veröffentlicht: (2026) -
LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
von: Sun, Minqiu, et al.
Veröffentlicht: (2026) -
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026) -
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
von: Mehboob, Talha, et al.
Veröffentlicht: (2025) -
Ira: Efficient Transaction Replay for Distributed Systems
von: Bhat, Adithya, et al.
Veröffentlicht: (2026)