mLR: Scalable Laminography Reconstruction based on Memoization
Fuente:
arXiv
Salvato in:
| Autori principali: | Ma, Bin, Nikitin, Viktor, Wang, Xi, Bicer, Tekin, Li, Dong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
di: Lin, Mao, et al.
Pubblicazione: (2026)
di: Lin, Mao, et al.
Pubblicazione: (2026)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
di: Ma, Bin, et al.
Pubblicazione: (2026)
di: Ma, Bin, et al.
Pubblicazione: (2026)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
di: Shan, Baodi, et al.
Pubblicazione: (2024)
di: Shan, Baodi, et al.
Pubblicazione: (2024)
Scalable GPU Performance Variability Analysis framework
di: Lahiry, Ankur, et al.
Pubblicazione: (2025)
di: Lahiry, Ankur, et al.
Pubblicazione: (2025)
Toward Scalable Docker-Based Emulations of Blockchain Networks for Research and Development
di: Pennino, Diego, et al.
Pubblicazione: (2024)
di: Pennino, Diego, et al.
Pubblicazione: (2024)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
di: Ramesh, Risshab Srinivas
Pubblicazione: (2024)
di: Ramesh, Risshab Srinivas
Pubblicazione: (2024)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
di: Curless, Brian, et al.
Pubblicazione: (2025)
di: Curless, Brian, et al.
Pubblicazione: (2025)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
di: Maurya, Avinash, et al.
Pubblicazione: (2026)
di: Maurya, Avinash, et al.
Pubblicazione: (2026)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
di: Wang, Yidi, et al.
Pubblicazione: (2024)
di: Wang, Yidi, et al.
Pubblicazione: (2024)
DIAL: Decentralized I/O AutoTuning via Learned Client-side Local Metrics for Parallel File System
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
di: Liu, Hang, et al.
Pubblicazione: (2025)
di: Liu, Hang, et al.
Pubblicazione: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
CARAT: Client-Side Adaptive RPC and Cache Co-Tuning for Parallel File Systems
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
di: Zhuang, Chen, et al.
Pubblicazione: (2024)
di: Zhuang, Chen, et al.
Pubblicazione: (2024)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
di: Raffin, Guillaume, et al.
Pubblicazione: (2024)
di: Raffin, Guillaume, et al.
Pubblicazione: (2024)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
FluxSieve: Unifying Streaming and Analytical Data Planes for Scalable Cloud Observability
di: Vogel, Adriano, et al.
Pubblicazione: (2026)
di: Vogel, Adriano, et al.
Pubblicazione: (2026)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
di: Zhao, Yanbo, et al.
Pubblicazione: (2025)
di: Zhao, Yanbo, et al.
Pubblicazione: (2025)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
di: Xu, Jingwei, et al.
Pubblicazione: (2025)
di: Xu, Jingwei, et al.
Pubblicazione: (2025)
An Experimental Study of Different Aggregation Schemes in Semi-Asynchronous Federated Learning
di: Li, Yunbo, et al.
Pubblicazione: (2024)
di: Li, Yunbo, et al.
Pubblicazione: (2024)
Optimal Parallel Scheduling under Concave Speedup Functions
di: Li, Chengzhang, et al.
Pubblicazione: (2025)
di: Li, Chengzhang, et al.
Pubblicazione: (2025)
How to Rent GPUs on a Budget
di: Li, Zhouzi, et al.
Pubblicazione: (2024)
di: Li, Zhouzi, et al.
Pubblicazione: (2024)
Asymptotically Optimal Scheduling of Multiple Parallelizable Job Classes
di: Berg, Benjamin, et al.
Pubblicazione: (2024)
di: Berg, Benjamin, et al.
Pubblicazione: (2024)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
di: Liu, Shifang, et al.
Pubblicazione: (2025)
di: Liu, Shifang, et al.
Pubblicazione: (2025)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
di: Wang, Chen, et al.
Pubblicazione: (2025)
di: Wang, Chen, et al.
Pubblicazione: (2025)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
di: Debnath, Shimul, et al.
Pubblicazione: (2026)
di: Debnath, Shimul, et al.
Pubblicazione: (2026)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
di: Wang, Xi, et al.
Pubblicazione: (2024)
di: Wang, Xi, et al.
Pubblicazione: (2024)
Shifting the Sweet Spot: High-Performance Matrix-Free Method for High-Order Elasticity
di: Chang, Dali, et al.
Pubblicazione: (2026)
di: Chang, Dali, et al.
Pubblicazione: (2026)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
di: Xu, Guanyu, et al.
Pubblicazione: (2025)
di: Xu, Guanyu, et al.
Pubblicazione: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
di: Ather, Hammad, et al.
Pubblicazione: (2024)
di: Ather, Hammad, et al.
Pubblicazione: (2024)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
di: Zhang, Yaozheng, et al.
Pubblicazione: (2025)
di: Zhang, Yaozheng, et al.
Pubblicazione: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
di: Wang, Yuxin, et al.
Pubblicazione: (2023)
di: Wang, Yuxin, et al.
Pubblicazione: (2023)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
di: Cornelius, Melanie, et al.
Pubblicazione: (2025)
di: Cornelius, Melanie, et al.
Pubblicazione: (2025)
Profiling and optimization of multi-card GPU machine learning jobs
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
WebAssembly and Unikernels: A Comparative Study for Serverless at the Edge
di: Besozzi, Valerio, et al.
Pubblicazione: (2025)
di: Besozzi, Valerio, et al.
Pubblicazione: (2025)
Bridding OT and PaaS in Edge-to-Cloud Continuum
di: Barrios, Carlos J, et al.
Pubblicazione: (2025)
di: Barrios, Carlos J, et al.
Pubblicazione: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
di: Karfakis, George, et al.
Pubblicazione: (2025)
di: Karfakis, George, et al.
Pubblicazione: (2025)
Documenti analoghi
-
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
di: Lin, Mao, et al.
Pubblicazione: (2026) -
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
di: Ma, Bin, et al.
Pubblicazione: (2026) -
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
di: Shan, Baodi, et al.
Pubblicazione: (2024) -
Scalable GPU Performance Variability Analysis framework
di: Lahiry, Ankur, et al.
Pubblicazione: (2025) -
Toward Scalable Docker-Based Emulations of Blockchain Networks for Research and Development
di: Pennino, Diego, et al.
Pubblicazione: (2024)