ML-based Adaptive Prefetching and Data Placement for US HEP Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Karanam, Venkat Sai Suman Lamba, Barla, Sarat Sasank, Ramamurthy, Byrav, Weitzel, Derek |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Open Science Data Federation -- operation and monitoring
von: Andrijauskas, Fabio, et al.
Veröffentlicht: (2026)
von: Andrijauskas, Fabio, et al.
Veröffentlicht: (2026)
Adventures with Grace Hopper AI Super Chip and the National Research Platform
von: Hurt, J. Alex, et al.
Veröffentlicht: (2024)
von: Hurt, J. Alex, et al.
Veröffentlicht: (2024)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026)
Prefetching in Deep Memory Hierarchies with NVRAM as Main Memory
von: Lurbe, Manel, et al.
Veröffentlicht: (2025)
von: Lurbe, Manel, et al.
Veröffentlicht: (2025)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
FedFetch: Faster Federated Learning with Adaptive Downstream Prefetching
von: Yan, Qifan, et al.
Veröffentlicht: (2025)
von: Yan, Qifan, et al.
Veröffentlicht: (2025)
Lessons Learned Migrating CUDA to SYCL: A HEP Case Study with ROOT RDataFrame
von: Chen, Jolly, et al.
Veröffentlicht: (2024)
von: Chen, Jolly, et al.
Veröffentlicht: (2024)
Generalized Data Placement Strategies for Racetrack Memories
von: Khan, Asif Ali, et al.
Veröffentlicht: (2019)
von: Khan, Asif Ali, et al.
Veröffentlicht: (2019)
PROBE: Co-Balancing Computation and Communication in MoE Inference via Real-Time Predictive Prefetching
von: Zhu, Qianchao, et al.
Veröffentlicht: (2026)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2026)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
von: Raj, Suman, et al.
Veröffentlicht: (2024)
von: Raj, Suman, et al.
Veröffentlicht: (2024)
Declarative Data Pipeline for Large Scale ML Services
von: Yang, Yunzhao, et al.
Veröffentlicht: (2025)
von: Yang, Yunzhao, et al.
Veröffentlicht: (2025)
NetSenseML: Network-Adaptive Compression for Efficient Distributed Machine Learning
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures
von: Suman, Shekhar, et al.
Veröffentlicht: (2026)
von: Suman, Shekhar, et al.
Veröffentlicht: (2026)
Placement of Microservices-based IoT Applications in Fog Computing: A Taxonomy and Future Directions
von: Pallewatta, Samodha, et al.
Veröffentlicht: (2022)
von: Pallewatta, Samodha, et al.
Veröffentlicht: (2022)
Snowpark: Performant, Secure, User-Friendly Data Engineering and AI/ML Next To Your Data
von: Baker, Brandon, et al.
Veröffentlicht: (2025)
von: Baker, Brandon, et al.
Veröffentlicht: (2025)
FLASH Viterbi: Fast and Adaptive Viterbi Decoding for Modern Data Systems
von: Deng, Ziheng, et al.
Veröffentlicht: (2025)
von: Deng, Ziheng, et al.
Veröffentlicht: (2025)
AeroDaaS: A Programmable Drones-as-a-Service Platform for Intelligent Aerial Systems
von: Astu, Kautuk, et al.
Veröffentlicht: (2026)
von: Astu, Kautuk, et al.
Veröffentlicht: (2026)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
Optimal Workload Placement on Multi-Instance GPUs
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
Multi-core & GPU-based Balanced Butterfly Counting in Signed Bipartite Graphs
von: Kiran, Mekala, et al.
Veröffentlicht: (2026)
von: Kiran, Mekala, et al.
Veröffentlicht: (2026)
HPAC-ML: A Programming Model for Embedding ML Surrogates in Scientific Applications
von: Fink, Zane, et al.
Veröffentlicht: (2024)
von: Fink, Zane, et al.
Veröffentlicht: (2024)
An Efficient and Adaptive Watermark Detection System with Tile-based Error Correction
von: Zhong, Xinrui, et al.
Veröffentlicht: (2025)
von: Zhong, Xinrui, et al.
Veröffentlicht: (2025)
Optimizing Service Placement in Edge-to-Cloud AR/VR Systems using a Multi-Objective Genetic Algorithm
von: Herabad, Mohammadsadeq Garshasbi, et al.
Veröffentlicht: (2024)
von: Herabad, Mohammadsadeq Garshasbi, et al.
Veröffentlicht: (2024)
The National Research Platform: Stretched, Multi-Tenant, Scientific Kubernetes Cluster
von: Weitzel, Derek, et al.
Veröffentlicht: (2025)
von: Weitzel, Derek, et al.
Veröffentlicht: (2025)
Energy-aware Distributed Microservice Request Placement at the Edge
von: Toczé, Klervie, et al.
Veröffentlicht: (2024)
von: Toczé, Klervie, et al.
Veröffentlicht: (2024)
Energy Metrics for Edge Microservice Request Placement Strategies
von: Toczé, Klervie, et al.
Veröffentlicht: (2025)
von: Toczé, Klervie, et al.
Veröffentlicht: (2025)
It's the People, Not the Placement: Rethinking Allocations in Post-Moore Clouds
von: Harith, Tejas, et al.
Veröffentlicht: (2025)
von: Harith, Tejas, et al.
Veröffentlicht: (2025)
ML-based Modeling to Predict I/O Performance on Different Storage Sub-systems
von: Xu, Yiheng, et al.
Veröffentlicht: (2023)
von: Xu, Yiheng, et al.
Veröffentlicht: (2023)
A Multi-Objective Framework for Optimizing GPU-Enabled VM Placement in Cloud Data Centers with Multi-Instance GPU Technology
von: Siavashi, Ahmad, et al.
Veröffentlicht: (2025)
von: Siavashi, Ahmad, et al.
Veröffentlicht: (2025)
Performance and Security Aware Distributed Service Placement in Fog Computing
von: Goudarzi, Mohammad, et al.
Veröffentlicht: (2026)
von: Goudarzi, Mohammad, et al.
Veröffentlicht: (2026)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
Scalable mRMR feature selection to handle high dimensional datasets: Vertical partitioning based Iterative MapReduce framework
von: Vivek, Yelleti, et al.
Veröffentlicht: (2022)
von: Vivek, Yelleti, et al.
Veröffentlicht: (2022)
A Bring-Your-Own-Model Approach for ML-Driven Storage Placement in Warehouse-Scale Computers
von: Yang, Chenxi, et al.
Veröffentlicht: (2025)
von: Yang, Chenxi, et al.
Veröffentlicht: (2025)
INSPIRIT: Optimizing Heterogeneous Task Scheduling through Adaptive Priority in Task-based Runtime Systems
von: Wang, Yiqing, et al.
Veröffentlicht: (2024)
von: Wang, Yiqing, et al.
Veröffentlicht: (2024)
POSEIDON : Efficient Function Placement at the Edge using Deep Reinforcement Learning
von: Jain, Prakhar, et al.
Veröffentlicht: (2024)
von: Jain, Prakhar, et al.
Veröffentlicht: (2024)
Power Aware Container Placement in Cloud Computing with Affinity and Cubic Power Model
von: Sarkar, Suvarthi, et al.
Veröffentlicht: (2024)
von: Sarkar, Suvarthi, et al.
Veröffentlicht: (2024)
A More Scalable Sparse Dynamic Data Exchange
von: Geyko, Andrew, et al.
Veröffentlicht: (2023)
von: Geyko, Andrew, et al.
Veröffentlicht: (2023)
Self-healing Nodes with Adaptive Data-Sharding
von: Thakur, Ayush, et al.
Veröffentlicht: (2024)
von: Thakur, Ayush, et al.
Veröffentlicht: (2024)
LEISA: A Scalable Microservice-based System for Efficient Livestock Data Sharing
von: Habib, Mahir, et al.
Veröffentlicht: (2025)
von: Habib, Mahir, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Open Science Data Federation -- operation and monitoring
von: Andrijauskas, Fabio, et al.
Veröffentlicht: (2026) -
Adventures with Grace Hopper AI Super Chip and the National Research Platform
von: Hurt, J. Alex, et al.
Veröffentlicht: (2024) -
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026) -
Prefetching in Deep Memory Hierarchies with NVRAM as Main Memory
von: Lurbe, Manel, et al.
Veröffentlicht: (2025) -
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)