FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mei, Junyi, Sun, Shixuan, Li, Chao, Xu, Cheng, Chen, Cheng, Liu, Yibo, Wang, Jing, Zhao, Cheng, Hou, Xiaofeng, Guo, Minyi, He, Bingsheng, Cong, Xiaoliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
von: Park, Seongyeon, et al.
Veröffentlicht: (2025)
von: Park, Seongyeon, et al.
Veröffentlicht: (2025)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
A GPU Accelerated Temporal Window-Based Random Walk Sampler
von: Salehin, Md Ashfaq, et al.
Veröffentlicht: (2026)
von: Salehin, Md Ashfaq, et al.
Veröffentlicht: (2026)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
PilotANN: Memory-Bounded GPU Acceleration for Vector Search
von: Gui, Yuntao, et al.
Veröffentlicht: (2025)
von: Gui, Yuntao, et al.
Veröffentlicht: (2025)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026)
Towards Fast Setup and High Throughput of GPU Serverless Computing
von: Zhao, Han, et al.
Veröffentlicht: (2024)
von: Zhao, Han, et al.
Veröffentlicht: (2024)
Bingo: Radix-based Bias Factorization for Random Walk on Dynamic Graphs
von: Wang, Pinhuan, et al.
Veröffentlicht: (2025)
von: Wang, Pinhuan, et al.
Veröffentlicht: (2025)
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
Decentralized Federated Averaging via Random Walk
von: Wang, Changheng, et al.
Veröffentlicht: (2025)
von: Wang, Changheng, et al.
Veröffentlicht: (2025)
Heterogeneity-Aware Memory Efficient Federated Learning via Progressive Layer Freezing
von: Yebo, Wu, et al.
Veröffentlicht: (2024)
von: Yebo, Wu, et al.
Veröffentlicht: (2024)
Understanding the Landscape of Ampere GPU Memory Errors
von: Zhu, Zhu, et al.
Veröffentlicht: (2025)
von: Zhu, Zhu, et al.
Veröffentlicht: (2025)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
MemAscend: System Memory Optimization for SSD-Offloaded LLM Fine-Tuning
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
Analysis and Optimized CXL-Attached Memory Allocation for Long-Context LLM Fine-Tuning
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
Byzantine-Tolerant Consensus in GPU-Inspired Shared Memory
von: Georgiou, Chryssis, et al.
Veröffentlicht: (2025)
von: Georgiou, Chryssis, et al.
Veröffentlicht: (2025)
Distributed Graph Embedding with Information-Oriented Random Walks
von: Fang, Peng, et al.
Veröffentlicht: (2023)
von: Fang, Peng, et al.
Veröffentlicht: (2023)
Incremental GNN Embedding Computation on Streaming Graphs
von: Wang, Qiange, et al.
Veröffentlicht: (2026)
von: Wang, Qiange, et al.
Veröffentlicht: (2026)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
SkyMemory: A LEO Edge Cache for Transformer Inference Optimization and Scale Out
von: Sandholm, Thomas, et al.
Veröffentlicht: (2025)
von: Sandholm, Thomas, et al.
Veröffentlicht: (2025)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
DASH: Deterministic Attention Scheduling for High-throughput Reproducible LLM Training
von: Qiang, Xinwei, et al.
Veröffentlicht: (2026)
von: Qiang, Xinwei, et al.
Veröffentlicht: (2026)
Efficient Accelerated Graph Edit Distance Computation on GPU
von: Dabah, Adel, et al.
Veröffentlicht: (2026)
von: Dabah, Adel, et al.
Veröffentlicht: (2026)
GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations
von: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Veröffentlicht: (2026)
von: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Veröffentlicht: (2026)
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
von: Kumar, Abhishek Vijaya, et al.
Veröffentlicht: (2024)
von: Kumar, Abhishek Vijaya, et al.
Veröffentlicht: (2024)
Towards the Distributed Large-scale k-NN Graph Construction by Graph Merge
von: Zhang, Cheng, et al.
Veröffentlicht: (2025)
von: Zhang, Cheng, et al.
Veröffentlicht: (2025)
GCAPS: GPU Context-Aware Preemptive Priority-based Scheduling for Real-Time Tasks
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
von: Tang, Peng, et al.
Veröffentlicht: (2024)
von: Tang, Peng, et al.
Veröffentlicht: (2024)
FastGraph: Optimized GPU-Enabled Algorithms for Fast Graph Building and Message Passing
von: Agarwal, Aarush, et al.
Veröffentlicht: (2025)
von: Agarwal, Aarush, et al.
Veröffentlicht: (2025)
SOLANET: Distributed Neighbor Graph Construction on GPU-Accelerated Systems
von: Iwabuchi, Keita, et al.
Veröffentlicht: (2026)
von: Iwabuchi, Keita, et al.
Veröffentlicht: (2026)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
AME: An Efficient Heterogeneous Agentic Memory Engine for Smartphones
von: Zhao, Xinkui, et al.
Veröffentlicht: (2025)
von: Zhao, Xinkui, et al.
Veröffentlicht: (2025)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
von: Chang, Zihan, et al.
Veröffentlicht: (2024)
von: Chang, Zihan, et al.
Veröffentlicht: (2024)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
von: Lin, Shouxu, et al.
Veröffentlicht: (2026)
von: Lin, Shouxu, et al.
Veröffentlicht: (2026)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
HeteroSTA: A CPU-GPU Heterogeneous Static Timing Analysis Engine with Holistic Industrial Design Support
von: Guo, Zizheng, et al.
Veröffentlicht: (2025)
von: Guo, Zizheng, et al.
Veröffentlicht: (2025)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
von: Park, Seongyeon, et al.
Veröffentlicht: (2025) -
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
von: Guo, Cong, et al.
Veröffentlicht: (2024) -
A GPU Accelerated Temporal Window-Based Random Walk Sampler
von: Salehin, Md Ashfaq, et al.
Veröffentlicht: (2026) -
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026) -
PilotANN: Memory-Bounded GPU Acceleration for Vector Search
von: Gui, Yuntao, et al.
Veröffentlicht: (2025)