FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
Fuente:
arXiv
Guardado en:
| Autores principales: | Park, Seongyeon, Song, Jaeyong, Shin, Changmin, Kim, Sukjin, Hong, Junguk, Lee, Jinho |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read Mapping
por: Park, Seongyeon, et al.
Publicado: (2024)
por: Park, Seongyeon, et al.
Publicado: (2024)
PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
por: Kim, Sukjin, et al.
Publicado: (2025)
por: Kim, Sukjin, et al.
Publicado: (2025)
GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
por: Song, Jaeyong, et al.
Publicado: (2026)
por: Song, Jaeyong, et al.
Publicado: (2026)
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
por: Song, Jaeyong, et al.
Publicado: (2026)
por: Song, Jaeyong, et al.
Publicado: (2026)
FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk Framework
por: Mei, Junyi, et al.
Publicado: (2024)
por: Mei, Junyi, et al.
Publicado: (2024)
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
por: Noh, Si Ung, et al.
Publicado: (2024)
por: Noh, Si Ung, et al.
Publicado: (2024)
GraNNDis: Efficient Unified Distributed Training Framework for Deep GNNs on Large Clusters
por: Song, Jaeyong, et al.
Publicado: (2023)
por: Song, Jaeyong, et al.
Publicado: (2023)
A GPU Accelerated Temporal Window-Based Random Walk Sampler
por: Salehin, Md Ashfaq, et al.
Publicado: (2026)
por: Salehin, Md Ashfaq, et al.
Publicado: (2026)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
por: Ding, Zhimin, et al.
Publicado: (2024)
por: Ding, Zhimin, et al.
Publicado: (2024)
Bingo: Radix-based Bias Factorization for Random Walk on Dynamic Graphs
por: Wang, Pinhuan, et al.
Publicado: (2025)
por: Wang, Pinhuan, et al.
Publicado: (2025)
GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems
por: Shan, Baodi, et al.
Publicado: (2026)
por: Shan, Baodi, et al.
Publicado: (2026)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
por: Qianli, Liu, et al.
Publicado: (2025)
por: Qianli, Liu, et al.
Publicado: (2025)
Decentralized Federated Averaging via Random Walk
por: Wang, Changheng, et al.
Publicado: (2025)
por: Wang, Changheng, et al.
Publicado: (2025)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
por: Lee, Munkyu, et al.
Publicado: (2024)
por: Lee, Munkyu, et al.
Publicado: (2024)
GTaP: A GPU-Resident Fork-Join Task-Parallel Runtime with a Pragma-Based Interface
por: Maeda, Yuki, et al.
Publicado: (2026)
por: Maeda, Yuki, et al.
Publicado: (2026)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
por: Lin, Yi-Chien, et al.
Publicado: (2024)
por: Lin, Yi-Chien, et al.
Publicado: (2024)
SmartQC: An Extensible DLT-Based Framework for Trusted Data Workflows in Smart Manufacturing
por: McGibney, Alan, et al.
Publicado: (2024)
por: McGibney, Alan, et al.
Publicado: (2024)
On the Runtime of Local Mutual Exclusion for Anonymous Dynamic Networks
por: Chaturvedi, Anya, et al.
Publicado: (2025)
por: Chaturvedi, Anya, et al.
Publicado: (2025)
Pipette: Automatic Fine-grained Large Language Model Training Configurator for Real-World Clusters
por: Yim, Jinkyu, et al.
Publicado: (2024)
por: Yim, Jinkyu, et al.
Publicado: (2024)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
por: Wang, Tianyu, et al.
Publicado: (2024)
por: Wang, Tianyu, et al.
Publicado: (2024)
Stardust: A Scalable and Extensible Simulator for the 3D Continuum
por: Pusztai, Thomas, et al.
Publicado: (2025)
por: Pusztai, Thomas, et al.
Publicado: (2025)
GPU-Accelerated Batch-Dynamic Subgraph Matching
por: Qiu, Linshan, et al.
Publicado: (2024)
por: Qiu, Linshan, et al.
Publicado: (2024)
GreenDyGNN: Runtime-Adaptive Energy-Efficient Communication for Distributed GNN Training
por: Niam, Arefin, et al.
Publicado: (2026)
por: Niam, Arefin, et al.
Publicado: (2026)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
por: Yarlagadda, Srihas, et al.
Publicado: (2025)
por: Yarlagadda, Srihas, et al.
Publicado: (2025)
Heat: Satellite's meat is GPU's poison
por: Yuan, Zhehu, et al.
Publicado: (2024)
por: Yuan, Zhehu, et al.
Publicado: (2024)
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
por: Yang, Zhuoping, et al.
Publicado: (2025)
por: Yang, Zhuoping, et al.
Publicado: (2025)
Efficient Accelerated Graph Edit Distance Computation on GPU
por: Dabah, Adel, et al.
Publicado: (2026)
por: Dabah, Adel, et al.
Publicado: (2026)
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
por: Jiang, Youhe, et al.
Publicado: (2026)
por: Jiang, Youhe, et al.
Publicado: (2026)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
por: Gu, Jianfeng, et al.
Publicado: (2025)
por: Gu, Jianfeng, et al.
Publicado: (2025)
Leveraging Mathematical Reasoning of LLMs for Efficient GPU Thread Mapping
por: Maureira, Jose, et al.
Publicado: (2026)
por: Maureira, Jose, et al.
Publicado: (2026)
A Framework for Fine-Grained Synchronization of Dependent GPU Kernels
por: Jangda, Abhinav, et al.
Publicado: (2023)
por: Jangda, Abhinav, et al.
Publicado: (2023)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
por: Yu, Minchen, et al.
Publicado: (2025)
por: Yu, Minchen, et al.
Publicado: (2025)
A Multi-Objective Framework for Optimizing GPU-Enabled VM Placement in Cloud Data Centers with Multi-Instance GPU Technology
por: Siavashi, Ahmad, et al.
Publicado: (2025)
por: Siavashi, Ahmad, et al.
Publicado: (2025)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
por: Zhang, WenZheng, et al.
Publicado: (2024)
por: Zhang, WenZheng, et al.
Publicado: (2024)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
por: Li, Zhonggen, et al.
Publicado: (2025)
por: Li, Zhonggen, et al.
Publicado: (2025)
Automatic Tracing in Task-Based Runtime Systems
por: Yadav, Rohan, et al.
Publicado: (2024)
por: Yadav, Rohan, et al.
Publicado: (2024)
An AI-Native Runtime for Multi-Wearable Environments
por: Min, Chulhong, et al.
Publicado: (2024)
por: Min, Chulhong, et al.
Publicado: (2024)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
por: Saurez, Enrique, et al.
Publicado: (2024)
por: Saurez, Enrique, et al.
Publicado: (2024)
DuaLip-GPU Technical Report
por: Dexter, Gregory, et al.
Publicado: (2026)
por: Dexter, Gregory, et al.
Publicado: (2026)
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
por: Huang, Jiajun, et al.
Publicado: (2023)
por: Huang, Jiajun, et al.
Publicado: (2023)
Ejemplares similares
-
AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read Mapping
por: Park, Seongyeon, et al.
Publicado: (2024) -
PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
por: Kim, Sukjin, et al.
Publicado: (2025) -
GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
por: Song, Jaeyong, et al.
Publicado: (2026) -
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
por: Song, Jaeyong, et al.
Publicado: (2026) -
FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk Framework
por: Mei, Junyi, et al.
Publicado: (2024)