InkStream: Real-time GNN Inference on Streaming Graphs via Incremental Update
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Dan, Li, Zhaoying, Mitra, Tulika |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2024)
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2024)
Ripple: Scalable Incremental GNN Inferencing on Large Streaming Graphs
von: Naman, Pranjal, et al.
Veröffentlicht: (2025)
von: Naman, Pranjal, et al.
Veröffentlicht: (2025)
Incremental GNN Embedding Computation on Streaming Graphs
von: Wang, Qiange, et al.
Veröffentlicht: (2026)
von: Wang, Qiange, et al.
Veröffentlicht: (2026)
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
von: Chen, David, et al.
Veröffentlicht: (2026)
von: Chen, David, et al.
Veröffentlicht: (2026)
Multi-DNN Inference of Sparse Models on Edge SoCs
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
Reliable Microservice Tail Latency Prediction via Decoupled Dual-Stream Learning and Gradient Modulation
von: Qian, Wenzhuo, et al.
Veröffentlicht: (2025)
von: Qian, Wenzhuo, et al.
Veröffentlicht: (2025)
Multi-Dimensional Autoscaling of Stream Processing Services on Edge Devices
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
Towards Stream-Based Monitoring for EVM Networks
von: Onica, Emanuel, et al.
Veröffentlicht: (2025)
von: Onica, Emanuel, et al.
Veröffentlicht: (2025)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
von: Xu, Guanyu, et al.
Veröffentlicht: (2025)
von: Xu, Guanyu, et al.
Veröffentlicht: (2025)
KVDirect: Distributed Disaggregated LLM Inference
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
SDT-GNN: Streaming-based Distributed Training Framework for Graph Neural Networks
von: Huang, Xin, et al.
Veröffentlicht: (2024)
von: Huang, Xin, et al.
Veröffentlicht: (2024)
DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference
von: Zhang, Yujie, et al.
Veröffentlicht: (2024)
von: Zhang, Yujie, et al.
Veröffentlicht: (2024)
MQ-GNN: A Multi-Queue Pipelined Architecture for Scalable and Efficient GNN Training
von: Ullah, Irfan, et al.
Veröffentlicht: (2026)
von: Ullah, Irfan, et al.
Veröffentlicht: (2026)
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference
von: Hamadanian, Pouya, et al.
Veröffentlicht: (2025)
von: Hamadanian, Pouya, et al.
Veröffentlicht: (2025)
IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
von: Ghafouri, Saeid, et al.
Veröffentlicht: (2023)
von: Ghafouri, Saeid, et al.
Veröffentlicht: (2023)
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
SProBench: Stream Processing Benchmark for High Performance Computing Infrastructure
von: Kulkarni, Apurv Deepak, et al.
Veröffentlicht: (2025)
von: Kulkarni, Apurv Deepak, et al.
Veröffentlicht: (2025)
Fake Runs, Real Fixes -- Analyzing xPU Performance Through Simulation
von: Zarkadas, Ioannis, et al.
Veröffentlicht: (2025)
von: Zarkadas, Ioannis, et al.
Veröffentlicht: (2025)
Distributed Matrix-Based Sampling for Graph Neural Network Training
von: Tripathy, Alok, et al.
Veröffentlicht: (2023)
von: Tripathy, Alok, et al.
Veröffentlicht: (2023)
Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
von: Maczan, Jędrzej
Veröffentlicht: (2026)
von: Maczan, Jędrzej
Veröffentlicht: (2026)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
von: Rose, Martin, et al.
Veröffentlicht: (2025)
von: Rose, Martin, et al.
Veröffentlicht: (2025)
Execution time budget assignment for mixed criticality systems
von: Khelassi, Mohamed Amine, et al.
Veröffentlicht: (2023)
von: Khelassi, Mohamed Amine, et al.
Veröffentlicht: (2023)
iSpLib: A Library for Accelerating Graph Neural Networks using Auto-tuned Sparse Operations
von: Anik, Md Saidul Hoque, et al.
Veröffentlicht: (2024)
von: Anik, Md Saidul Hoque, et al.
Veröffentlicht: (2024)
Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge
von: Grailoo, M., et al.
Veröffentlicht: (2026)
von: Grailoo, M., et al.
Veröffentlicht: (2026)
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
von: Sridharan, Srinivas, et al.
Veröffentlicht: (2026)
von: Sridharan, Srinivas, et al.
Veröffentlicht: (2026)
PyGim: An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Phantora: Maximizing Code Reuse in Simulation-based Machine Learning System Performance Estimation
von: Qin, Jianxing, et al.
Veröffentlicht: (2025)
von: Qin, Jianxing, et al.
Veröffentlicht: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
FluxSieve: Unifying Streaming and Analytical Data Planes for Scalable Cloud Observability
von: Vogel, Adriano, et al.
Veröffentlicht: (2026)
von: Vogel, Adriano, et al.
Veröffentlicht: (2026)
Towards a Proactive Autoscaling Framework for Data Stream Processing at the Edge using GRU and Transfer Learning
von: Armah, Eugene, et al.
Veröffentlicht: (2025)
von: Armah, Eugene, et al.
Veröffentlicht: (2025)
Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU
von: Ning, Zhenyu, et al.
Veröffentlicht: (2024)
von: Ning, Zhenyu, et al.
Veröffentlicht: (2024)
ISO: Overlap of Computation and Communication within Seqenence For LLM Inference
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction
von: Mumenin, Khondoker Mirazul, et al.
Veröffentlicht: (2025)
von: Mumenin, Khondoker Mirazul, et al.
Veröffentlicht: (2025)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2024) -
Ripple: Scalable Incremental GNN Inferencing on Large Streaming Graphs
von: Naman, Pranjal, et al.
Veröffentlicht: (2025) -
Incremental GNN Embedding Computation on Streaming Graphs
von: Wang, Qiange, et al.
Veröffentlicht: (2026) -
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
von: Chen, David, et al.
Veröffentlicht: (2026) -
Multi-DNN Inference of Sparse Models on Edge SoCs
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)