Keras Sig: Efficient Path Signature Computation on GPU in Keras 3
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Genet, Rémi, Inzirillo, Hugo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Remote KV Cache Reuse with GPU-native Video Codec
von: Mi, Liang, et al.
Veröffentlicht: (2026)
von: Mi, Liang, et al.
Veröffentlicht: (2026)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
Eliminating Multi-GPU Performance Taxes: A Systems Approach to Efficient Distributed LLMs
von: Trifan, Octavian Alexandru, et al.
Veröffentlicht: (2025)
von: Trifan, Octavian Alexandru, et al.
Veröffentlicht: (2025)
GeoT: Tensor Centric Library for Graph Neural Network via Efficient Segment Reduction on GPU
von: Yu, Zhongming, et al.
Veröffentlicht: (2024)
von: Yu, Zhongming, et al.
Veröffentlicht: (2024)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
Single-GPU GNN Systems: Traps and Pitfalls
von: Gong, Yidong, et al.
Veröffentlicht: (2024)
von: Gong, Yidong, et al.
Veröffentlicht: (2024)
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
von: Li, Jinhao, et al.
Veröffentlicht: (2023)
von: Li, Jinhao, et al.
Veröffentlicht: (2023)
Semantic-Aware Scheduling for GPU Clusters with Large Language Models
von: Wang, Zerui, et al.
Veröffentlicht: (2025)
von: Wang, Zerui, et al.
Veröffentlicht: (2025)
Iris: First-Class Multi-GPU Programming Experience in Triton
von: Awad, Muhammad, et al.
Veröffentlicht: (2025)
von: Awad, Muhammad, et al.
Veröffentlicht: (2025)
KernelFoundry: Hardware-aware evolutionary GPU kernel optimization
von: Wiedemann, Nina, et al.
Veröffentlicht: (2026)
von: Wiedemann, Nina, et al.
Veröffentlicht: (2026)
GNNBENCH: Fair and Productive Benchmarking for Single-GPU GNN System
von: Gong, Yidong, et al.
Veröffentlicht: (2024)
von: Gong, Yidong, et al.
Veröffentlicht: (2024)
ParallelKittens: Systematic and Practical Simplification of Multi-GPU AI Kernels
von: Sul, Stuart H., et al.
Veröffentlicht: (2025)
von: Sul, Stuart H., et al.
Veröffentlicht: (2025)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
von: Shi, Xiaoxiang, et al.
Veröffentlicht: (2025)
von: Shi, Xiaoxiang, et al.
Veröffentlicht: (2025)
Ilargi: a GPU Compatible Factorized ML Model Training Framework
von: Sun, Wenbo, et al.
Veröffentlicht: (2025)
von: Sun, Wenbo, et al.
Veröffentlicht: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
von: Zhang, Haolin, et al.
Veröffentlicht: (2025)
von: Zhang, Haolin, et al.
Veröffentlicht: (2025)
A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
von: Agrawal, Amey, et al.
Veröffentlicht: (2026)
von: Agrawal, Amey, et al.
Veröffentlicht: (2026)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
von: Wang, Chong, et al.
Veröffentlicht: (2026)
von: Wang, Chong, et al.
Veröffentlicht: (2026)
Communication and Computation Efficient Split Federated Learning in O-RAN
von: Gu, Shunxian, et al.
Veröffentlicht: (2025)
von: Gu, Shunxian, et al.
Veröffentlicht: (2025)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
von: Luo, Ziyue, et al.
Veröffentlicht: (2025)
von: Luo, Ziyue, et al.
Veröffentlicht: (2025)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
von: Yarlagadda, Srihas, et al.
Veröffentlicht: (2025)
von: Yarlagadda, Srihas, et al.
Veröffentlicht: (2025)
RAPID-Serve: Resource-efficient and Accelerated P/D Intra-GPU Disaggregation
von: Masood, Amna, et al.
Veröffentlicht: (2026)
von: Masood, Amna, et al.
Veröffentlicht: (2026)
FloatSOM: GPU-Accelerated, Distributed, Topology-Flexible Self-Organizing Maps
von: Xu, Tony, et al.
Veröffentlicht: (2026)
von: Xu, Tony, et al.
Veröffentlicht: (2026)
GPU-Accelerated Optimization of Transformer-Based Neural Networks for Real-Time Inference
von: Mukherjee, Soutrik, et al.
Veröffentlicht: (2026)
von: Mukherjee, Soutrik, et al.
Veröffentlicht: (2026)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
von: Gond, Raja, et al.
Veröffentlicht: (2025)
von: Gond, Raja, et al.
Veröffentlicht: (2025)
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025)
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025)
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
von: Shu, Zhihao, et al.
Veröffentlicht: (2026)
von: Shu, Zhihao, et al.
Veröffentlicht: (2026)
SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data
von: Kapadia, Shashank, et al.
Veröffentlicht: (2026)
von: Kapadia, Shashank, et al.
Veröffentlicht: (2026)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
FT K-means: A High-Performance K-means on GPU with Fault Tolerance
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
FIKIT: Priority-Based Real-time GPU Multi-tasking Scheduling with Kernel Identification
von: Wu, Wenqing
Veröffentlicht: (2023)
von: Wu, Wenqing
Veröffentlicht: (2023)
Computation and Communication Efficient Lightweighting Vertical Federated Learning for Smart Building IoT
von: Wang, Heqiang, et al.
Veröffentlicht: (2024)
von: Wang, Heqiang, et al.
Veröffentlicht: (2024)
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
von: Ning, Rui, et al.
Veröffentlicht: (2026)
von: Ning, Rui, et al.
Veröffentlicht: (2026)
Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
von: Zhao, Lingxiao, et al.
Veröffentlicht: (2025)
von: Zhao, Lingxiao, et al.
Veröffentlicht: (2025)
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
von: Tummalapalli, Pranay, et al.
Veröffentlicht: (2026)
von: Tummalapalli, Pranay, et al.
Veröffentlicht: (2026)
SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-Device Inference
von: Khare, Alind, et al.
Veröffentlicht: (2023)
von: Khare, Alind, et al.
Veröffentlicht: (2023)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Efficient Remote KV Cache Reuse with GPU-native Video Codec
von: Mi, Liang, et al.
Veröffentlicht: (2026) -
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
von: Griggs, Tyler, et al.
Veröffentlicht: (2024) -
Eliminating Multi-GPU Performance Taxes: A Systems Approach to Efficient Distributed LLMs
von: Trifan, Octavian Alexandru, et al.
Veröffentlicht: (2025) -
GeoT: Tensor Centric Library for Graph Neural Network via Efficient Segment Reduction on GPU
von: Yu, Zhongming, et al.
Veröffentlicht: (2024) -
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)