KORAL: Knowledge Graph Guided LLM Reasoning for SSD Operational Analysis
Fuente:
arXiv
Guardado en:
| Autores principales: | Akewar, Mayur, Madireddy, Sandeep, Luo, Dongsheng, Bhimani, Janki |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
por: Luo, Xinhao, et al.
Publicado: (2025)
por: Luo, Xinhao, et al.
Publicado: (2025)
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
por: Da, Wei, et al.
Publicado: (2025)
por: Da, Wei, et al.
Publicado: (2025)
Evaluating the Efficacy of LLM-Based Reasoning for Multiobjective HPC Job Scheduling
por: Jadhav, Prachi, et al.
Publicado: (2025)
por: Jadhav, Prachi, et al.
Publicado: (2025)
Zipage: Maintain High Request Concurrency for LLM Reasoning through Compressed PagedAttention
por: Liao, Mengqi, et al.
Publicado: (2026)
por: Liao, Mengqi, et al.
Publicado: (2026)
B-PASTE: Beam-Aware Pattern-Guided Speculative Execution for Resource-Constrained LLM Agents
por: Song, Yanfei
Publicado: (2026)
por: Song, Yanfei
Publicado: (2026)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
por: Liu, Xing, et al.
Publicado: (2025)
por: Liu, Xing, et al.
Publicado: (2025)
Accelerated Digital Twin Learning for Edge AI: A Comparison of FPGA and Mobile GPU
por: Xu, Bin, et al.
Publicado: (2025)
por: Xu, Bin, et al.
Publicado: (2025)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
por: Li, Zhonggen, et al.
Publicado: (2025)
por: Li, Zhonggen, et al.
Publicado: (2025)
Simplifying Root Cause Analysis in Kubernetes with StateGraph and LLM
por: Xiang, Yong, et al.
Publicado: (2025)
por: Xiang, Yong, et al.
Publicado: (2025)
MemAscend: System Memory Optimization for SSD-Offloaded LLM Fine-Tuning
por: Liaw, Yong-Cheng, et al.
Publicado: (2025)
por: Liaw, Yong-Cheng, et al.
Publicado: (2025)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
por: Chen, Aodong, et al.
Publicado: (2023)
por: Chen, Aodong, et al.
Publicado: (2023)
Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware
por: Khalil, Alex, et al.
Publicado: (2025)
por: Khalil, Alex, et al.
Publicado: (2025)
Transforming Future Data Center Operations and Management via Physical AI
por: Cao, Zhiwei, et al.
Publicado: (2025)
por: Cao, Zhiwei, et al.
Publicado: (2025)
xLLM Technical Report
por: Liu, Tongxuan, et al.
Publicado: (2025)
por: Liu, Tongxuan, et al.
Publicado: (2025)
Elastic On-Device LLM Service
por: Yin, Wangsong, et al.
Publicado: (2024)
por: Yin, Wangsong, et al.
Publicado: (2024)
Idiosyncrasies of Programmable Caching Engines
por: Peixoto, José, et al.
Publicado: (2026)
por: Peixoto, José, et al.
Publicado: (2026)
A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
por: Xi, Shaoke, et al.
Publicado: (2026)
por: Xi, Shaoke, et al.
Publicado: (2026)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
por: Zhang, Ziyang, et al.
Publicado: (2025)
por: Zhang, Ziyang, et al.
Publicado: (2025)
AI Benchmarks and Datasets for LLM Evaluation
por: Ivanov, Todor, et al.
Publicado: (2024)
por: Ivanov, Todor, et al.
Publicado: (2024)
PolyKAN: Efficient Fused GPU Operators for Polynomial Kolmogorov-Arnold Network Variants
por: Yu, Mingkun, et al.
Publicado: (2025)
por: Yu, Mingkun, et al.
Publicado: (2025)
Learning Provably Correct Distributed Protocols Without Human Knowledge
por: Hui, Yujie, et al.
Publicado: (2026)
por: Hui, Yujie, et al.
Publicado: (2026)
EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
por: Kubwimana, Benjamin, et al.
Publicado: (2025)
por: Kubwimana, Benjamin, et al.
Publicado: (2025)
Revisiting Parameter Server in LLM Post-Training
por: Wan, Xinyi, et al.
Publicado: (2026)
por: Wan, Xinyi, et al.
Publicado: (2026)
Accelerating LLM Inference with Precomputed Query Storage
por: Park, Jay H., et al.
Publicado: (2025)
por: Park, Jay H., et al.
Publicado: (2025)
High-Throughput LLM inference on Heterogeneous Clusters
por: Xiong, Yi, et al.
Publicado: (2025)
por: Xiong, Yi, et al.
Publicado: (2025)
Byzantine-Robust Decentralized Coordination of LLM Agents
por: Jo, Yongrae, et al.
Publicado: (2025)
por: Jo, Yongrae, et al.
Publicado: (2025)
Tutoring LLM into a Better CUDA Optimizer
por: Brabec, Matyáš, et al.
Publicado: (2025)
por: Brabec, Matyáš, et al.
Publicado: (2025)
A Parallel CPU-GPU Framework for Batching Heuristic Operations in Depth-First Heuristic Search
por: Futuhi, Ehsan, et al.
Publicado: (2025)
por: Futuhi, Ehsan, et al.
Publicado: (2025)
LLM Inference Serving: Survey of Recent Advances and Opportunities
por: Li, Baolin, et al.
Publicado: (2024)
por: Li, Baolin, et al.
Publicado: (2024)
Decentralized AI: Permissionless LLM Inference on POKT Network
por: Olshansky, Daniel, et al.
Publicado: (2024)
por: Olshansky, Daniel, et al.
Publicado: (2024)
A Hashgraph-Inspired Consensus Mechanism for Reliable Multi-Model Reasoning
por: Ogunsina, Kolawole E., et al.
Publicado: (2025)
por: Ogunsina, Kolawole E., et al.
Publicado: (2025)
SemanticForge: Repository-Level Code Generation through Semantic Knowledge Graphs and Constraint Satisfaction
por: Zhang, Wuyang, et al.
Publicado: (2025)
por: Zhang, Wuyang, et al.
Publicado: (2025)
ECCENTRIC: Edge-Cloud Collaboration Framework for Distributed Inference Using Knowledge Adaptation
por: Kamani, Mohammad Mahdi, et al.
Publicado: (2025)
por: Kamani, Mohammad Mahdi, et al.
Publicado: (2025)
LAPS: A Length-Aware-Prefill LLM Serving System
por: She, Jianshu, et al.
Publicado: (2026)
por: She, Jianshu, et al.
Publicado: (2026)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
por: Wagenländer, Marcel, et al.
Publicado: (2026)
por: Wagenländer, Marcel, et al.
Publicado: (2026)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
por: Hankendi, Can, et al.
Publicado: (2026)
por: Hankendi, Can, et al.
Publicado: (2026)
PRAGMA: A Profiling-Reasoned Multi-Agent Framework for Automatic Kernel Optimization
por: Lei, Kelun, et al.
Publicado: (2025)
por: Lei, Kelun, et al.
Publicado: (2025)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
por: Lyu, Hongtao, et al.
Publicado: (2025)
por: Lyu, Hongtao, et al.
Publicado: (2025)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
por: Zhang, Ping, et al.
Publicado: (2024)
por: Zhang, Ping, et al.
Publicado: (2024)
LLM as HPC Expert: Extending RAG Architecture for HPC Data
por: Miyashita, Yusuke, et al.
Publicado: (2024)
por: Miyashita, Yusuke, et al.
Publicado: (2024)
Ejemplares similares
-
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
por: Luo, Xinhao, et al.
Publicado: (2025) -
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
por: Da, Wei, et al.
Publicado: (2025) -
Evaluating the Efficacy of LLM-Based Reasoning for Multiobjective HPC Job Scheduling
por: Jadhav, Prachi, et al.
Publicado: (2025) -
Zipage: Maintain High Request Concurrency for LLM Reasoning through Compressed PagedAttention
por: Liao, Mengqi, et al.
Publicado: (2026) -
B-PASTE: Beam-Aware Pattern-Guided Speculative Execution for Resource-Constrained LLM Agents
por: Song, Yanfei
Publicado: (2026)