AI Benchmarks and Datasets for LLM Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ivanov, Todor, Penchev, Valeri |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025)
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025)
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
von: Chandrasekar, Ashok, et al.
Veröffentlicht: (2026)
von: Chandrasekar, Ashok, et al.
Veröffentlicht: (2026)
Decentralized AI: Permissionless LLM Inference on POKT Network
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
LLM & HPC:Benchmarking DeepSeek's Performance in High-Performance Computing Tasks
von: Nader, Noujoud, et al.
Veröffentlicht: (2025)
von: Nader, Noujoud, et al.
Veröffentlicht: (2025)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
Evaluating the Efficacy of LLM-Based Reasoning for Multiobjective HPC Job Scheduling
von: Jadhav, Prachi, et al.
Veröffentlicht: (2025)
von: Jadhav, Prachi, et al.
Veröffentlicht: (2025)
Benchmarking Federated Learning in Edge Computing Environments: A Systematic Review and Performance Evaluation
von: Aribe Jr., Sales, et al.
Veröffentlicht: (2026)
von: Aribe Jr., Sales, et al.
Veröffentlicht: (2026)
Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware
von: Khalil, Alex, et al.
Veröffentlicht: (2025)
von: Khalil, Alex, et al.
Veröffentlicht: (2025)
MLCommons Cloud Masking Benchmark with Early Stopping
von: Chennamsetti, Varshitha, et al.
Veröffentlicht: (2023)
von: Chennamsetti, Varshitha, et al.
Veröffentlicht: (2023)
Elastic On-Device LLM Service
von: Yin, Wangsong, et al.
Veröffentlicht: (2024)
von: Yin, Wangsong, et al.
Veröffentlicht: (2024)
xLLM Technical Report
von: Liu, Tongxuan, et al.
Veröffentlicht: (2025)
von: Liu, Tongxuan, et al.
Veröffentlicht: (2025)
A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
von: Xi, Shaoke, et al.
Veröffentlicht: (2026)
von: Xi, Shaoke, et al.
Veröffentlicht: (2026)
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
von: Abramovich, Talor, et al.
Veröffentlicht: (2026)
von: Abramovich, Talor, et al.
Veröffentlicht: (2026)
Accelerating LLM Inference with Precomputed Query Storage
von: Park, Jay H., et al.
Veröffentlicht: (2025)
von: Park, Jay H., et al.
Veröffentlicht: (2025)
High-Throughput LLM inference on Heterogeneous Clusters
von: Xiong, Yi, et al.
Veröffentlicht: (2025)
von: Xiong, Yi, et al.
Veröffentlicht: (2025)
Byzantine-Robust Decentralized Coordination of LLM Agents
von: Jo, Yongrae, et al.
Veröffentlicht: (2025)
von: Jo, Yongrae, et al.
Veröffentlicht: (2025)
Revisiting Parameter Server in LLM Post-Training
von: Wan, Xinyi, et al.
Veröffentlicht: (2026)
von: Wan, Xinyi, et al.
Veröffentlicht: (2026)
Tutoring LLM into a Better CUDA Optimizer
von: Brabec, Matyáš, et al.
Veröffentlicht: (2025)
von: Brabec, Matyáš, et al.
Veröffentlicht: (2025)
Benchmarking of CPU-intensive Stream Data Processing in The Edge Computing Systems
von: Szydlo, Tomasz, et al.
Veröffentlicht: (2025)
von: Szydlo, Tomasz, et al.
Veröffentlicht: (2025)
CloudEval-YAML: A Practical Benchmark for Cloud Configuration Generation
von: Xu, Yifei, et al.
Veröffentlicht: (2023)
von: Xu, Yifei, et al.
Veröffentlicht: (2023)
LLM Inference Serving: Survey of Recent Advances and Opportunities
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
Multi-IaC-Eval: Benchmarking Cloud Infrastructure as Code Across Multiple Formats
von: Davidson, Sam, et al.
Veröffentlicht: (2025)
von: Davidson, Sam, et al.
Veröffentlicht: (2025)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
von: Zhang, Ping, et al.
Veröffentlicht: (2024)
von: Zhang, Ping, et al.
Veröffentlicht: (2024)
LLM as HPC Expert: Extending RAG Architecture for HPC Data
von: Miyashita, Yusuke, et al.
Veröffentlicht: (2024)
von: Miyashita, Yusuke, et al.
Veröffentlicht: (2024)
Evaluating Kubernetes Performance for GenAI Inference: From Automatic Speech Recognition to LLM Summarization
von: Malleni, Sai Sindhur, et al.
Veröffentlicht: (2026)
von: Malleni, Sai Sindhur, et al.
Veröffentlicht: (2026)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
LAPS: A Length-Aware-Prefill LLM Serving System
von: She, Jianshu, et al.
Veröffentlicht: (2026)
von: She, Jianshu, et al.
Veröffentlicht: (2026)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding
von: McDanel, Bradley, et al.
Veröffentlicht: (2025)
von: McDanel, Bradley, et al.
Veröffentlicht: (2025)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
von: Hao, Zixu, et al.
Veröffentlicht: (2025)
von: Hao, Zixu, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Compound AI Systems
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025)
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025)
GPU-Virt-Bench: A Comprehensive Benchmarking Framework for Software-Based GPU Virtualization Systems
von: VG, Jithin, et al.
Veröffentlicht: (2025)
von: VG, Jithin, et al.
Veröffentlicht: (2025)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
von: Huang, Tao, et al.
Veröffentlicht: (2024)
von: Huang, Tao, et al.
Veröffentlicht: (2024)
HiveMind: OS-Inspired Scheduling for Concurrent LLM Agent Workloads
von: Agyemang, Justice Owusu, et al.
Veröffentlicht: (2026)
von: Agyemang, Justice Owusu, et al.
Veröffentlicht: (2026)
TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training
von: Ye, Chenhao, et al.
Veröffentlicht: (2026)
von: Ye, Chenhao, et al.
Veröffentlicht: (2026)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
von: Kim, Joon Ha, et al.
Veröffentlicht: (2026)
von: Kim, Joon Ha, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025) -
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
von: Chandrasekar, Ashok, et al.
Veröffentlicht: (2026) -
Decentralized AI: Permissionless LLM Inference on POKT Network
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024) -
LLM & HPC:Benchmarking DeepSeek's Performance in High-Performance Computing Tasks
von: Nader, Noujoud, et al.
Veröffentlicht: (2025) -
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)