Spark-LLM-Eval: A Distributed Framework for Statistically Rigorous Large Language Model Evaluation
Fuente:
arXiv
Guardado en:
| Autor principal: | Mitra, Subhadip |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Addressing tokens dynamic generation, propagation, storage and renewal to secure the GlideinWMS pilot based jobs and system
por: Coimbra, Bruno Moreira, et al.
Publicado: (2025)
por: Coimbra, Bruno Moreira, et al.
Publicado: (2025)
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
por: Gao, Yuxuan, et al.
Publicado: (2026)
por: Gao, Yuxuan, et al.
Publicado: (2026)
Deploy, Calibrate, Monitor, Heal -- No Human Required: An Autonomous AI SRE Agent for Elasticsearch
por: Mukkolakkal, Muhamed Ramees Cheriya
Publicado: (2026)
por: Mukkolakkal, Muhamed Ramees Cheriya
Publicado: (2026)
AutoBench: Automating LLM Evaluation through Reciprocal Peer Assessment
por: Loi, Dario, et al.
Publicado: (2025)
por: Loi, Dario, et al.
Publicado: (2025)
Evaluating Large Language Models for Workload Mapping and Scheduling in Heterogeneous HPC Systems
por: Sharma, Aasish Kumar, et al.
Publicado: (2025)
por: Sharma, Aasish Kumar, et al.
Publicado: (2025)
Combining Serverless and High-Performance Computing Paradigms to support ML Data-Intensive Applications
por: Staylor, Mills, et al.
Publicado: (2025)
por: Staylor, Mills, et al.
Publicado: (2025)
Deep RC: A Scalable Data Engineering and Deep Learning Pipeline
por: Sarker, Arup Kumar, et al.
Publicado: (2025)
por: Sarker, Arup Kumar, et al.
Publicado: (2025)
Design and Implementation of an Analysis Pipeline for Heterogeneous Data
por: Sarker, Arup Kumar, et al.
Publicado: (2024)
por: Sarker, Arup Kumar, et al.
Publicado: (2024)
StepCache: Step-Level Reuse with Lightweight Verification and Selective Patching for LLM Serving
por: Nouri, Azam
Publicado: (2026)
por: Nouri, Azam
Publicado: (2026)
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
por: Zehra, Sehar, et al.
Publicado: (2025)
por: Zehra, Sehar, et al.
Publicado: (2025)
N2N: A Parallel Framework for Large-Scale MILP under Distributed Memory
por: Wang, Longfei, et al.
Publicado: (2025)
por: Wang, Longfei, et al.
Publicado: (2025)
DAGER: Exact Gradient Inversion for Large Language Models
por: Petrov, Ivo, et al.
Publicado: (2024)
por: Petrov, Ivo, et al.
Publicado: (2024)
Scalable Co-Clustering for Large-Scale Data through Dynamic Partitioning and Hierarchical Merging
por: Wu, Zihan, et al.
Publicado: (2024)
por: Wu, Zihan, et al.
Publicado: (2024)
Challenges of Heterogeneity in Big Data: A Comparative Study of Classification in Large-Scale Structured and Unstructured Domains
por: Eduardo, González Trigueros Jesús, et al.
Publicado: (2025)
por: Eduardo, González Trigueros Jesús, et al.
Publicado: (2025)
GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration
por: Sarker, Yeahia, et al.
Publicado: (2026)
por: Sarker, Yeahia, et al.
Publicado: (2026)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
por: Li, Xiangchen, et al.
Publicado: (2026)
por: Li, Xiangchen, et al.
Publicado: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
por: Li, Xiangchen, et al.
Publicado: (2026)
por: Li, Xiangchen, et al.
Publicado: (2026)
Cost Trade-offs of Reasoning and Non-Reasoning Large Language Models in Text-to-SQL
por: Deochake, Saurabh, et al.
Publicado: (2025)
por: Deochake, Saurabh, et al.
Publicado: (2025)
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
por: Will, Jonathan, et al.
Publicado: (2025)
por: Will, Jonathan, et al.
Publicado: (2025)
HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment
por: Jang, Yoonjin, et al.
Publicado: (2026)
por: Jang, Yoonjin, et al.
Publicado: (2026)
Efficient Construction of Large Search Spaces for Auto-Tuning
por: Willemsen, Floris-Jan, et al.
Publicado: (2025)
por: Willemsen, Floris-Jan, et al.
Publicado: (2025)
Learning Interpretable Scheduling Algorithms for Data Processing Clusters
por: Hu, Zhibo, et al.
Publicado: (2024)
por: Hu, Zhibo, et al.
Publicado: (2024)
CooperLLM: Cloud-Edge-End Cooperative Federated Fine-tuning for LLMs via ZOO-based Gradient Correction
por: Sun, He, et al.
Publicado: (2026)
por: Sun, He, et al.
Publicado: (2026)
Using Containers to Speed Up Development, to Run Integration Tests and to Teach About Distributed Systems
por: Mambelli, Marco, et al.
Publicado: (2025)
por: Mambelli, Marco, et al.
Publicado: (2025)
SLA Management in Reconfigurable Multi-Agent RAG: A Systems Approach to Question Answering
por: Iannelli, Michael, et al.
Publicado: (2024)
por: Iannelli, Michael, et al.
Publicado: (2024)
Deadline-Aware Joint Task Scheduling and Offloading in Mobile Edge Computing Systems
por: Nguyen, Ngoc Hung, et al.
Publicado: (2025)
por: Nguyen, Ngoc Hung, et al.
Publicado: (2025)
Scalability Optimization in Cloud-Based AI Inference Services: Strategies for Real-Time Load Balancing and Automated Scaling
por: Jin, Yihong, et al.
Publicado: (2025)
por: Jin, Yihong, et al.
Publicado: (2025)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
por: Kolluru, Saicharan
Publicado: (2025)
por: Kolluru, Saicharan
Publicado: (2025)
Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling
por: Luo, Zizhang, et al.
Publicado: (2026)
por: Luo, Zizhang, et al.
Publicado: (2026)
Operational Memory Architecture for Kubernetes:Preserving Causal Context Across the Evidence Horizon
por: Khan, Shamsher
Publicado: (2026)
por: Khan, Shamsher
Publicado: (2026)
Flash-Fusion: Enabling Expressive, Low-Latency Queries on IoT Sensor Streams with LLMs
por: Patherya, Kausar, et al.
Publicado: (2025)
por: Patherya, Kausar, et al.
Publicado: (2025)
Cost-Aware Logging: Measuring the Financial Impact of Excessive Log Retention in Small-Scale Cloud Deployments
por: Putra, Jody Almaida
Publicado: (2026)
por: Putra, Jody Almaida
Publicado: (2026)
AAFLOW: Scalable Patterns for Agentic AI Workflows
por: Sarker, Arup Kumar, et al.
Publicado: (2026)
por: Sarker, Arup Kumar, et al.
Publicado: (2026)
GraphEval36K: Benchmarking Coding and Reasoning Capabilities of Large Language Models on Graph Datasets
por: Wu, Qiming, et al.
Publicado: (2024)
por: Wu, Qiming, et al.
Publicado: (2024)
Evaluating the Overhead of the Performance Profiler Cloudprofiler With MooBench
por: Yang, Shinhyung, et al.
Publicado: (2024)
por: Yang, Shinhyung, et al.
Publicado: (2024)
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
por: Li, Zonghang, et al.
Publicado: (2025)
por: Li, Zonghang, et al.
Publicado: (2025)
Heuristic Search Space Partitioning for Low-Latency Multi-Tenant Cloud Queries
por: Pathak, Prashant Kumar, et al.
Publicado: (2026)
por: Pathak, Prashant Kumar, et al.
Publicado: (2026)
Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation
por: Koc, Vincent
Publicado: (2025)
por: Koc, Vincent
Publicado: (2025)
Retrieval and Augmentation of Domain Knowledge for Text-to-SQL Semantic Parsing
por: Patwardhan, Manasi, et al.
Publicado: (2025)
por: Patwardhan, Manasi, et al.
Publicado: (2025)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
por: Kamath, Aditya K, et al.
Publicado: (2024)
por: Kamath, Aditya K, et al.
Publicado: (2024)
Ejemplares similares
-
Addressing tokens dynamic generation, propagation, storage and renewal to secure the GlideinWMS pilot based jobs and system
por: Coimbra, Bruno Moreira, et al.
Publicado: (2025) -
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
por: Gao, Yuxuan, et al.
Publicado: (2026) -
Deploy, Calibrate, Monitor, Heal -- No Human Required: An Autonomous AI SRE Agent for Elasticsearch
por: Mukkolakkal, Muhamed Ramees Cheriya
Publicado: (2026) -
AutoBench: Automating LLM Evaluation through Reciprocal Peer Assessment
por: Loi, Dario, et al.
Publicado: (2025) -
Evaluating Large Language Models for Workload Mapping and Scheduling in Heterogeneous HPC Systems
por: Sharma, Aasish Kumar, et al.
Publicado: (2025)