TeleEval-OS: Performance evaluations of large language models for operations scheduling
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yanyan, Wang, Yingying, Liang, Junli, Xu, Yin, Liu, Yunlong, Xu, Yiming, Jiang, Zhengwang, Li, Zhehe, Li, Fei, Zhao, Long, Xu, Kuang, Song, Qi, Li, Xiangyang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient scheduling in redundancy systems with general service times
by: Anton, Elene, et al.
Published: (2022)
by: Anton, Elene, et al.
Published: (2022)
Profiling Apple Silicon Performance for ML Training
by: Feng, Dahua, et al.
Published: (2025)
by: Feng, Dahua, et al.
Published: (2025)
Response time in a pair of processor sharing queues with Join-the-Shortest-Queue scheduling
by: Bor, Julianna, et al.
Published: (2024)
by: Bor, Julianna, et al.
Published: (2024)
CAPSim: A Fast CPU Performance Simulator Using Attention-based Predictor
by: Xu, Buqing, et al.
Published: (2025)
by: Xu, Buqing, et al.
Published: (2025)
Anatomizing Deep Learning Inference in Web Browsers
by: Wang, Qipeng, et al.
Published: (2024)
by: Wang, Qipeng, et al.
Published: (2024)
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
by: Yin, Wangsong, et al.
Published: (2025)
by: Yin, Wangsong, et al.
Published: (2025)
Closed-Loop Integrated Sensing, Communication, and Control for Efficient Drone Flight
by: Li, Jingli, et al.
Published: (2026)
by: Li, Jingli, et al.
Published: (2026)
DSO: A GPU Energy Efficiency Optimizer by Fusing Dynamic and Static Information
by: Wang, Qiang, et al.
Published: (2024)
by: Wang, Qiang, et al.
Published: (2024)
The Theoretical Limit of Radar Target Detection
by: Xu, Dazhuan, et al.
Published: (2021)
by: Xu, Dazhuan, et al.
Published: (2021)
Two-Timescale Dynamic Service Deployment and Task Scheduling with Spatiotemporal Collaboration in Mobile Edge Networks
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Systematic Performance Evaluation Framework for LEO Mega-Constellation Satellite Networks
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
by: Huang, Haochen, et al.
Published: (2025)
by: Huang, Haochen, et al.
Published: (2025)
SynthEval: A Framework for Detailed Utility and Privacy Evaluation of Tabular Synthetic Data
by: Lautrup, Anton Danholt, et al.
Published: (2024)
by: Lautrup, Anton Danholt, et al.
Published: (2024)
Mosaic: Cross-Modal Clustering for Efficient Video Understanding
by: Wang, Tuowei, et al.
Published: (2026)
by: Wang, Tuowei, et al.
Published: (2026)
Spatiotemporal Non-Uniformity-Aware Online Task Scheduling in Collaborative Edge Computing for Industrial Internet of Things
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
by: Zhao, Yanbo, et al.
Published: (2025)
by: Zhao, Yanbo, et al.
Published: (2025)
ZO2: Scalable Zeroth-Order Fine-Tuning for Extremely Large Language Models with Limited GPU Memory
by: Wang, Liangyu, et al.
Published: (2025)
by: Wang, Liangyu, et al.
Published: (2025)
Exploiting Scheduling Flexibility via State-Based Scheduling When Guaranteeing Worst-Case Services
by: Xu, Yike, et al.
Published: (2026)
by: Xu, Yike, et al.
Published: (2026)
The Price of Interoperability: Exploring Cross-Chain Bridges and Their Economic Consequences
by: Cao, Yiyue, et al.
Published: (2026)
by: Cao, Yiyue, et al.
Published: (2026)
Attributing the System's Overall Effect to its Components
by: Wang, Chenxi, et al.
Published: (2026)
by: Wang, Chenxi, et al.
Published: (2026)
Supercharging Packet-level Network Simulation of Large Model Training via Memoization and Fast-Forwarding
by: Long, Fei, et al.
Published: (2026)
by: Long, Fei, et al.
Published: (2026)
DeepContext: A Context-aware, Cross-platform, and Cross-framework Tool for Performance Profiling and Analysis of Deep Learning Workloads
by: Zhao, Qidong, et al.
Published: (2024)
by: Zhao, Qidong, et al.
Published: (2024)
Redundant Array Computation Elimination
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Towards Efficient Multi-Scale Deformable Attention on NPU
by: Huang, Chenghuan, et al.
Published: (2025)
by: Huang, Chenghuan, et al.
Published: (2025)
DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMs
by: Chen, Mingkai, et al.
Published: (2024)
by: Chen, Mingkai, et al.
Published: (2024)
RIS-Assisted Received Adaptive Spatial Modulation for Wireless Communications
by: Zhang, Chaorong, et al.
Published: (2024)
by: Zhang, Chaorong, et al.
Published: (2024)
RAGPerf: An End-to-End Benchmarking Framework for Retrieval-Augmented Generation Systems
by: Li, Shaobo, et al.
Published: (2026)
by: Li, Shaobo, et al.
Published: (2026)
Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
AGIPC: Adaptive In-Solve Algebraic Coarsening for GPU IPC
by: Wang, Xuan, et al.
Published: (2026)
by: Wang, Xuan, et al.
Published: (2026)
Intent-driven scheduling of backup jobs
by: Dutta, Souvik, et al.
Published: (2024)
by: Dutta, Souvik, et al.
Published: (2024)
Accelerating Vertical Federated Learning
by: Cai, Dongqi, et al.
Published: (2022)
by: Cai, Dongqi, et al.
Published: (2022)
AI Load Dynamics--A Power Electronics Perspective
by: Li, Yuzhuo, et al.
Published: (2025)
by: Li, Yuzhuo, et al.
Published: (2025)
Optimization of Armv9 architecture general large language model inference performance based on Llama.cpp
by: Chen, Longhao, et al.
Published: (2024)
by: Chen, Longhao, et al.
Published: (2024)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
by: Fu, Zizhuo, et al.
Published: (2025)
by: Fu, Zizhuo, et al.
Published: (2025)
CEBench: A Benchmarking Toolkit for the Cost-Effectiveness of LLM Pipelines
by: Sun, Wenbo, et al.
Published: (2024)
by: Sun, Wenbo, et al.
Published: (2024)
Optimal local storage policy based on stochastic intensities and its large scale behavior
by: Carrasco, Matias, et al.
Published: (2024)
by: Carrasco, Matias, et al.
Published: (2024)
Tuning Fast Memory Size based on Modeling of Page Migration for Tiered Memory
by: Chen, Shangye, et al.
Published: (2024)
by: Chen, Shangye, et al.
Published: (2024)
PrETi: Predicting Execution Time in Early Stage with LLVM and Machine Learning
by: Xu, Risheng, et al.
Published: (2025)
by: Xu, Risheng, et al.
Published: (2025)
SysOM-AI: Continuous Cross-Layer Performance Diagnosis for Production AI Training
by: Zheng, Yusheng, et al.
Published: (2026)
by: Zheng, Yusheng, et al.
Published: (2026)
Optimizing Winograd Convolution on ARMv8 processors
by: Gui, Haoyuan, et al.
Published: (2024)
by: Gui, Haoyuan, et al.
Published: (2024)
Similar Items
-
Efficient scheduling in redundancy systems with general service times
by: Anton, Elene, et al.
Published: (2022) -
Profiling Apple Silicon Performance for ML Training
by: Feng, Dahua, et al.
Published: (2025) -
Response time in a pair of processor sharing queues with Join-the-Shortest-Queue scheduling
by: Bor, Julianna, et al.
Published: (2024) -
CAPSim: A Fast CPU Performance Simulator Using Attention-based Predictor
by: Xu, Buqing, et al.
Published: (2025) -
Anatomizing Deep Learning Inference in Web Browsers
by: Wang, Qipeng, et al.
Published: (2024)