Task Scheduling for Efficient Inference of Large Language Models on Single Moderate GPU Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Wenxiang, Pan, Xinglin, Shi, Shaohuai, Wang, Xuan, Chu, Xiaowen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
by: Pan, Xinglin, et al.
Published: (2024)
by: Pan, Xinglin, et al.
Published: (2024)
Integrating Optimal Transport and Structural Inference Models for GRN Inference from Single-cell Data
by: Tong, Tsz Pan, et al.
Published: (2024)
by: Tong, Tsz Pan, et al.
Published: (2024)
Automated machine learning for physics-informed convolutional neural networks
by: Zhou, Wanyun, et al.
Published: (2024)
by: Zhou, Wanyun, et al.
Published: (2024)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
by: Pan, Xinglin, et al.
Published: (2025)
by: Pan, Xinglin, et al.
Published: (2025)
A Memory Efficient Adjoint Method to Enable Billion Parameter Optimization on a Single GPU in Dynamic Problems
by: Herrmann, Leon, et al.
Published: (2025)
by: Herrmann, Leon, et al.
Published: (2025)
Population-Based Search Method Using Uncertainty-related Pareto Front for Robust Multi-objective Optimization
by: Xu, Lihong, et al.
Published: (2025)
by: Xu, Lihong, et al.
Published: (2025)
Two-Stage Optimization for Efficient V2G Coordination in Distribution Power System
by: Tian, Pengchao, et al.
Published: (2024)
by: Tian, Pengchao, et al.
Published: (2024)
ZiGong 1.0: A Large Language Model for Financial Credit
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing
by: Lv, Liuzhenghao, et al.
Published: (2024)
by: Lv, Liuzhenghao, et al.
Published: (2024)
Constructing Mechanical Design Agent Based on Large Language Models
by: Lu, Jiaxing, et al.
Published: (2024)
by: Lu, Jiaxing, et al.
Published: (2024)
Unleashing Expert Opinion from Social Media for Stock Prediction
by: Zhou, Wanyun, et al.
Published: (2025)
by: Zhou, Wanyun, et al.
Published: (2025)
Efficient Computation of Redundancy Matrices for Moderately Redundant Truss and Frame Structures
by: Tkachuk, Anton, et al.
Published: (2023)
by: Tkachuk, Anton, et al.
Published: (2023)
Enhancing Evolutionary Solver Efficiency for NP Hard Single Machine Scheduling Problems
by: Alromema, Mohammed, et al.
Published: (2024)
by: Alromema, Mohammed, et al.
Published: (2024)
Energy-Adaptive Checkpoint-Free Intermittent Inference for Low Power Energy Harvesting Systems
by: Islam, Sahidul, et al.
Published: (2025)
by: Islam, Sahidul, et al.
Published: (2025)
Mitigating Procrastination in Spatial Crowdsourcing Via Efficient Scheduling Algorithm
by: Debnath, Naren, et al.
Published: (2024)
by: Debnath, Naren, et al.
Published: (2024)
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training
by: Lin, Wenxiang, et al.
Published: (2026)
by: Lin, Wenxiang, et al.
Published: (2026)
Evaluation and Benchmarking Suite for Financial Large Language Models and Agents
by: Lin, Shengyuan, et al.
Published: (2026)
by: Lin, Shengyuan, et al.
Published: (2026)
OpenTM: An Open-source, Single-GPU, Large-scale Thermal Microstructure Design Framework
by: Quan, Yuchen, et al.
Published: (2024)
by: Quan, Yuchen, et al.
Published: (2024)
A Matrix-Free Galerkin Multigrid Solver and Failure-Mode Screen for Single-GPU 3D SIMP Linear Systems
by: Yang, Shaoliang, et al.
Published: (2026)
by: Yang, Shaoliang, et al.
Published: (2026)
Kolmogorov-Arnold Network for Gene Regulatory Network Inference
by: Tong, Tsz Pan, et al.
Published: (2025)
by: Tong, Tsz Pan, et al.
Published: (2025)
Position: LLM Inference Should Be Evaluated as Energy-to-Token Production
by: Liu, Xiang, et al.
Published: (2026)
by: Liu, Xiang, et al.
Published: (2026)
Efficient Road Renovation Scheduling under Uncertainty using Lower Bound Pruning
by: Bosch, Robbert, et al.
Published: (2026)
by: Bosch, Robbert, et al.
Published: (2026)
DeltaLag: Learning Dynamic Lead-Lag Patterns in Financial Markets
by: Zhou, Wanyun, et al.
Published: (2025)
by: Zhou, Wanyun, et al.
Published: (2025)
Connectomics Informed by Large Language Models
by: Thompson, Elinor, et al.
Published: (2025)
by: Thompson, Elinor, et al.
Published: (2025)
Can Large Language Models Effectively Process and Execute Financial Trading Instructions?
by: Kang, Yu, et al.
Published: (2024)
by: Kang, Yu, et al.
Published: (2024)
UCFE: A User-Centric Financial Expertise Benchmark for Large Language Models
by: Yang, Yuzhe, et al.
Published: (2024)
by: Yang, Yuzhe, et al.
Published: (2024)
PINNsAgent: Automated PDE Surrogation with Large Language Models
by: Wuwu, Qingpo, et al.
Published: (2025)
by: Wuwu, Qingpo, et al.
Published: (2025)
Physics-Aware Compression of Plasma Distribution Functions with GPU-Accelerated Gaussian Mixture Models
by: Hu, Andong, et al.
Published: (2025)
by: Hu, Andong, et al.
Published: (2025)
DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering
by: Li, Yinsheng, et al.
Published: (2025)
by: Li, Yinsheng, et al.
Published: (2025)
MFMDQwen: Multilingual Financial Misinformation Detection Based on Large Language Model
by: Liu, Zhiwei, et al.
Published: (2026)
by: Liu, Zhiwei, et al.
Published: (2026)
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications
by: Huang, Jimin, et al.
Published: (2024)
by: Huang, Jimin, et al.
Published: (2024)
Benchmarking Large Language Models for Polymer Property Predictions
by: Gupta, Sonakshi, et al.
Published: (2025)
by: Gupta, Sonakshi, et al.
Published: (2025)
Using Large Language Models for Solving Thermodynamic Problems
by: Loubet, Rebecca, et al.
Published: (2025)
by: Loubet, Rebecca, et al.
Published: (2025)
Finetuning Large Language Model as an Effective Symbolic Regressor
by: Hua, Yingfan, et al.
Published: (2025)
by: Hua, Yingfan, et al.
Published: (2025)
An Efficient Bayesian Framework for Inverse Problems via Optimization and Inversion: Surrogate Modeling, Parameter Inference, and Uncertainty Quantification
by: Chiappetta, Mihaela, et al.
Published: (2026)
by: Chiappetta, Mihaela, et al.
Published: (2026)
Generalized Scattering Matrix Synthesis: Independent Region Decomposition for Hybrid Antenna--Scatterer Systems
by: Shi, Chenbo, et al.
Published: (2025)
by: Shi, Chenbo, et al.
Published: (2025)
Linear and Non-Linear Models for Master Scheduling of Dynamic Resources Product Mix
by: Mohammed, Ayman R., et al.
Published: (2024)
by: Mohammed, Ayman R., et al.
Published: (2024)
From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models
by: Kuang, Ziyan, et al.
Published: (2025)
by: Kuang, Ziyan, et al.
Published: (2025)
Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks
by: Wang, Julian Junyan, et al.
Published: (2025)
by: Wang, Julian Junyan, et al.
Published: (2025)
LaMP-Val: Large Language Models Empower Personalized Valuation in Auction
by: Sun, Jie, et al.
Published: (2024)
by: Sun, Jie, et al.
Published: (2024)
Similar Items
-
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
by: Pan, Xinglin, et al.
Published: (2024) -
Integrating Optimal Transport and Structural Inference Models for GRN Inference from Single-cell Data
by: Tong, Tsz Pan, et al.
Published: (2024) -
Automated machine learning for physics-informed convolutional neural networks
by: Zhou, Wanyun, et al.
Published: (2024) -
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
by: Pan, Xinglin, et al.
Published: (2025) -
A Memory Efficient Adjoint Method to Enable Billion Parameter Optimization on a Single GPU in Dynamic Problems
by: Herrmann, Leon, et al.
Published: (2025)