Performance Prediction for Large Systems via Text-to-Text Regression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Akhauri, Yash, Lewandowski, Bryan, Lin, Cheng-Hsi, Reyes, Adrian N., Forbes, Grant C., Wongpanich, Arissa, Yang, Bangding, Abdelfattah, Mohamed S., Perel, Sagi, Song, Xingyou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912452989419520
author Akhauri, Yash
Lewandowski, Bryan
Lin, Cheng-Hsi
Reyes, Adrian N.
Forbes, Grant C.
Wongpanich, Arissa
Yang, Bangding
Abdelfattah, Mohamed S.
Perel, Sagi
Song, Xingyou
author_facet Akhauri, Yash
Lewandowski, Bryan
Lin, Cheng-Hsi
Reyes, Adrian N.
Forbes, Grant C.
Wongpanich, Arissa
Yang, Bangding
Abdelfattah, Mohamed S.
Perel, Sagi
Song, Xingyou
contents In many industries, predicting metric outcomes of large systems is a fundamental problem, driven largely by traditional tabular regression. However, such methods struggle on complex systems data in the wild such as configuration files or system logs, where feature engineering is often infeasible. We propose text-to-text regression as a general, scalable alternative. For predicting resource efficiency on Borg, Google's massive compute cluster scheduling system, a 60M parameter encoder-decoder, trained from random initialization, achieves up to a near perfect 0.99 (0.9 average) rank correlation across the entire fleet, and 100x lower MSE than tabular approaches. The model also easily adapts to new tasks in only 500 few-shot examples and captures the densities of complex outcome distributions. Ablation studies highlight the importance of using encoders, increasing sequence length, and the model's inherent uncertainty quantification. These findings pave the way for universal simulators of real-world outcomes.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21718
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Performance Prediction for Large Systems via Text-to-Text Regression
Akhauri, Yash
Lewandowski, Bryan
Lin, Cheng-Hsi
Reyes, Adrian N.
Forbes, Grant C.
Wongpanich, Arissa
Yang, Bangding
Abdelfattah, Mohamed S.
Perel, Sagi
Song, Xingyou
Machine Learning
Artificial Intelligence
Performance
Software Engineering
Systems and Control
In many industries, predicting metric outcomes of large systems is a fundamental problem, driven largely by traditional tabular regression. However, such methods struggle on complex systems data in the wild such as configuration files or system logs, where feature engineering is often infeasible. We propose text-to-text regression as a general, scalable alternative. For predicting resource efficiency on Borg, Google's massive compute cluster scheduling system, a 60M parameter encoder-decoder, trained from random initialization, achieves up to a near perfect 0.99 (0.9 average) rank correlation across the entire fleet, and 100x lower MSE than tabular approaches. The model also easily adapts to new tasks in only 500 few-shot examples and captures the densities of complex outcome distributions. Ablation studies highlight the importance of using encoders, increasing sequence length, and the model's inherent uncertainty quantification. These findings pave the way for universal simulators of real-world outcomes.
title Performance Prediction for Large Systems via Text-to-Text Regression
topic Machine Learning
Artificial Intelligence
Performance
Software Engineering
Systems and Control
url https://arxiv.org/abs/2506.21718