ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Stephan, Cohen, Ben, Goswami, Mononito, Shen, Junhong, Khwaja, Emaad, Liu, Chenghao, Asker, David, Abou-Amal, Othmane, Talwalkar, Ameet |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toto 2.0: Time Series Forecasting Enters the Scaling Era
by: Khwaja, Emaad, et al.
Published: (2026)
by: Khwaja, Emaad, et al.
Published: (2026)
Toto: Time Series Optimized Transformer for Observability
by: Cohen, Ben, et al.
Published: (2024)
by: Cohen, Ben, et al.
Published: (2024)
This Time is Different: An Observability Perspective on Time Series Foundation Models
by: Cohen, Ben, et al.
Published: (2025)
by: Cohen, Ben, et al.
Published: (2025)
UPS: Efficiently Building Foundation Models for PDE Solving via Cross-Modal Adaptation
by: Shen, Junhong, et al.
Published: (2024)
by: Shen, Junhong, et al.
Published: (2024)
RECODE: Reasoning Through Code Generation for Visual Question Answering
by: Shen, Junhong, et al.
Published: (2025)
by: Shen, Junhong, et al.
Published: (2025)
TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale
by: Gwiazda, Malgorzata, et al.
Published: (2026)
by: Gwiazda, Malgorzata, et al.
Published: (2026)
Specialized Foundation Models Struggle to Beat Supervised Baselines
by: Xu, Zongzhe, et al.
Published: (2024)
by: Xu, Zongzhe, et al.
Published: (2024)
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents
by: Cai, Yifu, et al.
Published: (2025)
by: Cai, Yifu, et al.
Published: (2025)
CodePDE: An Inference Framework for LLM-driven PDE Solver Generation
by: Li, Shanda, et al.
Published: (2025)
by: Li, Shanda, et al.
Published: (2025)
ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data
by: Shen, Junhong, et al.
Published: (2024)
by: Shen, Junhong, et al.
Published: (2024)
The Impact of Element Ordering on LM Agent Performance
by: Chi, Wayne, et al.
Published: (2024)
by: Chi, Wayne, et al.
Published: (2024)
TimeSeriesExam: A time series understanding exam
by: Cai, Yifu, et al.
Published: (2024)
by: Cai, Yifu, et al.
Published: (2024)
HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
by: Ben-Ami, Dan, et al.
Published: (2025)
by: Ben-Ami, Dan, et al.
Published: (2025)
Exploring Representations and Interventions in Time Series Foundation Models
by: Wiliński, Michał, et al.
Published: (2024)
by: Wiliński, Michał, et al.
Published: (2024)
Towards Long-Context Time Series Foundation Models
by: Żukowska, Nina, et al.
Published: (2024)
by: Żukowska, Nina, et al.
Published: (2024)
Sample Complexity and Representation Ability of Test-time Scaling Paradigms
by: Huang, Baihe, et al.
Published: (2025)
by: Huang, Baihe, et al.
Published: (2025)
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
by: Chen, Valerie, et al.
Published: (2025)
by: Chen, Valerie, et al.
Published: (2025)
Why Do Decision Makers (Not) Use AI? A Cross-Domain Analysis of Factors Impacting AI Adoption
by: Yu, Rebecca, et al.
Published: (2025)
by: Yu, Rebecca, et al.
Published: (2025)
Agreement-Based Cascading for Efficient Inference
by: Kolawole, Steven, et al.
Published: (2024)
by: Kolawole, Steven, et al.
Published: (2024)
When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback
by: Pan, Jane, et al.
Published: (2025)
by: Pan, Jane, et al.
Published: (2025)
Layered Diffusion Model for One-Shot High Resolution Text-to-Image Synthesis
by: Khwaja, Emaad, et al.
Published: (2024)
by: Khwaja, Emaad, et al.
Published: (2024)
Implicit Reasoning in Deep Time Series Forecasting
by: Potosnak, Willa, et al.
Published: (2024)
by: Potosnak, Willa, et al.
Published: (2024)
TSAQA: Time Series Analysis Question And Answering Benchmark
by: Jing, Baoyu, et al.
Published: (2026)
by: Jing, Baoyu, et al.
Published: (2026)
Understanding Optimization in Deep Learning with Central Flows
by: Cohen, Jeremy M., et al.
Published: (2024)
by: Cohen, Jeremy M., et al.
Published: (2024)
Provably tuning the ElasticNet across instances
by: Balcan, Maria-Florina, et al.
Published: (2022)
by: Balcan, Maria-Florina, et al.
Published: (2022)
Learning to Relax: Setting Solver Parameters Across a Sequence of Linear System Instances
by: Khodak, Mikhail, et al.
Published: (2023)
by: Khodak, Mikhail, et al.
Published: (2023)
Where Does My Model Underperform? A Human Evaluation of Slice Discovery Algorithms
by: Johnson, Nari, et al.
Published: (2023)
by: Johnson, Nari, et al.
Published: (2023)
Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting
by: Garza, Azul, et al.
Published: (2026)
by: Garza, Azul, et al.
Published: (2026)
Testing Question Answering Software with Context-Driven Question Generation
by: Liu, Shuang, et al.
Published: (2025)
by: Liu, Shuang, et al.
Published: (2025)
SpIDER: Spatially Informed Dense Embedding Retrieval for Software Issue Localization
by: Chaudhari, Shravan, et al.
Published: (2025)
by: Chaudhari, Shravan, et al.
Published: (2025)
AQuA: A Benchmarking Tool for Label Quality Assessment
by: Goswami, Mononito, et al.
Published: (2023)
by: Goswami, Mononito, et al.
Published: (2023)
The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers
by: Mozannar, Hussein, et al.
Published: (2024)
by: Mozannar, Hussein, et al.
Published: (2024)
Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants
by: Chen, Valerie, et al.
Published: (2026)
by: Chen, Valerie, et al.
Published: (2026)
CoMind: Towards Community-Driven Agents for Machine Learning Engineering
by: Li, Sijie, et al.
Published: (2025)
by: Li, Sijie, et al.
Published: (2025)
FrontierCO: Real-World and Large-Scale Evaluation of Machine Learning Solvers for Combinatorial Optimization
by: Feng, Shengyu, et al.
Published: (2025)
by: Feng, Shengyu, et al.
Published: (2025)
UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks
by: Nguyen, Jason, et al.
Published: (2026)
by: Nguyen, Jason, et al.
Published: (2026)
Toward Automated Security Risk Detection in Large Software Using Call Graph Analysis
by: Pecka, Nicholas, et al.
Published: (2025)
by: Pecka, Nicholas, et al.
Published: (2025)
MICA: Multivariate Infini Compressive Attention for Time Series Forecasting
by: Potosnak, Willa, et al.
Published: (2026)
by: Potosnak, Willa, et al.
Published: (2026)
Investigating Compositional Reasoning in Time Series Foundation Models
by: Potosnak, Willa, et al.
Published: (2025)
by: Potosnak, Willa, et al.
Published: (2025)
Do LLMs exhibit human-like response biases? A case study in survey design
by: Tjuatja, Lindia, et al.
Published: (2023)
by: Tjuatja, Lindia, et al.
Published: (2023)
Similar Items
-
Toto 2.0: Time Series Forecasting Enters the Scaling Era
by: Khwaja, Emaad, et al.
Published: (2026) -
Toto: Time Series Optimized Transformer for Observability
by: Cohen, Ben, et al.
Published: (2024) -
This Time is Different: An Observability Perspective on Time Series Foundation Models
by: Cohen, Ben, et al.
Published: (2025) -
UPS: Efficiently Building Foundation Models for PDE Solving via Cross-Modal Adaptation
by: Shen, Junhong, et al.
Published: (2024) -
RECODE: Reasoning Through Code Generation for Visual Question Answering
by: Shen, Junhong, et al.
Published: (2025)