TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Weng, Muyan, Cao, Defu, Yang, Wei, Sharma, Yashaswi, Liu, Yan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adaptive Collaboration with Humans: Metacognitive Policy Optimization for Multi-Agent LLMs with Continual Learning
von: Yang, Wei, et al.
Veröffentlicht: (2026)
von: Yang, Wei, et al.
Veröffentlicht: (2026)
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
von: Cai, Mu, et al.
Veröffentlicht: (2024)
von: Cai, Mu, et al.
Veröffentlicht: (2024)
When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
von: Ye, Wen, et al.
Veröffentlicht: (2025)
von: Ye, Wen, et al.
Veröffentlicht: (2025)
TS-Reasoner: Domain-Oriented Time Series Inference Agents for Reasoning and Automated Analysis
von: Ye, Wen, et al.
Veröffentlicht: (2024)
von: Ye, Wen, et al.
Veröffentlicht: (2024)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
TimeDiT: General-purpose Diffusion Transformers for Time Series Foundation Model
von: Cao, Defu, et al.
Veröffentlicht: (2024)
von: Cao, Defu, et al.
Veröffentlicht: (2024)
Conversational Time Series Foundation Models: Towards Explainable and Effective Forecasting
von: Cao, Defu, et al.
Veröffentlicht: (2025)
von: Cao, Defu, et al.
Veröffentlicht: (2025)
TimeCAP: Learning to Contextualize, Augment, and Predict Time Series Events with Large Language Model Agents
von: Lee, Geon, et al.
Veröffentlicht: (2025)
von: Lee, Geon, et al.
Veröffentlicht: (2025)
TSI-Bench: Benchmarking Time Series Imputation
von: Du, Wenjie, et al.
Veröffentlicht: (2024)
von: Du, Wenjie, et al.
Veröffentlicht: (2024)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
von: He, Wei, et al.
Veröffentlicht: (2025)
von: He, Wei, et al.
Veröffentlicht: (2025)
Foundation Models for Demand Forecasting via Dual-Strategy Ensembling
von: Yang, Wei, et al.
Veröffentlicht: (2025)
von: Yang, Wei, et al.
Veröffentlicht: (2025)
WorkstreamBench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
von: Yen, Thomson, et al.
Veröffentlicht: (2026)
von: Yen, Thomson, et al.
Veröffentlicht: (2026)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
von: Yin, Sheng, et al.
Veröffentlicht: (2024)
von: Yin, Sheng, et al.
Veröffentlicht: (2024)
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
von: Cui, Fan, et al.
Veröffentlicht: (2026)
von: Cui, Fan, et al.
Veröffentlicht: (2026)
MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents
von: Wang, Shouju, et al.
Veröffentlicht: (2026)
von: Wang, Shouju, et al.
Veröffentlicht: (2026)
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
von: Liu, Yi, et al.
Veröffentlicht: (2026)
von: Liu, Yi, et al.
Veröffentlicht: (2026)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
von: Long, Xiang, et al.
Veröffentlicht: (2026)
von: Long, Xiang, et al.
Veröffentlicht: (2026)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
ESL-Bench: An Event-Driven Synthetic Longitudinal Benchmark for Health Agents
von: Li, Chao, et al.
Veröffentlicht: (2026)
von: Li, Chao, et al.
Veröffentlicht: (2026)
AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
von: Shang, Yu, et al.
Veröffentlicht: (2025)
von: Shang, Yu, et al.
Veröffentlicht: (2025)
DataSciBench: An LLM Agent Benchmark for Data Science
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
HEARTS: Benchmarking LLM Reasoning on Health Time Series
von: Li, Sirui, et al.
Veröffentlicht: (2026)
von: Li, Sirui, et al.
Veröffentlicht: (2026)
Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents
von: Shen, Yiting, et al.
Veröffentlicht: (2026)
von: Shen, Yiting, et al.
Veröffentlicht: (2026)
A Survey of Reasoning and Agentic Systems in Time Series with Large Language Models
von: Chang, Ching, et al.
Veröffentlicht: (2025)
von: Chang, Ching, et al.
Veröffentlicht: (2025)
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
von: Wang, Luyuan, et al.
Veröffentlicht: (2024)
von: Wang, Luyuan, et al.
Veröffentlicht: (2024)
SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
von: Zhou, Yifan, et al.
Veröffentlicht: (2026)
von: Zhou, Yifan, et al.
Veröffentlicht: (2026)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
von: Guo, Zhengkang, et al.
Veröffentlicht: (2026)
von: Guo, Zhengkang, et al.
Veröffentlicht: (2026)
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
von: Merrill, Mike A., et al.
Veröffentlicht: (2026)
von: Merrill, Mike A., et al.
Veröffentlicht: (2026)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
von: Zheng, Junhao, et al.
Veröffentlicht: (2025)
von: Zheng, Junhao, et al.
Veröffentlicht: (2025)
HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks
von: Bedi, Suhana, et al.
Veröffentlicht: (2026)
von: Bedi, Suhana, et al.
Veröffentlicht: (2026)
MIRAI: Evaluating LLM Agents for Event Forecasting
von: Ye, Chenchen, et al.
Veröffentlicht: (2024)
von: Ye, Chenchen, et al.
Veröffentlicht: (2024)
SPA-Bench: A Comprehensive Benchmark for SmartPhone Agent Evaluation
von: Chen, Jingxuan, et al.
Veröffentlicht: (2024)
von: Chen, Jingxuan, et al.
Veröffentlicht: (2024)
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
von: Dong, Haonan, et al.
Veröffentlicht: (2026)
von: Dong, Haonan, et al.
Veröffentlicht: (2026)
An Electrocardiogram Multi-task Benchmark with Comprehensive Evaluations and Insightful Findings
von: Xu, Yuhao, et al.
Veröffentlicht: (2025)
von: Xu, Yuhao, et al.
Veröffentlicht: (2025)
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks
von: Schmidt, Jan-Philipp
Veröffentlicht: (2026)
von: Schmidt, Jan-Philipp
Veröffentlicht: (2026)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
von: Lee, Gyubok, et al.
Veröffentlicht: (2025)
von: Lee, Gyubok, et al.
Veröffentlicht: (2025)
Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
von: Zhang, Yuzhe, et al.
Veröffentlicht: (2026)
von: Zhang, Yuzhe, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Adaptive Collaboration with Humans: Metacognitive Policy Optimization for Multi-Agent LLMs with Continual Learning
von: Yang, Wei, et al.
Veröffentlicht: (2026) -
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
von: Cai, Mu, et al.
Veröffentlicht: (2024) -
When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
von: Ye, Wen, et al.
Veröffentlicht: (2025) -
TS-Reasoner: Domain-Oriented Time Series Inference Agents for Reasoning and Automated Analysis
von: Ye, Wen, et al.
Veröffentlicht: (2024) -
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
von: Deng, Shihan, et al.
Veröffentlicht: (2024)