TimeSeriesExam: A time series understanding exam
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cai, Yifu, Choudhry, Arjun, Goswami, Mononito, Dubrawski, Artur |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale
von: Gwiazda, Malgorzata, et al.
Veröffentlicht: (2026)
von: Gwiazda, Malgorzata, et al.
Veröffentlicht: (2026)
MOMENT: A Family of Open Time-series Foundation Models
von: Goswami, Mononito, et al.
Veröffentlicht: (2024)
von: Goswami, Mononito, et al.
Veröffentlicht: (2024)
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents
von: Cai, Yifu, et al.
Veröffentlicht: (2025)
von: Cai, Yifu, et al.
Veröffentlicht: (2025)
AQuA: A Benchmarking Tool for Label Quality Assessment
von: Goswami, Mononito, et al.
Veröffentlicht: (2023)
von: Goswami, Mononito, et al.
Veröffentlicht: (2023)
STAMP: Spatial-Temporal Adapter with Multi-Head Pooling
von: Shook, Brad, et al.
Veröffentlicht: (2025)
von: Shook, Brad, et al.
Veröffentlicht: (2025)
Non-Stationarity in the Embedding Space of Time Series Foundation Models
von: Choi, Jinmyeong, et al.
Veröffentlicht: (2026)
von: Choi, Jinmyeong, et al.
Veröffentlicht: (2026)
Exploring Representations and Interventions in Time Series Foundation Models
von: Wiliński, Michał, et al.
Veröffentlicht: (2024)
von: Wiliński, Michał, et al.
Veröffentlicht: (2024)
Towards Long-Context Time Series Foundation Models
von: Żukowska, Nina, et al.
Veröffentlicht: (2024)
von: Żukowska, Nina, et al.
Veröffentlicht: (2024)
The Ever-Evolving Science Exam
von: Wang, Junying, et al.
Veröffentlicht: (2025)
von: Wang, Junying, et al.
Veröffentlicht: (2025)
A SAT-based approach to rigorous verification of Bayesian networks
von: Stępka, Ignacy, et al.
Veröffentlicht: (2024)
von: Stępka, Ignacy, et al.
Veröffentlicht: (2024)
Reinforcement learning fine-tuning of language model for instruction following and math reasoning
von: Han, Yifu, et al.
Veröffentlicht: (2025)
von: Han, Yifu, et al.
Veröffentlicht: (2025)
Humanity's Last Exam
von: Phan, Long, et al.
Veröffentlicht: (2025)
von: Phan, Long, et al.
Veröffentlicht: (2025)
RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)
von: Cuclea, Luca-Ncolae, et al.
Veröffentlicht: (2026)
von: Cuclea, Luca-Ncolae, et al.
Veröffentlicht: (2026)
Implicit Reasoning in Deep Time Series Forecasting
von: Potosnak, Willa, et al.
Veröffentlicht: (2024)
von: Potosnak, Willa, et al.
Veröffentlicht: (2024)
EnviroExam: Benchmarking Environmental Science Knowledge of Large Language Models
von: Huang, Yu, et al.
Veröffentlicht: (2024)
von: Huang, Yu, et al.
Veröffentlicht: (2024)
The Potential of LLMs in Medical Education: Generating Questions and Answers for Qualification Exams
von: Zhu, Yunqi, et al.
Veröffentlicht: (2024)
von: Zhu, Yunqi, et al.
Veröffentlicht: (2024)
LLM Olympiad: Why Model Evaluation Needs a Sealed Exam
von: Cruz, Jan Christian Blaise, et al.
Veröffentlicht: (2026)
von: Cruz, Jan Christian Blaise, et al.
Veröffentlicht: (2026)
Automated Analysis of Learning Outcomes and Exam Questions Based on Bloom's Taxonomy
von: Kumar, Ramya, et al.
Veröffentlicht: (2025)
von: Kumar, Ramya, et al.
Veröffentlicht: (2025)
HALT-RAG: A Task-Adaptable Framework for Hallucination Detection with Calibrated NLI Ensembles and Abstention
von: Goswami, Saumya, et al.
Veröffentlicht: (2025)
von: Goswami, Saumya, et al.
Veröffentlicht: (2025)
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
von: Riachi, Roland, et al.
Veröffentlicht: (2025)
von: Riachi, Roland, et al.
Veröffentlicht: (2025)
Alvorada-Bench: Can Language Models Solve Brazilian University Entrance Exams?
von: Godoy, Henrique
Veröffentlicht: (2025)
von: Godoy, Henrique
Veröffentlicht: (2025)
Reasoning Models Ace the CFA Exams
von: Patel, Jaisal, et al.
Veröffentlicht: (2025)
von: Patel, Jaisal, et al.
Veröffentlicht: (2025)
Evaluating ChatGPT-4 Vision on Brazil's National Undergraduate Computer Science Exam
von: Mendonça, Nabor C.
Veröffentlicht: (2024)
von: Mendonça, Nabor C.
Veröffentlicht: (2024)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
Query Disambiguation via Answer-Free Context: Doubling Performance on Humanity's Last Exam
von: Majurski, Michael, et al.
Veröffentlicht: (2026)
von: Majurski, Michael, et al.
Veröffentlicht: (2026)
Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
Machine-Assisted Grading of Nationwide School-Leaving Essay Exams with LLMs and Statistical NLP
von: Karjus, Andres, et al.
Veröffentlicht: (2026)
von: Karjus, Andres, et al.
Veröffentlicht: (2026)
Towards Interpretable Time Series Foundation Models
von: Boileau, Matthieu, et al.
Veröffentlicht: (2025)
von: Boileau, Matthieu, et al.
Veröffentlicht: (2025)
LLM-as-a-Judge for Time Series Explanations
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning
von: Kong, Yaxuan, et al.
Veröffentlicht: (2026)
von: Kong, Yaxuan, et al.
Veröffentlicht: (2026)
Improving Fairness in LLMs Through Testing-Time Adversaries
von: Gregio, Isabela Pereira, et al.
Veröffentlicht: (2025)
von: Gregio, Isabela Pereira, et al.
Veröffentlicht: (2025)
Empowering Smaller Models: Tuning LLaMA and Gemma with Chain-of-Thought for Ukrainian Exam Tasks
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
Designing Reliable LLM-Assisted Rubric Scoring for Constructed Responses: Evidence from Physics Exams
von: Tang, Xiuxiu, et al.
Veröffentlicht: (2026)
von: Tang, Xiuxiu, et al.
Veröffentlicht: (2026)
TimeSense:Making Large Language Models Proficient in Time-Series Analysis
von: Zhang, Zhirui, et al.
Veröffentlicht: (2025)
von: Zhang, Zhirui, et al.
Veröffentlicht: (2025)
From Transformers to LLMs: A Systematic Survey of Efficiency Considerations in NLP
von: Ansar, Wazib, et al.
Veröffentlicht: (2024)
von: Ansar, Wazib, et al.
Veröffentlicht: (2024)
MTBench: A Multimodal Time Series Benchmark for Temporal Reasoning and Question Answering
von: Chen, Jialin, et al.
Veröffentlicht: (2025)
von: Chen, Jialin, et al.
Veröffentlicht: (2025)
EMTSF:Extraordinary Mixture of SOTA Models for Time Series Forecasting
von: Alharthi, Musleh, et al.
Veröffentlicht: (2025)
von: Alharthi, Musleh, et al.
Veröffentlicht: (2025)
Small but Mighty: Enhancing Time Series Forecasting with Lightweight LLMs
von: Fan, Haoran, et al.
Veröffentlicht: (2025)
von: Fan, Haoran, et al.
Veröffentlicht: (2025)
Think While You Write: Hypothesis Verification Promotes Faithful Knowledge-to-Text Generation
von: Qiu, Yifu, et al.
Veröffentlicht: (2023)
von: Qiu, Yifu, et al.
Veröffentlicht: (2023)
Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale
von: Gwiazda, Malgorzata, et al.
Veröffentlicht: (2026) -
MOMENT: A Family of Open Time-series Foundation Models
von: Goswami, Mononito, et al.
Veröffentlicht: (2024) -
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents
von: Cai, Yifu, et al.
Veröffentlicht: (2025) -
AQuA: A Benchmarking Tool for Label Quality Assessment
von: Goswami, Mononito, et al.
Veröffentlicht: (2023) -
STAMP: Spatial-Temporal Adapter with Multi-Head Pooling
von: Shook, Brad, et al.
Veröffentlicht: (2025)