TRAVELER: A Benchmark for Evaluating Temporal Reasoning across Vague, Implicit and Explicit References
Fuente:
arXiv
Saved in:
| Main Authors: | Kenneweg, Svenja, Deigmöller, Jörg, Cimiano, Philipp, Eggert, Julian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Factorized Probabilistic Model of the Semantics of Vague Temporal Adverbials Relative to Different Event Types
by: Kenneweg, Svenja, et al.
Published: (2025)
by: Kenneweg, Svenja, et al.
Published: (2025)
Retrieval Augmented Generation Systems: Automatic Dataset Creation, Evaluation and Boolean Agent Setup
by: Kenneweg, Tristan, et al.
Published: (2024)
by: Kenneweg, Tristan, et al.
Published: (2024)
MixRea: Benchmarking Explicit-Implicit Reasoning in Large Language Models
by: Cai, Yuanqing, et al.
Published: (2026)
by: Cai, Yuanqing, et al.
Published: (2026)
CompoST: A Benchmark for Analyzing the Ability of LLMs To Compositionally Interpret Questions in a QALD Setting
by: Schmidt, David Maria, et al.
Published: (2025)
by: Schmidt, David Maria, et al.
Published: (2025)
A Grounded Memory System For Smart Personal Assistants
by: Ocker, Felix, et al.
Published: (2025)
by: Ocker, Felix, et al.
Published: (2025)
Awakening LLMs' Reasoning Potential: A Fine-Grained Pipeline to Evaluate and Mitigate Vague Perception
by: Ling, Zipeng, et al.
Published: (2025)
by: Ling, Zipeng, et al.
Published: (2025)
Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It's Best to Relate Perspectives!
by: Heinisch, Philipp, et al.
Published: (2023)
by: Heinisch, Philipp, et al.
Published: (2023)
ThreatCore: A Benchmark for Explicit and Implicit Threat Detection
by: Bruni, Davide, et al.
Published: (2026)
by: Bruni, Davide, et al.
Published: (2026)
The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics
by: Orlikowski, Matthias, et al.
Published: (2023)
by: Orlikowski, Matthias, et al.
Published: (2023)
Pointing out the Shortcomings of Relation Extraction Models with Semantically Motivated Adversarials
by: Nolano, Gennaro, et al.
Published: (2024)
by: Nolano, Gennaro, et al.
Published: (2024)
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
by: Fatemi, Bahare, et al.
Published: (2024)
by: Fatemi, Bahare, et al.
Published: (2024)
VAQUUM: Are Vague Quantifiers Grounded in Visual Data?
by: Wong, Hugh Mee, et al.
Published: (2025)
by: Wong, Hugh Mee, et al.
Published: (2025)
Argument Summarization and its Evaluation in the Era of Large Language Models
by: Altemeyer, Moritz, et al.
Published: (2025)
by: Altemeyer, Moritz, et al.
Published: (2025)
Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
by: Fleisig, Eve, et al.
Published: (2025)
by: Fleisig, Eve, et al.
Published: (2025)
From Argumentation to Deliberation: Perspectivized Stance Vectors for Fine-grained (Dis)agreement Analysis
by: Plenz, Moritz, et al.
Published: (2025)
by: Plenz, Moritz, et al.
Published: (2025)
Modeling Clinical Uncertainty in Radiology Reports: from Explicit Uncertainty Markers to Implicit Reasoning Pathways
by: Rabaey, Paloma, et al.
Published: (2025)
by: Rabaey, Paloma, et al.
Published: (2025)
From Implicit to Explicit: Token-Efficient Logical Supervision for Mathematical Reasoning in LLMs
by: Wang, Shaojie, et al.
Published: (2026)
by: Wang, Shaojie, et al.
Published: (2026)
Implicit Probabilistic Reasoning Does Not Reflect Explicit Answers in Large Language Models
by: Mondal, Manuel, et al.
Published: (2024)
by: Mondal, Manuel, et al.
Published: (2024)
Evaluating Metrics for Bias in Word Embeddings
by: Schröder, Sarah, et al.
Published: (2021)
by: Schröder, Sarah, et al.
Published: (2021)
Unlocking Fine-Grained Translation Quality Estimation in LRMs through Synergistically Evolving Implicit and Explicit Reasoning
by: Dang, Renfei, et al.
Published: (2026)
by: Dang, Renfei, et al.
Published: (2026)
Lexicalization Is All You Need: Examining the Impact of Lexical Knowledge in a Compositional QALD System
by: Schmidt, David Maria, et al.
Published: (2024)
by: Schmidt, David Maria, et al.
Published: (2024)
Explicit vs. Implicit Biographies: Evaluating and Adapting LLM Information Extraction on Wikidata-Derived Texts
by: Stramiglio, Alessandra, et al.
Published: (2025)
by: Stramiglio, Alessandra, et al.
Published: (2025)
Debiasing Sentence Embedders through Contrastive Word Pairs
by: Kenneweg, Philip, et al.
Published: (2024)
by: Kenneweg, Philip, et al.
Published: (2024)
MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions
by: Huang, Zeyu, et al.
Published: (2025)
by: Huang, Zeyu, et al.
Published: (2025)
BaZi-Based Character Simulation Benchmark: Evaluating AI on Temporal and Persona Reasoning
by: Zheng, Siyuan, et al.
Published: (2025)
by: Zheng, Siyuan, et al.
Published: (2025)
BaziQA-Benchmark: Evaluating Symbolic and Temporally Compositional Reasoning in Large Language Models
by: Chen, Jiangxi, et al.
Published: (2026)
by: Chen, Jiangxi, et al.
Published: (2026)
When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation
by: Jiang, Xunyi, et al.
Published: (2025)
by: Jiang, Xunyi, et al.
Published: (2025)
Resolving Word Vagueness with Scenario-guided Adapter for Natural Language Inference
by: Liu, Yonghao, et al.
Published: (2024)
by: Liu, Yonghao, et al.
Published: (2024)
LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
by: Tang, Weizhi, et al.
Published: (2024)
by: Tang, Weizhi, et al.
Published: (2024)
Modeling the Quality of Dialogical Explanations
by: Alshomary, Milad, et al.
Published: (2024)
by: Alshomary, Milad, et al.
Published: (2024)
What Causes the Failure of Explicit to Implicit Discourse Relation Recognition?
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
Measuring Implicit Bias in Explicitly Unbiased Large Language Models
by: Bai, Xuechunzi, et al.
Published: (2024)
by: Bai, Xuechunzi, et al.
Published: (2024)
Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties
by: Wang, Zhenglin, et al.
Published: (2025)
by: Wang, Zhenglin, et al.
Published: (2025)
TRAM: Benchmarking Temporal Reasoning for Large Language Models
by: Wang, Yuqing, et al.
Published: (2023)
by: Wang, Yuqing, et al.
Published: (2023)
Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
by: Bhattacharya, Debarpan, et al.
Published: (2025)
by: Bhattacharya, Debarpan, et al.
Published: (2025)
Explicit, Implicit, and Scattered: Revisiting Event Extraction to Capture Complex Arguments
by: Sharif, Omar, et al.
Published: (2024)
by: Sharif, Omar, et al.
Published: (2024)
Retrieving Implicit and Explicit Emotional Events Using Large Language Models
by: Hu, Guimin, et al.
Published: (2024)
by: Hu, Guimin, et al.
Published: (2024)
Vague Knowledge: Information without Transitivity and Partitions
by: Xiao, Kerry
Published: (2025)
by: Xiao, Kerry
Published: (2025)
Making Implicit Premises Explicit in Logical Understanding of Enthymemes
by: Feng, Xuyao, et al.
Published: (2026)
by: Feng, Xuyao, et al.
Published: (2026)
Explicit v.s. Implicit Memory: Exploring Multi-hop Complex Reasoning Over Personalized Information
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
Similar Items
-
A Factorized Probabilistic Model of the Semantics of Vague Temporal Adverbials Relative to Different Event Types
by: Kenneweg, Svenja, et al.
Published: (2025) -
Retrieval Augmented Generation Systems: Automatic Dataset Creation, Evaluation and Boolean Agent Setup
by: Kenneweg, Tristan, et al.
Published: (2024) -
MixRea: Benchmarking Explicit-Implicit Reasoning in Large Language Models
by: Cai, Yuanqing, et al.
Published: (2026) -
CompoST: A Benchmark for Analyzing the Ability of LLMs To Compositionally Interpret Questions in a QALD Setting
by: Schmidt, David Maria, et al.
Published: (2025) -
A Grounded Memory System For Smart Personal Assistants
by: Ocker, Felix, et al.
Published: (2025)