Remember This Event That Year? Assessing Temporal Information and Reasoning in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Beniwal, Himanshu, Patel, Dishant, D, Kowsik Nandagopan, Ladia, Hritik, Yadav, Ankit, Singh, Mayank |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cross-lingual Editing in Multilingual Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2026)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2026)
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
von: Sheth, Rajvee, et al.
Veröffentlicht: (2025)
von: Sheth, Rajvee, et al.
Veröffentlicht: (2025)
Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework
von: Sheth, Rajvee, et al.
Veröffentlicht: (2024)
von: Sheth, Rajvee, et al.
Veröffentlicht: (2024)
Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models across Modalities
von: Sheth, Rajvee, et al.
Veröffentlicht: (2025)
von: Sheth, Rajvee, et al.
Veröffentlicht: (2025)
One Instruction Does Not Fit All: How Well Do Embeddings Align Personas and Instructions in Low-Resource Indian Languages?
von: Shah, Arya, et al.
Veröffentlicht: (2026)
von: Shah, Arya, et al.
Veröffentlicht: (2026)
Cause and Effect: Can Large Language Models Truly Understand Causality?
von: Ashwani, Swagata, et al.
Veröffentlicht: (2024)
von: Ashwani, Swagata, et al.
Veröffentlicht: (2024)
UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
von: Bansal, Hritik, et al.
Veröffentlicht: (2023)
von: Bansal, Hritik, et al.
Veröffentlicht: (2023)
HALO: An Ontology for Representing and Categorizing Hallucinations in Large Language Models
von: Nananukul, Navapat, et al.
Veröffentlicht: (2023)
von: Nananukul, Navapat, et al.
Veröffentlicht: (2023)
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
Error Taxonomy-Guided Prompt Optimization
von: Singh, Mayank, et al.
Veröffentlicht: (2026)
von: Singh, Mayank, et al.
Veröffentlicht: (2026)
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop
von: Afzal, Anum, et al.
Veröffentlicht: (2024)
von: Afzal, Anum, et al.
Veröffentlicht: (2024)
A Comprehensive Evaluation on Event Reasoning of Large Language Models
von: Tao, Zhengwei, et al.
Veröffentlicht: (2024)
von: Tao, Zhengwei, et al.
Veröffentlicht: (2024)
Navigating Semantic Relations: Challenges for Language Models in Abstract Common-Sense Reasoning
von: Gawin, Cole, et al.
Veröffentlicht: (2025)
von: Gawin, Cole, et al.
Veröffentlicht: (2025)
No Universal Prompt: Unifying Reasoning through Adaptive Prompting for Temporal Table Reasoning
von: Rajgaria, Abhishek, et al.
Veröffentlicht: (2025)
von: Rajgaria, Abhishek, et al.
Veröffentlicht: (2025)
LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
AdapTime: Enabling Adaptive Temporal Reasoning in Large Language Models
von: Deng, Yimin, et al.
Veröffentlicht: (2026)
von: Deng, Yimin, et al.
Veröffentlicht: (2026)
DEPART: DEcomposing PARiTy across Multilingual LLMs
von: Uppadhyay, Manan, et al.
Veröffentlicht: (2026)
von: Uppadhyay, Manan, et al.
Veröffentlicht: (2026)
Large Language Models-guided Dynamic Adaptation for Temporal Knowledge Graph Reasoning
von: Wang, Jiapu, et al.
Veröffentlicht: (2024)
von: Wang, Jiapu, et al.
Veröffentlicht: (2024)
Assessing and Enhancing the Robustness of Large Language Models with Task Structure Variations for Logical Reasoning
von: Bao, Qiming, et al.
Veröffentlicht: (2023)
von: Bao, Qiming, et al.
Veröffentlicht: (2023)
An Evaluation of Estimative Uncertainty in Large Language Models
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
Assessing Large Language Models on Climate Information
von: Bulian, Jannis, et al.
Veröffentlicht: (2023)
von: Bulian, Jannis, et al.
Veröffentlicht: (2023)
Multilingual Information Retrieval with a Monolingual Knowledge Base
von: Zhuang, Yingying, et al.
Veröffentlicht: (2025)
von: Zhuang, Yingying, et al.
Veröffentlicht: (2025)
Narrative-of-Thought: Improving Temporal Reasoning of Large Language Models via Recounted Narratives
von: Zhang, Xinliang Frederick, et al.
Veröffentlicht: (2024)
von: Zhang, Xinliang Frederick, et al.
Veröffentlicht: (2024)
TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
von: Chu, Zheng, et al.
Veröffentlicht: (2023)
von: Chu, Zheng, et al.
Veröffentlicht: (2023)
What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?
von: Bhatia, Gagan, et al.
Veröffentlicht: (2026)
von: Bhatia, Gagan, et al.
Veröffentlicht: (2026)
Defining and Evaluating Decision and Composite Risk in Language Models Applied to Natural Language Inference
von: Shen, Ke, et al.
Veröffentlicht: (2024)
von: Shen, Ke, et al.
Veröffentlicht: (2024)
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Towards Explainable Temporal Reasoning in Large Language Models: A Structure-Aware Generative Framework
von: Jiang, Zihao, et al.
Veröffentlicht: (2025)
von: Jiang, Zihao, et al.
Veröffentlicht: (2025)
SemEval-2026 Task 12: Abductive Event Reasoning: Towards Real-World Event Causal Inference for Large Language Models
von: Cao, Pengfei, et al.
Veröffentlicht: (2026)
von: Cao, Pengfei, et al.
Veröffentlicht: (2026)
Towards a More Inclusive AI: Progress and Perspectives in Large Language Model Training for the Sámi Language
von: Paul, Ronny, et al.
Veröffentlicht: (2024)
von: Paul, Ronny, et al.
Veröffentlicht: (2024)
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
von: Suvarna, Ashima, et al.
Veröffentlicht: (2026)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2026)
Assessing and Understanding Creativity in Large Language Models
von: Zhao, Yunpu, et al.
Veröffentlicht: (2024)
von: Zhao, Yunpu, et al.
Veröffentlicht: (2024)
Assessing Political Bias in Large Language Models
von: Rettenberger, Luca, et al.
Veröffentlicht: (2024)
von: Rettenberger, Luca, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Cross-lingual Editing in Multilingual Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024) -
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
von: Yadav, Ankit, et al.
Veröffentlicht: (2024) -
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2026) -
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
von: Sheth, Rajvee, et al.
Veröffentlicht: (2025) -
Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)