Date Fragments: A Hidden Bottleneck of Tokenization for Temporal Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Bhatia, Gagan, Peyrard, Maxime, Zhao, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?
by: Bhatia, Gagan, et al.
Published: (2026)
by: Bhatia, Gagan, et al.
Published: (2026)
DateLogicQA: Benchmarking Temporal Biases in Large Language Models
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
Dallah: A Dialect-Aware Multimodal Large Language Model for Arabic
by: Alwajih, Fakhraddin, et al.
Published: (2024)
by: Alwajih, Fakhraddin, et al.
Published: (2024)
Agentic AI: The Era of Semantic Decoding
by: Peyrard, Maxime, et al.
Published: (2024)
by: Peyrard, Maxime, et al.
Published: (2024)
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
by: Geng, Saibo, et al.
Published: (2023)
by: Geng, Saibo, et al.
Published: (2023)
FinTral: A Family of GPT-4 Level Multimodal Financial Large Language Models
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
Heterogeneity in Formal Linguistic Competence of Language Models: Is Data the Real Bottleneck?
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2026)
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2026)
Peacock: A Family of Arabic Multimodal Large Language Models and Benchmarks
by: Alwajih, Fakhraddin, et al.
Published: (2024)
by: Alwajih, Fakhraddin, et al.
Published: (2024)
Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
by: Bhatia, Gagan, et al.
Published: (2025)
by: Bhatia, Gagan, et al.
Published: (2025)
A Study on Hidden Layer Distillation for Large Language Model Pre-Training
by: Guigon, Maxime, et al.
Published: (2026)
by: Guigon, Maxime, et al.
Published: (2026)
RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human Mobility
by: He, Haoyu, et al.
Published: (2025)
by: He, Haoyu, et al.
Published: (2025)
Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning
by: Wang, Yucheng, et al.
Published: (2025)
by: Wang, Yucheng, et al.
Published: (2025)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
Qalam : A Multimodal LLM for Arabic Optical Character and Handwriting Recognition
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
Efficient Reasoning with Hidden Thinking
by: Shen, Xuan, et al.
Published: (2025)
by: Shen, Xuan, et al.
Published: (2025)
The Tokenization Bottleneck: How Vocabulary Extension Improves Chemistry Representation Learning in Pretrained Language Models
by: Kalamkar, Prathamesh, et al.
Published: (2025)
by: Kalamkar, Prathamesh, et al.
Published: (2025)
State over Tokens: Characterizing the Role of Reasoning Tokens
by: Levy, Mosh, et al.
Published: (2025)
by: Levy, Mosh, et al.
Published: (2025)
Tokenization Constraints in LLMs: A Study of Symbolic and Arithmetic Reasoning Limits
by: Zhang, Xiang, et al.
Published: (2025)
by: Zhang, Xiang, et al.
Published: (2025)
Evaluating Language Model Agency through Negotiations
by: Davidson, Tim R., et al.
Published: (2024)
by: Davidson, Tim R., et al.
Published: (2024)
Extending Token Computation for LLM Reasoning
by: Liao, Bingli, et al.
Published: (2024)
by: Liao, Bingli, et al.
Published: (2024)
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
by: Monea, Giovanni, et al.
Published: (2023)
by: Monea, Giovanni, et al.
Published: (2023)
KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical Reasoning
by: Sun, Wei, et al.
Published: (2025)
by: Sun, Wei, et al.
Published: (2025)
How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning
by: Chen, Haoyang, et al.
Published: (2026)
by: Chen, Haoyang, et al.
Published: (2026)
Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
by: Zhang, Anqi, et al.
Published: (2025)
by: Zhang, Anqi, et al.
Published: (2025)
Token-Budget-Aware LLM Reasoning
by: Han, Tingxu, et al.
Published: (2024)
by: Han, Tingxu, et al.
Published: (2024)
TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
by: Wei, Shaohang, et al.
Published: (2025)
by: Wei, Shaohang, et al.
Published: (2025)
STaRR: Spatial-Temporal Token-Dynamics-Aware Responsive Remasking for Diffusion Language Models
by: Sun, Xinhao, et al.
Published: (2025)
by: Sun, Xinhao, et al.
Published: (2025)
Say Anything but This: When Tokenizer Betrays Reasoning in LLMs
by: Ayoobi, Navid, et al.
Published: (2026)
by: Ayoobi, Navid, et al.
Published: (2026)
AdapTime: Enabling Adaptive Temporal Reasoning in Large Language Models
by: Deng, Yimin, et al.
Published: (2026)
by: Deng, Yimin, et al.
Published: (2026)
Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models
by: Guo, Zhenyuan, et al.
Published: (2026)
by: Guo, Zhenyuan, et al.
Published: (2026)
Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
by: Qian, Chen, et al.
Published: (2025)
by: Qian, Chen, et al.
Published: (2025)
Large Language Models-guided Dynamic Adaptation for Temporal Knowledge Graph Reasoning
by: Wang, Jiapu, et al.
Published: (2024)
by: Wang, Jiapu, et al.
Published: (2024)
Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN
by: Xu, Yao, et al.
Published: (2025)
by: Xu, Yao, et al.
Published: (2025)
Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models
by: Yin, Huifeng, et al.
Published: (2025)
by: Yin, Huifeng, et al.
Published: (2025)
MTBench: A Multimodal Time Series Benchmark for Temporal Reasoning and Question Answering
by: Chen, Jialin, et al.
Published: (2025)
by: Chen, Jialin, et al.
Published: (2025)
Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal
by: Yuan, Aojie, et al.
Published: (2026)
by: Yuan, Aojie, et al.
Published: (2026)
Joint Multi-Facts Reasoning Network For Complex Temporal Question Answering Over Knowledge Graph
by: Huang, Rikui, et al.
Published: (2024)
by: Huang, Rikui, et al.
Published: (2024)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
Similar Items
-
What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?
by: Bhatia, Gagan, et al.
Published: (2026) -
DateLogicQA: Benchmarking Temporal Biases in Large Language Models
by: Bhatia, Gagan, et al.
Published: (2024) -
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
by: Méloux, Maxime, et al.
Published: (2025) -
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
by: Méloux, Maxime, et al.
Published: (2025) -
Dallah: A Dialect-Aware Multimodal Large Language Model for Arabic
by: Alwajih, Fakhraddin, et al.
Published: (2024)