Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Chuyifei, Cui, Hongyu, Huang, Xiaowen, Sang, Jitao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
Below-Chance Blindness: Prompted Underperformance in Small LLMs Produces Positional Bias Rather than Answer Avoidance
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Evaluating Long Range Dependency Handling in Code Generation LLMs
by: Assogba, Yannick, et al.
Published: (2024)
by: Assogba, Yannick, et al.
Published: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Contextual Integrity in LLMs via Reasoning and Reinforcement Learning
by: Lan, Guangchen, et al.
Published: (2025)
by: Lan, Guangchen, et al.
Published: (2025)
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
by: Nadali, Alireza, et al.
Published: (2026)
by: Nadali, Alireza, et al.
Published: (2026)
Efficient Solutions For An Intriguing Failure of LLMs: Long Context Window Does Not Mean LLMs Can Analyze Long Sequences Flawlessly
by: Hosseini, Peyman, et al.
Published: (2024)
by: Hosseini, Peyman, et al.
Published: (2024)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
by: Le, Van-Truong
Published: (2026)
by: Le, Van-Truong
Published: (2026)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
by: Vieira, Inês, et al.
Published: (2026)
by: Vieira, Inês, et al.
Published: (2026)
Detecting Data Contamination in LLMs via In-Context Learning
by: Zawalski, Michał, et al.
Published: (2025)
by: Zawalski, Michał, et al.
Published: (2025)
Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs
by: Paulsen, Norman
Published: (2025)
by: Paulsen, Norman
Published: (2025)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
by: Adapala, Sai Teja Reddy
Published: (2025)
by: Adapala, Sai Teja Reddy
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Measuring Reasoning Utility in LLMs via Conditional Entropy Reduction
by: Guo, Xu
Published: (2025)
by: Guo, Xu
Published: (2025)
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks
by: Tahir, Munief Hassan, et al.
Published: (2024)
by: Tahir, Munief Hassan, et al.
Published: (2024)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Self-Supervised Position Debiasing for Large Language Models
by: Liu, Zhongkun, et al.
Published: (2024)
by: Liu, Zhongkun, et al.
Published: (2024)
Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
by: Williamson, Dane, et al.
Published: (2025)
by: Williamson, Dane, et al.
Published: (2025)
KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
by: Saha, Soumadeep, et al.
Published: (2025)
by: Saha, Soumadeep, et al.
Published: (2025)
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning
by: Zhang, Zhaowei, et al.
Published: (2026)
by: Zhang, Zhaowei, et al.
Published: (2026)
Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
by: Pai, Aaditya
Published: (2026)
by: Pai, Aaditya
Published: (2026)
Enhancing In-Context Learning via Implicit Demonstration Augmentation
by: Zhou, Xiaoling, et al.
Published: (2024)
by: Zhou, Xiaoling, et al.
Published: (2024)
Adaptive Interviewing for Persona Simulation in LLMs: Evidence-Grounded Reasoning Improves Decision Alignment
by: Su, Ruoxi, et al.
Published: (2026)
by: Su, Ruoxi, et al.
Published: (2026)
GMoE: Empowering LLMs Fine-Tuning via MoE Graph Collaboration
by: Bai, Ting, et al.
Published: (2024)
by: Bai, Ting, et al.
Published: (2024)
MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
by: Marcuzzo, Matteo, et al.
Published: (2025)
by: Marcuzzo, Matteo, et al.
Published: (2025)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
by: Er, Yakup Abrek, et al.
Published: (2025)
by: Er, Yakup Abrek, et al.
Published: (2025)
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
by: Kubica, Dominick, et al.
Published: (2025)
by: Kubica, Dominick, et al.
Published: (2025)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
by: Wu, Zhengxuan, et al.
Published: (2025)
by: Wu, Zhengxuan, et al.
Published: (2025)
LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation
by: Bishop, Jennifer A, et al.
Published: (2023)
by: Bishop, Jennifer A, et al.
Published: (2023)
Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
by: Chen, Haolin, et al.
Published: (2024)
by: Chen, Haolin, et al.
Published: (2024)
RAVR: Reference-Answer-guided Variational Reasoning for Large Language Models
by: Lin, Tianqianjin, et al.
Published: (2025)
by: Lin, Tianqianjin, et al.
Published: (2025)
RAudit: A Blind Auditing Protocol for Large Language Model Reasoning
by: Chang, Edward Y., et al.
Published: (2026)
by: Chang, Edward Y., et al.
Published: (2026)
UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
by: Shafique, Muhammad Ali, et al.
Published: (2026)
by: Shafique, Muhammad Ali, et al.
Published: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs
by: Oliveira, Rafael C. T.
Published: (2026)
by: Oliveira, Rafael C. T.
Published: (2026)
IntentGrasp: A Comprehensive Benchmark for Intent Understanding
by: Yin, Yuwei, et al.
Published: (2026)
by: Yin, Yuwei, et al.
Published: (2026)
Similar Items
-
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
by: Manikantan, Kawshik, et al.
Published: (2024) -
Below-Chance Blindness: Prompted Underperformance in Small LLMs Produces Positional Bias Rather than Answer Avoidance
by: Cacioli, Jon-Paul
Published: (2026) -
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025) -
Evaluating Long Range Dependency Handling in Code Generation LLMs
by: Assogba, Yannick, et al.
Published: (2024) -
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)