Even GPT-5.2 Can't Count to Five: The Case for Zero-Error Horizons in Trustworthy LLMs
Fuente:
arXiv
Saved in:
| Main Author: | Sato, Ryoma |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can We Count on LLMs? The Fixed-Effect Fallacy and Claims of GPT-4 Capabilities
by: Ball, Thomas, et al.
Published: (2024)
by: Ball, Thomas, et al.
Published: (2024)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
by: Manvi, Rohin, et al.
Published: (2024)
by: Manvi, Rohin, et al.
Published: (2024)
User-Side Realization
by: Sato, Ryoma
Published: (2024)
by: Sato, Ryoma
Published: (2024)
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
by: Riachi, Roland, et al.
Published: (2025)
by: Riachi, Roland, et al.
Published: (2025)
When Can Transformers Count to n?
by: Yehudai, Gilad, et al.
Published: (2024)
by: Yehudai, Gilad, et al.
Published: (2024)
Harmonic LLMs are Trustworthy
by: Kersting, Nicholas S., et al.
Published: (2024)
by: Kersting, Nicholas S., et al.
Published: (2024)
CHILL at SemEval-2025 Task 2: You Can't Just Throw Entities and Hope -- Make Your LLM to Get Them Right
by: Lee, Jaebok, et al.
Published: (2025)
by: Lee, Jaebok, et al.
Published: (2025)
Can we trust the evaluation on ChatGPT?
by: Aiyappa, Rachith, et al.
Published: (2023)
by: Aiyappa, Rachith, et al.
Published: (2023)
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
by: Ming, Yifei, et al.
Published: (2024)
by: Ming, Yifei, et al.
Published: (2024)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
Can't Remember Details in Long Documents? You Need Some R&R
by: Agrawal, Devanshu, et al.
Published: (2024)
by: Agrawal, Devanshu, et al.
Published: (2024)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
by: Yu, Erxin, et al.
Published: (2025)
by: Yu, Erxin, et al.
Published: (2025)
Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
by: Jia, Furong, et al.
Published: (2025)
by: Jia, Furong, et al.
Published: (2025)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
by: Lee, Heekyung, et al.
Published: (2025)
by: Lee, Heekyung, et al.
Published: (2025)
Can LLMs Follow Simple Rules?
by: Mu, Norman, et al.
Published: (2023)
by: Mu, Norman, et al.
Published: (2023)
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
by: Chen, Junying, et al.
Published: (2024)
by: Chen, Junying, et al.
Published: (2024)
HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs
by: Chen, Junying, et al.
Published: (2023)
by: Chen, Junying, et al.
Published: (2023)
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
by: Wang, Boshi, et al.
Published: (2024)
by: Wang, Boshi, et al.
Published: (2024)
LLMs are not Zero-Shot Reasoners for Biomedical Information Extraction
by: Nagar, Aishik, et al.
Published: (2024)
by: Nagar, Aishik, et al.
Published: (2024)
BgGPT 1.0: Extending English-centric LLMs to other languages
by: Alexandrov, Anton, et al.
Published: (2024)
by: Alexandrov, Anton, et al.
Published: (2024)
ContextGPT: Infusing LLMs Knowledge into Neuro-Symbolic Activity Recognition Models
by: Arrotta, Luca, et al.
Published: (2024)
by: Arrotta, Luca, et al.
Published: (2024)
Can GPT Improve the State of Prior Authorization via Guideline Based Automated Question Answering?
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
Can Post-Training Transform LLMs into Causal Reasoners?
by: Chen, Junqi, et al.
Published: (2026)
by: Chen, Junqi, et al.
Published: (2026)
Can GRPO Help LLMs Transcend Their Pretraining Origin?
by: Ni, Kangqi, et al.
Published: (2025)
by: Ni, Kangqi, et al.
Published: (2025)
Can LLMs Convert Graphs to Text-Attributed Graphs?
by: Wang, Zehong, et al.
Published: (2024)
by: Wang, Zehong, et al.
Published: (2024)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
by: Park, Jungsoo, et al.
Published: (2025)
by: Park, Jungsoo, et al.
Published: (2025)
Single layer tiny Co$^4$ outpaces GPT-2 and GPT-BERT
by: Zain, Noor Ul, et al.
Published: (2025)
by: Zain, Noor Ul, et al.
Published: (2025)
Enough Coin Flips Can Make LLMs Act Bayesian
by: Gupta, Ritwik, et al.
Published: (2025)
by: Gupta, Ritwik, et al.
Published: (2025)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
by: Kirchhof, Michael, et al.
Published: (2025)
by: Kirchhof, Michael, et al.
Published: (2025)
Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
by: Xiong, Boya, et al.
Published: (2025)
by: Xiong, Boya, et al.
Published: (2025)
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
by: Hans, Abhimanyu, et al.
Published: (2024)
by: Hans, Abhimanyu, et al.
Published: (2024)
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
by: Tseng, Albert, et al.
Published: (2024)
by: Tseng, Albert, et al.
Published: (2024)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
by: Wu, Zhengxuan, et al.
Published: (2025)
by: Wu, Zhengxuan, et al.
Published: (2025)
Revisiting Chain-of-Thought Prompting: Zero-shot Can Be Stronger than Few-shot
by: Cheng, Xiang, et al.
Published: (2025)
by: Cheng, Xiang, et al.
Published: (2025)
Zero-Shot End-to-End Relation Extraction in Chinese: A Comparative Study of Gemini, LLaMA and ChatGPT
by: Du, Shaoshuai, et al.
Published: (2025)
by: Du, Shaoshuai, et al.
Published: (2025)
LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report
by: Zhao, Justin, et al.
Published: (2024)
by: Zhao, Justin, et al.
Published: (2024)
Training-free Graph Neural Networks and the Power of Labels as Features
by: Sato, Ryoma
Published: (2024)
by: Sato, Ryoma
Published: (2024)
TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models
by: Mo, Yichuan, et al.
Published: (2026)
by: Mo, Yichuan, et al.
Published: (2026)
Similar Items
-
Can We Count on LLMs? The Fixed-Effect Fallacy and Claims of GPT-4 Capabilities
by: Ball, Thomas, et al.
Published: (2024) -
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
by: Sun, Yiyou, et al.
Published: (2025) -
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
by: Manvi, Rohin, et al.
Published: (2024) -
User-Side Realization
by: Sato, Ryoma
Published: (2024) -
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
by: Riachi, Roland, et al.
Published: (2025)