When Not to Answer: Evaluating Prompts on GPT Models for Effective Abstention in Unanswerable Math Word Problems
Fuente:
arXiv
Saved in:
| Main Authors: | Saadat, Asir, Sogir, Tasmia Binte, Chowdhury, Md Taukir Azam, Aziz, Syem |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VisionTrap: Unanswerable Questions On Visual Data
by: Saadat, Asir, et al.
Published: (2025)
by: Saadat, Asir, et al.
Published: (2025)
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation
by: Ouyang, Jialin
Published: (2025)
by: Ouyang, Jialin
Published: (2025)
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
by: Sun, Yuhong, et al.
Published: (2024)
by: Sun, Yuhong, et al.
Published: (2024)
Contextual Breach: Assessing the Robustness of Transformer-based QA Models
by: Saadat, Asir, et al.
Published: (2024)
by: Saadat, Asir, et al.
Published: (2024)
Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems
by: Madhusudhan, Nishanth, et al.
Published: (2026)
by: Madhusudhan, Nishanth, et al.
Published: (2026)
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
by: Abdaljalil, Samir, et al.
Published: (2026)
by: Abdaljalil, Samir, et al.
Published: (2026)
Effective Skill Unlearning through Intervention and Abstention
by: Li, Yongce, et al.
Published: (2025)
by: Li, Yongce, et al.
Published: (2025)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
by: Adib, Shefayat E Shams, et al.
Published: (2026)
by: Adib, Shefayat E Shams, et al.
Published: (2026)
Augmenting Math Word Problems via Iterative Question Composing
by: Liu, Haoxiong, et al.
Published: (2024)
by: Liu, Haoxiong, et al.
Published: (2024)
Leveraging Large Language Models for Bengali Math Word Problem Solving with Chain of Thought Reasoning
by: Paul, Bidyarthi, et al.
Published: (2025)
by: Paul, Bidyarthi, et al.
Published: (2025)
Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models
by: Madhusudhan, Nishanth, et al.
Published: (2024)
by: Madhusudhan, Nishanth, et al.
Published: (2024)
Answering the Unanswerable Is to Err Knowingly: Analyzing and Mitigating Abstention Failures in Large Reasoning Models
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Structured Reasoning with Tree-of-Thoughts for Bengali Math Word Problems
by: Mahmood, Aurprita, et al.
Published: (2025)
by: Mahmood, Aurprita, et al.
Published: (2025)
Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams
by: Caraeni, Adriana, et al.
Published: (2024)
by: Caraeni, Adriana, et al.
Published: (2024)
Empowering Bengali Education with AI: Solving Bengali Math Word Problems through Transformer Models
by: Era, Jalisha Jashim, et al.
Published: (2025)
by: Era, Jalisha Jashim, et al.
Published: (2025)
Prompting and Fine-tuning Large Language Models for Automated Code Review Comment Generation
by: Haider, Md. Asif, et al.
Published: (2024)
by: Haider, Md. Asif, et al.
Published: (2024)
Geometry-Calibrated Conformal Abstention for Language Models
by: Xu, Rui, et al.
Published: (2026)
by: Xu, Rui, et al.
Published: (2026)
Ontology-Guided Query Expansion for Biomedical Document Retrieval using Large Language Models
by: Nazi, Zabir Al, et al.
Published: (2025)
by: Nazi, Zabir Al, et al.
Published: (2025)
MathChat: Converse to Tackle Challenging Math Problems with LLM Agents
by: Wu, Yiran, et al.
Published: (2023)
by: Wu, Yiran, et al.
Published: (2023)
MKA: Leveraging Cross-Lingual Consensus for Model Abstention
by: Duwal, Sharad
Published: (2025)
by: Duwal, Sharad
Published: (2025)
Adversarial Math Word Problem Generation
by: Xie, Roy, et al.
Published: (2024)
by: Xie, Roy, et al.
Published: (2024)
Fill in the Blank: Exploring and Enhancing LLM Capabilities for Backward Reasoning in Math Word Problems
by: Deb, Aniruddha, et al.
Published: (2023)
by: Deb, Aniruddha, et al.
Published: (2023)
Data Augmentation with In-Context Learning and Comparative Evaluation in Math Word Problem Solving
by: Yigit, Gulsum, et al.
Published: (2024)
by: Yigit, Gulsum, et al.
Published: (2024)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
Elementary Math Word Problem Generation using Large Language Models
by: Ariyarathne, Nimesh, et al.
Published: (2025)
by: Ariyarathne, Nimesh, et al.
Published: (2025)
What Makes Math Word Problems Challenging for LLMs?
by: Srivatsa, KV Aditya, et al.
Published: (2024)
by: Srivatsa, KV Aditya, et al.
Published: (2024)
Expression Syntax Information Bottleneck for Math Word Problems
by: Xiong, Jing, et al.
Published: (2023)
by: Xiong, Jing, et al.
Published: (2023)
Self-consistent Reasoning For Solving Math Word Problems
by: Xiong, Jing, et al.
Published: (2022)
by: Xiong, Jing, et al.
Published: (2022)
When to Speak, When to Abstain: Contrastive Decoding with Abstention
by: Kim, Hyuhng Joon, et al.
Published: (2024)
by: Kim, Hyuhng Joon, et al.
Published: (2024)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
by: Rahman, Subhey Sadi, et al.
Published: (2025)
by: Rahman, Subhey Sadi, et al.
Published: (2025)
ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation
by: Zhao, Raoyuan, et al.
Published: (2026)
by: Zhao, Raoyuan, et al.
Published: (2026)
Rewarding Intellectual Humility Learning When Not To Answer In Large Language Models
by: Jha, Abha, et al.
Published: (2026)
by: Jha, Abha, et al.
Published: (2026)
Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffs
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
Learning by Analogy: Enhancing Few-Shot Prompting for Math Word Problem Solving with Computational Graph-Based Retrieval
by: Yang, Xiaocong, et al.
Published: (2024)
by: Yang, Xiaocong, et al.
Published: (2024)
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
by: Baan, Joris, et al.
Published: (2026)
by: Baan, Joris, et al.
Published: (2026)
Solving Math Word Problems via Cooperative Reasoning induced Language Models
by: Zhu, Xinyu, et al.
Published: (2022)
by: Zhu, Xinyu, et al.
Published: (2022)
EvoGPT-f: An Evolutionary GPT Framework for Benchmarking Formal Math Languages
by: Mercer, Johnathan
Published: (2024)
by: Mercer, Johnathan
Published: (2024)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
by: Yao, Siyang, et al.
Published: (2026)
by: Yao, Siyang, et al.
Published: (2026)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
by: Mahdavi, Sadegh, et al.
Published: (2025)
by: Mahdavi, Sadegh, et al.
Published: (2025)
Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs
by: Mahran, Mariam, et al.
Published: (2025)
by: Mahran, Mariam, et al.
Published: (2025)
Similar Items
-
VisionTrap: Unanswerable Questions On Visual Data
by: Saadat, Asir, et al.
Published: (2025) -
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation
by: Ouyang, Jialin
Published: (2025) -
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
by: Sun, Yuhong, et al.
Published: (2024) -
Contextual Breach: Assessing the Robustness of Transformer-based QA Models
by: Saadat, Asir, et al.
Published: (2024) -
Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems
by: Madhusudhan, Nishanth, et al.
Published: (2026)