Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Haq, Saiful, Chhaya, Niyati, Pandey, Piyush, Bhattacharya, Pushpak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles
von: Bhattacharya, Antara Raaghavi, et al.
Veröffentlicht: (2025)
von: Bhattacharya, Antara Raaghavi, et al.
Veröffentlicht: (2025)
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
Analysis of instruction-based LLMs' capabilities to score and judge text-input problems in an academic setting
von: Ramirez-Garcia, Valeria, et al.
Veröffentlicht: (2025)
von: Ramirez-Garcia, Valeria, et al.
Veröffentlicht: (2025)
ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model
von: Lan, Wuyang, et al.
Veröffentlicht: (2025)
von: Lan, Wuyang, et al.
Veröffentlicht: (2025)
Can LLMs perform structured graph reasoning?
von: Agrawal, Palaash, et al.
Veröffentlicht: (2024)
von: Agrawal, Palaash, et al.
Veröffentlicht: (2024)
MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication
von: Sambara, Sraavya, et al.
Veröffentlicht: (2026)
von: Sambara, Sraavya, et al.
Veröffentlicht: (2026)
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M
von: Pant, Piyush
Veröffentlicht: (2025)
von: Pant, Piyush
Veröffentlicht: (2025)
Are LLMs classical or nonmonotonic reasoners? Lessons from generics
von: Leidinger, Alina, et al.
Veröffentlicht: (2024)
von: Leidinger, Alina, et al.
Veröffentlicht: (2024)
Can formal argumentative reasoning enhance LLMs performances?
von: Castagna, Federico, et al.
Veröffentlicht: (2024)
von: Castagna, Federico, et al.
Veröffentlicht: (2024)
Mental Disorder Classification via Temporal Representation of Text
von: Kumar, Raja, et al.
Veröffentlicht: (2024)
von: Kumar, Raja, et al.
Veröffentlicht: (2024)
SwissNYF: Tool Grounded LLM Agents for Black Box Setting
von: Kumar, Somnath Sendhil, et al.
Veröffentlicht: (2024)
von: Kumar, Somnath Sendhil, et al.
Veröffentlicht: (2024)
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
Critical-Questions-of-Thought: Steering LLM reasoning with Argumentative Querying
von: Castagna, Federico, et al.
Veröffentlicht: (2024)
von: Castagna, Federico, et al.
Veröffentlicht: (2024)
Plan with Code: Comparing approaches for robust NL to DSL generation
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
Discerning minds or generic tutors? Evaluating instructional guidance capabilities in Socratic LLMs
von: Liu, Ying, et al.
Veröffentlicht: (2025)
von: Liu, Ying, et al.
Veröffentlicht: (2025)
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?
von: Lei, Zhikai, et al.
Veröffentlicht: (2024)
von: Lei, Zhikai, et al.
Veröffentlicht: (2024)
MARS: toward more efficient multi-agent collaboration for LLM reasoning
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?
von: Sil, Pritam, et al.
Veröffentlicht: (2026)
von: Sil, Pritam, et al.
Veröffentlicht: (2026)
Are complicated loss functions necessary for teaching LLMs to reason?
von: Carrino, Gabriele, et al.
Veröffentlicht: (2026)
von: Carrino, Gabriele, et al.
Veröffentlicht: (2026)
Improving User Behavior Prediction: Leveraging Annotator Metadata in Supervised Machine Learning Models
von: Ng, Lynnette Hui Xian, et al.
Veröffentlicht: (2025)
von: Ng, Lynnette Hui Xian, et al.
Veröffentlicht: (2025)
Broaden your SCOPE! Efficient Multi-turn Conversation Planning for LLMs with Semantic Space
von: Chen, Zhiliang, et al.
Veröffentlicht: (2025)
von: Chen, Zhiliang, et al.
Veröffentlicht: (2025)
Next Token Perception Score: Analytical Assessment of your LLM Perception Skills
von: Cheng, Yu-Ang, et al.
Veröffentlicht: (2025)
von: Cheng, Yu-Ang, et al.
Veröffentlicht: (2025)
Leveraging LLMs for Bangla Grammar Error Correction:Error Categorization, Synthetic Data, and Model Evaluation
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2024)
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2024)
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
von: Khan, Haidar, et al.
Veröffentlicht: (2025)
von: Khan, Haidar, et al.
Veröffentlicht: (2025)
How new data permeates LLM knowledge and how to dilute it
von: Sun, Chen, et al.
Veröffentlicht: (2025)
von: Sun, Chen, et al.
Veröffentlicht: (2025)
Recon, Answer, Verify: Agents in Search of Truth
von: Shukla, Satyam, et al.
Veröffentlicht: (2025)
von: Shukla, Satyam, et al.
Veröffentlicht: (2025)
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
von: Alyahya, Hisham A., et al.
Veröffentlicht: (2025)
von: Alyahya, Hisham A., et al.
Veröffentlicht: (2025)
Can LLM-Augmented autonomous agents cooperate?, An evaluation of their cooperative capabilities through Melting Pot
von: Mosquera, Manuel, et al.
Veröffentlicht: (2024)
von: Mosquera, Manuel, et al.
Veröffentlicht: (2024)
An evaluation of LLM code generation capabilities through graded exercises
von: Jiménez, Álvaro Barbero
Veröffentlicht: (2024)
von: Jiménez, Álvaro Barbero
Veröffentlicht: (2024)
Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs
von: Kim, Junsol, et al.
Veröffentlicht: (2026)
von: Kim, Junsol, et al.
Veröffentlicht: (2026)
Towards LLM-based optimization compilers. Can LLMs learn how to apply a single peephole optimization? Reasoning is all LLMs need!
von: Fang, Xiangxin, et al.
Veröffentlicht: (2024)
von: Fang, Xiangxin, et al.
Veröffentlicht: (2024)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
LuxVeri at GenAI Detection Task 3: Cross-Domain Detection of AI-Generated Text Using Inverse Perplexity-Weighted Ensemble of Fine-Tuned Transformer Models
von: Mobin, Md Kamrujjaman, et al.
Veröffentlicht: (2025)
von: Mobin, Md Kamrujjaman, et al.
Veröffentlicht: (2025)
LuxVeri at GenAI Detection Task 1: Inverse Perplexity Weighted Ensemble for Robust Detection of AI-Generated Text across English and Multilingual Contexts
von: Mobin, Md Kamrujjaman, et al.
Veröffentlicht: (2025)
von: Mobin, Md Kamrujjaman, et al.
Veröffentlicht: (2025)
PyBangla at BLP-2025 Task 2: Enhancing Bangla-to-Python Code Generation with Iterative Self-Correction and Multilingual Agents
von: Islam, Jahidul, et al.
Veröffentlicht: (2025)
von: Islam, Jahidul, et al.
Veröffentlicht: (2025)
Looks can be Deceptive: Distinguishing Repetition Disfluency from Reduplication
von: Ahmad, Arif, et al.
Veröffentlicht: (2024)
von: Ahmad, Arif, et al.
Veröffentlicht: (2024)
Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
von: Hegde, Pradyoth, et al.
Veröffentlicht: (2025)
von: Hegde, Pradyoth, et al.
Veröffentlicht: (2025)
Meaningless is better: hashing bias-inducing words in LLM prompts improves performance in logical reasoning and statistical learning
von: Chadimová, Milena, et al.
Veröffentlicht: (2024)
von: Chadimová, Milena, et al.
Veröffentlicht: (2024)
IoT-LLM: a framework for enhancing Large Language Model reasoning from real-world sensor data
von: An, Tuo, et al.
Veröffentlicht: (2024)
von: An, Tuo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles
von: Bhattacharya, Antara Raaghavi, et al.
Veröffentlicht: (2025) -
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025) -
Analysis of instruction-based LLMs' capabilities to score and judge text-input problems in an academic setting
von: Ramirez-Garcia, Valeria, et al.
Veröffentlicht: (2025) -
ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model
von: Lan, Wuyang, et al.
Veröffentlicht: (2025) -
Can LLMs perform structured graph reasoning?
von: Agrawal, Palaash, et al.
Veröffentlicht: (2024)