Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Haq, Saiful, Chhaya, Niyati, Pandey, Piyush, Bhattacharya, Pushpak |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles
di: Bhattacharya, Antara Raaghavi, et al.
Pubblicazione: (2025)
di: Bhattacharya, Antara Raaghavi, et al.
Pubblicazione: (2025)
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian
di: Syromiatnikov, Mykyta, et al.
Pubblicazione: (2025)
di: Syromiatnikov, Mykyta, et al.
Pubblicazione: (2025)
Analysis of instruction-based LLMs' capabilities to score and judge text-input problems in an academic setting
di: Ramirez-Garcia, Valeria, et al.
Pubblicazione: (2025)
di: Ramirez-Garcia, Valeria, et al.
Pubblicazione: (2025)
ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model
di: Lan, Wuyang, et al.
Pubblicazione: (2025)
di: Lan, Wuyang, et al.
Pubblicazione: (2025)
Can LLMs perform structured graph reasoning?
di: Agrawal, Palaash, et al.
Pubblicazione: (2024)
di: Agrawal, Palaash, et al.
Pubblicazione: (2024)
MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication
di: Sambara, Sraavya, et al.
Pubblicazione: (2026)
di: Sambara, Sraavya, et al.
Pubblicazione: (2026)
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M
di: Pant, Piyush
Pubblicazione: (2025)
di: Pant, Piyush
Pubblicazione: (2025)
Are LLMs classical or nonmonotonic reasoners? Lessons from generics
di: Leidinger, Alina, et al.
Pubblicazione: (2024)
di: Leidinger, Alina, et al.
Pubblicazione: (2024)
Can formal argumentative reasoning enhance LLMs performances?
di: Castagna, Federico, et al.
Pubblicazione: (2024)
di: Castagna, Federico, et al.
Pubblicazione: (2024)
Mental Disorder Classification via Temporal Representation of Text
di: Kumar, Raja, et al.
Pubblicazione: (2024)
di: Kumar, Raja, et al.
Pubblicazione: (2024)
SwissNYF: Tool Grounded LLM Agents for Black Box Setting
di: Kumar, Somnath Sendhil, et al.
Pubblicazione: (2024)
di: Kumar, Somnath Sendhil, et al.
Pubblicazione: (2024)
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
Critical-Questions-of-Thought: Steering LLM reasoning with Argumentative Querying
di: Castagna, Federico, et al.
Pubblicazione: (2024)
di: Castagna, Federico, et al.
Pubblicazione: (2024)
Plan with Code: Comparing approaches for robust NL to DSL generation
di: Bassamzadeh, Nastaran, et al.
Pubblicazione: (2024)
di: Bassamzadeh, Nastaran, et al.
Pubblicazione: (2024)
A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation
di: Bassamzadeh, Nastaran, et al.
Pubblicazione: (2024)
di: Bassamzadeh, Nastaran, et al.
Pubblicazione: (2024)
Discerning minds or generic tutors? Evaluating instructional guidance capabilities in Socratic LLMs
di: Liu, Ying, et al.
Pubblicazione: (2025)
di: Liu, Ying, et al.
Pubblicazione: (2025)
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?
di: Lei, Zhikai, et al.
Pubblicazione: (2024)
di: Lei, Zhikai, et al.
Pubblicazione: (2024)
MARS: toward more efficient multi-agent collaboration for LLM reasoning
di: Wang, Xiao, et al.
Pubblicazione: (2025)
di: Wang, Xiao, et al.
Pubblicazione: (2025)
Are complicated loss functions necessary for teaching LLMs to reason?
di: Carrino, Gabriele, et al.
Pubblicazione: (2026)
di: Carrino, Gabriele, et al.
Pubblicazione: (2026)
How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?
di: Sil, Pritam, et al.
Pubblicazione: (2026)
di: Sil, Pritam, et al.
Pubblicazione: (2026)
Improving User Behavior Prediction: Leveraging Annotator Metadata in Supervised Machine Learning Models
di: Ng, Lynnette Hui Xian, et al.
Pubblicazione: (2025)
di: Ng, Lynnette Hui Xian, et al.
Pubblicazione: (2025)
Broaden your SCOPE! Efficient Multi-turn Conversation Planning for LLMs with Semantic Space
di: Chen, Zhiliang, et al.
Pubblicazione: (2025)
di: Chen, Zhiliang, et al.
Pubblicazione: (2025)
Next Token Perception Score: Analytical Assessment of your LLM Perception Skills
di: Cheng, Yu-Ang, et al.
Pubblicazione: (2025)
di: Cheng, Yu-Ang, et al.
Pubblicazione: (2025)
Leveraging LLMs for Bangla Grammar Error Correction:Error Categorization, Synthetic Data, and Model Evaluation
di: Bhattacharyya, Pramit, et al.
Pubblicazione: (2024)
di: Bhattacharyya, Pramit, et al.
Pubblicazione: (2024)
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
di: Khan, Haidar, et al.
Pubblicazione: (2025)
di: Khan, Haidar, et al.
Pubblicazione: (2025)
How new data permeates LLM knowledge and how to dilute it
di: Sun, Chen, et al.
Pubblicazione: (2025)
di: Sun, Chen, et al.
Pubblicazione: (2025)
Recon, Answer, Verify: Agents in Search of Truth
di: Shukla, Satyam, et al.
Pubblicazione: (2025)
di: Shukla, Satyam, et al.
Pubblicazione: (2025)
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
di: Alyahya, Hisham A., et al.
Pubblicazione: (2025)
di: Alyahya, Hisham A., et al.
Pubblicazione: (2025)
Can LLM-Augmented autonomous agents cooperate?, An evaluation of their cooperative capabilities through Melting Pot
di: Mosquera, Manuel, et al.
Pubblicazione: (2024)
di: Mosquera, Manuel, et al.
Pubblicazione: (2024)
An evaluation of LLM code generation capabilities through graded exercises
di: Jiménez, Álvaro Barbero
Pubblicazione: (2024)
di: Jiménez, Álvaro Barbero
Pubblicazione: (2024)
Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs
di: Kim, Junsol, et al.
Pubblicazione: (2026)
di: Kim, Junsol, et al.
Pubblicazione: (2026)
Towards LLM-based optimization compilers. Can LLMs learn how to apply a single peephole optimization? Reasoning is all LLMs need!
di: Fang, Xiangxin, et al.
Pubblicazione: (2024)
di: Fang, Xiangxin, et al.
Pubblicazione: (2024)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
di: Yu, Ping, et al.
Pubblicazione: (2025)
di: Yu, Ping, et al.
Pubblicazione: (2025)
LuxVeri at GenAI Detection Task 3: Cross-Domain Detection of AI-Generated Text Using Inverse Perplexity-Weighted Ensemble of Fine-Tuned Transformer Models
di: Mobin, Md Kamrujjaman, et al.
Pubblicazione: (2025)
di: Mobin, Md Kamrujjaman, et al.
Pubblicazione: (2025)
LuxVeri at GenAI Detection Task 1: Inverse Perplexity Weighted Ensemble for Robust Detection of AI-Generated Text across English and Multilingual Contexts
di: Mobin, Md Kamrujjaman, et al.
Pubblicazione: (2025)
di: Mobin, Md Kamrujjaman, et al.
Pubblicazione: (2025)
PyBangla at BLP-2025 Task 2: Enhancing Bangla-to-Python Code Generation with Iterative Self-Correction and Multilingual Agents
di: Islam, Jahidul, et al.
Pubblicazione: (2025)
di: Islam, Jahidul, et al.
Pubblicazione: (2025)
Looks can be Deceptive: Distinguishing Repetition Disfluency from Reduplication
di: Ahmad, Arif, et al.
Pubblicazione: (2024)
di: Ahmad, Arif, et al.
Pubblicazione: (2024)
Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
di: Hegde, Pradyoth, et al.
Pubblicazione: (2025)
di: Hegde, Pradyoth, et al.
Pubblicazione: (2025)
Meaningless is better: hashing bias-inducing words in LLM prompts improves performance in logical reasoning and statistical learning
di: Chadimová, Milena, et al.
Pubblicazione: (2024)
di: Chadimová, Milena, et al.
Pubblicazione: (2024)
IoT-LLM: a framework for enhancing Large Language Model reasoning from real-world sensor data
di: An, Tuo, et al.
Pubblicazione: (2024)
di: An, Tuo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles
di: Bhattacharya, Antara Raaghavi, et al.
Pubblicazione: (2025) -
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian
di: Syromiatnikov, Mykyta, et al.
Pubblicazione: (2025) -
Analysis of instruction-based LLMs' capabilities to score and judge text-input problems in an academic setting
di: Ramirez-Garcia, Valeria, et al.
Pubblicazione: (2025) -
ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model
di: Lan, Wuyang, et al.
Pubblicazione: (2025) -
Can LLMs perform structured graph reasoning?
di: Agrawal, Palaash, et al.
Pubblicazione: (2024)