Gauging Overprecision in LLMs: An Empirical Study
Fuente:
arXiv
Saved in:
| Main Authors: | Bahaj, Adil, Rahimi, Hamed, Chetouani, Mohamed, Ghogho, Mounir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
by: Bahaj, Adil, et al.
Published: (2025)
by: Bahaj, Adil, et al.
Published: (2025)
AsthmaBot: Multi-modal, Multi-Lingual Retrieval Augmented Generation For Asthma Patient Support
by: Bahaj, Adil, et al.
Published: (2024)
by: Bahaj, Adil, et al.
Published: (2024)
MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering
by: Bahaj, Adil, et al.
Published: (2025)
by: Bahaj, Adil, et al.
Published: (2025)
USER-VLM 360: Personalized Vision Language Models with User-aware Tuning for Social Human-Robot Interactions
by: Rahimi, Hamed, et al.
Published: (2025)
by: Rahimi, Hamed, et al.
Published: (2025)
Skill Demand Forecasting Using Temporal Knowledge Graph Embeddings
by: Fettach, Yousra, et al.
Published: (2025)
by: Fettach, Yousra, et al.
Published: (2025)
A multilingual training strategy for low resource Text to Speech
by: Amalas, Asma, et al.
Published: (2024)
by: Amalas, Asma, et al.
Published: (2024)
Cross-Lingual Multi-Granularity Framework for Interpretable Parkinson's Disease Diagnosis from Speech
by: Tougui, Ilias, et al.
Published: (2025)
by: Tougui, Ilias, et al.
Published: (2025)
Demographic User Modeling for Social Robotics with Multimodal Pre-trained Models
by: Rahimi, Hamed, et al.
Published: (2025)
by: Rahimi, Hamed, et al.
Published: (2025)
Revisiting Neighborhood Aggregation in Graph Neural Networks for Node Classification using Statistical Signal Processing
by: Ghogho, Mounir
Published: (2024)
by: Ghogho, Mounir
Published: (2024)
Reasoning LLMs for User-Aware Multimodal Conversational Agents
by: Rahimi, Hamed, et al.
Published: (2025)
by: Rahimi, Hamed, et al.
Published: (2025)
Designing Robust Software Sensors for Nonlinear Systems via Neural Networks and Adaptive Sliding Mode Control
by: Farkane, Ayoub, et al.
Published: (2025)
by: Farkane, Ayoub, et al.
Published: (2025)
IntentVLM: Open-Vocabulary Intention Recognition through Forward-Inverse Modeling with Video-Language Models
by: Rahimi, Hamed, et al.
Published: (2026)
by: Rahimi, Hamed, et al.
Published: (2026)
An Empirical Study of In-context Learning in LLMs for Machine Translation
by: Chitale, Pranjal A., et al.
Published: (2024)
by: Chitale, Pranjal A., et al.
Published: (2024)
Balancing Stability and Plasticity in Sequentially Trained Early-Exiting Neural Networks
by: Zniber, Alaa, et al.
Published: (2026)
by: Zniber, Alaa, et al.
Published: (2026)
HARMONI: Multimodal Personalization of Multi-User Human-Robot Interactions with LLMs
by: Malécot, Jeanne, et al.
Published: (2026)
by: Malécot, Jeanne, et al.
Published: (2026)
Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare
by: Christophe, Clément, et al.
Published: (2026)
by: Christophe, Clément, et al.
Published: (2026)
I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
by: Grislain, Clemence, et al.
Published: (2025)
by: Grislain, Clemence, et al.
Published: (2025)
Strong Memory, Weak Control: An Empirical Study of Executive Functioning in LLMs
by: de Langis, Karin, et al.
Published: (2025)
by: de Langis, Karin, et al.
Published: (2025)
Explore the Potential of LLMs in Misinformation Detection: An Empirical Study
by: Chen, Mengyang, et al.
Published: (2023)
by: Chen, Mengyang, et al.
Published: (2023)
Are Social Sentiments Inherent in LLMs? An Empirical Study on Extraction of Inter-demographic Sentiments
by: Tanaka, Kunitomo, et al.
Published: (2024)
by: Tanaka, Kunitomo, et al.
Published: (2024)
Entropy-Regularized Adjoint Matching for Offline Reinforcement Learning
by: Ghanem, Abdelghani, et al.
Published: (2026)
by: Ghanem, Abdelghani, et al.
Published: (2026)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
by: Xiong, Miao, et al.
Published: (2023)
by: Xiong, Miao, et al.
Published: (2023)
Can Reasoning LLMs Enhance Clinical Document Classification?
by: Mustafa, Akram, et al.
Published: (2025)
by: Mustafa, Akram, et al.
Published: (2025)
Encoding Predictability and Legibility for Style-Conditioned Diffusion Policy
by: Crétides, Adrien Jacquet, et al.
Published: (2026)
by: Crétides, Adrien Jacquet, et al.
Published: (2026)
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
Too Late to Train, Too Early To Use? A Study on Necessity and Viability of Low-Resource Bengali LLMs
by: Mahfuz, Tamzeed, et al.
Published: (2024)
by: Mahfuz, Tamzeed, et al.
Published: (2024)
What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metric
by: Kerkouri, Mohamed Amine, et al.
Published: (2026)
by: Kerkouri, Mohamed Amine, et al.
Published: (2026)
Can Many-Shot In-Context Learning Help LLMs as Evaluators? A Preliminary Empirical Study
by: Song, Mingyang, et al.
Published: (2024)
by: Song, Mingyang, et al.
Published: (2024)
Evaluating Hierarchical Clinical Document Classification Using Reasoning-Based LLMs
by: Mustafa, Akram, et al.
Published: (2025)
by: Mustafa, Akram, et al.
Published: (2025)
Abdelhak at SemEval-2024 Task 9 : Decoding Brainteasers, The Efficacy of Dedicated Models Versus ChatGPT
by: Kelious, Abdelhak, et al.
Published: (2024)
by: Kelious, Abdelhak, et al.
Published: (2024)
Aligning (Medical) LLMs for (Counterfactual) Fairness
by: Poulain, Raphael, et al.
Published: (2024)
by: Poulain, Raphael, et al.
Published: (2024)
Sample Design Engineering: An Empirical Study of What Makes Good Downstream Fine-Tuning Samples for LLMs
by: Guo, Biyang, et al.
Published: (2024)
by: Guo, Biyang, et al.
Published: (2024)
The Thinking Spectrum: An Empirical Study of Tunable Reasoning in LLMs through Model Merging
by: Lan, Xiaochong, et al.
Published: (2025)
by: Lan, Xiaochong, et al.
Published: (2025)
L-ReLF: A Framework for Lexical Dataset Creation
by: Sedrati, Anass, et al.
Published: (2026)
by: Sedrati, Anass, et al.
Published: (2026)
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning
by: Chen, Zui, et al.
Published: (2024)
by: Chen, Zui, et al.
Published: (2024)
Early Exiting Predictive Coding Neural Networks for Edge AI
by: Zniber, Alaa, et al.
Published: (2023)
by: Zniber, Alaa, et al.
Published: (2023)
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
by: Bavaresco, Anna, et al.
Published: (2024)
by: Bavaresco, Anna, et al.
Published: (2024)
Enabling Scalable Evaluation of Bias Patterns in Medical LLMs
by: Fayyaz, Hamed, et al.
Published: (2024)
by: Fayyaz, Hamed, et al.
Published: (2024)
Do Large Language Models Rank Fairly? An Empirical Study on the Fairness of LLMs as Rankers
by: Wang, Yuan, et al.
Published: (2024)
by: Wang, Yuan, et al.
Published: (2024)
Multi-Agent LLMs for Generating Research Limitations
by: Azher, Ibrahim Al, et al.
Published: (2025)
by: Azher, Ibrahim Al, et al.
Published: (2025)
Similar Items
-
PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
by: Bahaj, Adil, et al.
Published: (2025) -
AsthmaBot: Multi-modal, Multi-Lingual Retrieval Augmented Generation For Asthma Patient Support
by: Bahaj, Adil, et al.
Published: (2024) -
MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering
by: Bahaj, Adil, et al.
Published: (2025) -
USER-VLM 360: Personalized Vision Language Models with User-aware Tuning for Social Human-Robot Interactions
by: Rahimi, Hamed, et al.
Published: (2025) -
Skill Demand Forecasting Using Temporal Knowledge Graph Embeddings
by: Fettach, Yousra, et al.
Published: (2025)