Don't Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Jinzhe, Li, Gengxu, Chang, Yi, Wu, Yuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Refining Critical Thinking in LLM Code Generation: A Faulty Premise-based Evaluation Framework
di: Li, Jialin, et al.
Pubblicazione: (2025)
di: Li, Jialin, et al.
Pubblicazione: (2025)
Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning
di: Qin, Yuehan, et al.
Pubblicazione: (2025)
di: Qin, Yuehan, et al.
Pubblicazione: (2025)
Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability
di: Yang, Haiqi, et al.
Pubblicazione: (2025)
di: Yang, Haiqi, et al.
Pubblicazione: (2025)
Premise Order Matters in Reasoning with Large Language Models
di: Chen, Xinyun, et al.
Pubblicazione: (2024)
di: Chen, Xinyun, et al.
Pubblicazione: (2024)
RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises
di: Zhai, Zenan, et al.
Pubblicazione: (2025)
di: Zhai, Zenan, et al.
Pubblicazione: (2025)
Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language Models
di: Yuan, Hongbang, et al.
Pubblicazione: (2024)
di: Yuan, Hongbang, et al.
Pubblicazione: (2024)
Length-Controlled Margin-Based Preference Optimization without Reference Model
di: Li, Gengxu, et al.
Pubblicazione: (2025)
di: Li, Gengxu, et al.
Pubblicazione: (2025)
Secure On-Premise Deployment of Open-Weights Large Language Models in Radiology: An Isolation-First Architecture with Prospective Pilot Evaluation
di: Nowak, Sebastian, et al.
Pubblicazione: (2026)
di: Nowak, Sebastian, et al.
Pubblicazione: (2026)
DecoPrompt : Decoding Prompts Reduces Hallucinations when Large Language Models Meet False Premises
di: Xu, Nan, et al.
Pubblicazione: (2024)
di: Xu, Nan, et al.
Pubblicazione: (2024)
MedAide: Leveraging Large Language Models for On-Premise Medical Assistance on Edge Devices
di: Basit, Abdul, et al.
Pubblicazione: (2024)
di: Basit, Abdul, et al.
Pubblicazione: (2024)
Abductive Inference in Retrieval-Augmented Language Models: Generating and Validating Missing Premises
di: Lin, Shiyin
Pubblicazione: (2025)
di: Lin, Shiyin
Pubblicazione: (2025)
Judge Before Answer: Can MLLM Discern the False Premise in Question?
di: Li, Jidong, et al.
Pubblicazione: (2025)
di: Li, Jidong, et al.
Pubblicazione: (2025)
Making Implicit Premises Explicit in Logical Understanding of Enthymemes
di: Feng, Xuyao, et al.
Pubblicazione: (2026)
di: Feng, Xuyao, et al.
Pubblicazione: (2026)
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
di: Fan, Chenrui, et al.
Pubblicazione: (2025)
di: Fan, Chenrui, et al.
Pubblicazione: (2025)
Learning an Effective Premise Retrieval Model for Efficient Mathematical Formalization
di: Tao, Yicheng, et al.
Pubblicazione: (2025)
di: Tao, Yicheng, et al.
Pubblicazione: (2025)
From Hypothesis to Premises: LLM-based Backward Logical Reasoning with Selective Symbolic Translation
di: Li, Qingchuan, et al.
Pubblicazione: (2025)
di: Li, Qingchuan, et al.
Pubblicazione: (2025)
VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions
di: Su, Hung-Ting, et al.
Pubblicazione: (2026)
di: Su, Hung-Ting, et al.
Pubblicazione: (2026)
MultiHoax: A Dataset of Multi-hop False-Premise Questions
di: Shafiei, Mohammadamin, et al.
Pubblicazione: (2025)
di: Shafiei, Mohammadamin, et al.
Pubblicazione: (2025)
Combining Textual and Structural Information for Premise Selection in Lean
di: Petrovčič, Job, et al.
Pubblicazione: (2025)
di: Petrovčič, Job, et al.
Pubblicazione: (2025)
PAVE: Premise-Aware Validation and Editing for Retrieval-Augmented LLMs
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
Self-Evolving Critique Abilities in Large Language Models
di: Tang, Zhengyang, et al.
Pubblicazione: (2025)
di: Tang, Zhengyang, et al.
Pubblicazione: (2025)
Premise-Augmented Reasoning Chains Improve Error Identification in Math reasoning with LLMs
di: Mukherjee, Sagnik, et al.
Pubblicazione: (2025)
di: Mukherjee, Sagnik, et al.
Pubblicazione: (2025)
KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions
di: Zhu, Yanxu, et al.
Pubblicazione: (2024)
di: Zhu, Yanxu, et al.
Pubblicazione: (2024)
From Sycophancy to Sensemaking: Premise Governance for Human-AI Decision Making
di: Jain, Raunak
Pubblicazione: (2026)
di: Jain, Raunak
Pubblicazione: (2026)
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
di: Yan, Shaotian, et al.
Pubblicazione: (2025)
di: Yan, Shaotian, et al.
Pubblicazione: (2025)
Align, Don't Divide: Revisiting the LoRA Architecture in Multi-Task Learning
di: Liu, Jinda, et al.
Pubblicazione: (2025)
di: Liu, Jinda, et al.
Pubblicazione: (2025)
MoPS: Modular Story Premise Synthesis for Open-Ended Automatic Story Generation
di: Ma, Yan, et al.
Pubblicazione: (2024)
di: Ma, Yan, et al.
Pubblicazione: (2024)
Large Language Model Evaluation via Matrix Nuclear-Norm
di: Li, Yahan, et al.
Pubblicazione: (2024)
di: Li, Yahan, et al.
Pubblicazione: (2024)
R-Tuning: Instructing Large Language Models to Say `I Don't Know'
di: Zhang, Hanning, et al.
Pubblicazione: (2023)
di: Zhang, Hanning, et al.
Pubblicazione: (2023)
Don't Buy it! Reassessing the Ad Understanding Abilities of Contrastive Multimodal Models
di: Bavaresco, A., et al.
Pubblicazione: (2024)
di: Bavaresco, A., et al.
Pubblicazione: (2024)
StructFlowBench: A Structured Flow Benchmark for Multi-turn Instruction Following
di: Li, Jinnan, et al.
Pubblicazione: (2025)
di: Li, Jinnan, et al.
Pubblicazione: (2025)
Exploring the Limits of Fine-grained LLM-based Physics Inference via Premise Removal Interventions
di: Meadows, Jordan, et al.
Pubblicazione: (2024)
di: Meadows, Jordan, et al.
Pubblicazione: (2024)
From Articles to Premises: Building PrimeFacts, an Extraction Methodology and Resource for Fact-Checking Evidence
di: Sahitaj, Premtim, et al.
Pubblicazione: (2026)
di: Sahitaj, Premtim, et al.
Pubblicazione: (2026)
XTRUST: On the Multilingual Trustworthiness of Large Language Models
di: Li, Yahan, et al.
Pubblicazione: (2024)
di: Li, Yahan, et al.
Pubblicazione: (2024)
Remask, Don't Replace: Token-to-Mask Refinement in Diffusion Large Language Models
di: Yao, Lin
Pubblicazione: (2026)
di: Yao, Lin
Pubblicazione: (2026)
EPM-RL: Reinforcement Learning for On-Premise Product Mapping in E-Commerce
di: Yu, Minhyeong, et al.
Pubblicazione: (2026)
di: Yu, Minhyeong, et al.
Pubblicazione: (2026)
Show, Don't Tell: Evaluating Large Language Models Beyond Textual Understanding with ChildPlay
di: de Carvalho, Gonçalo Hora, et al.
Pubblicazione: (2024)
di: de Carvalho, Gonçalo Hora, et al.
Pubblicazione: (2024)
THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models
di: Li, Zhiyuan, et al.
Pubblicazione: (2025)
di: Li, Zhiyuan, et al.
Pubblicazione: (2025)
Thought-Path Contrastive Learning via Premise-Oriented Data Augmentation for Logical Reading Comprehension
di: Wang, Chenxu, et al.
Pubblicazione: (2024)
di: Wang, Chenxu, et al.
Pubblicazione: (2024)
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
di: Hariri, Mohsen, et al.
Pubblicazione: (2025)
di: Hariri, Mohsen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Refining Critical Thinking in LLM Code Generation: A Faulty Premise-based Evaluation Framework
di: Li, Jialin, et al.
Pubblicazione: (2025) -
Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning
di: Qin, Yuehan, et al.
Pubblicazione: (2025) -
Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability
di: Yang, Haiqi, et al.
Pubblicazione: (2025) -
Premise Order Matters in Reasoning with Large Language Models
di: Chen, Xinyun, et al.
Pubblicazione: (2024) -
RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises
di: Zhai, Zenan, et al.
Pubblicazione: (2025)