Reasoning about concepts with LLMs: Inconsistencies abound
Fuente:
arXiv
Guardado en:
| Autores principales: | Uceda-Sosa, Rosario, Ramamurthy, Karthikeyan Natesan, Chang, Maria, Singh, Moninder |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Cross-Examiner: Evaluating Consistency of Large Language Model-Generated Explanations
por: Villa, Danielle, et al.
Publicado: (2025)
por: Villa, Danielle, et al.
Publicado: (2025)
Ranking Large Language Models without Ground Truth
por: Dhurandhar, Amit, et al.
Publicado: (2024)
por: Dhurandhar, Amit, et al.
Publicado: (2024)
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
por: Achintalwar, Swapnaja, et al.
Publicado: (2024)
por: Achintalwar, Swapnaja, et al.
Publicado: (2024)
AI Steerability 360: A Toolkit for Steering Large Language Models
por: Miehling, Erik, et al.
Publicado: (2026)
por: Miehling, Erik, et al.
Publicado: (2026)
Mitigating Misalignment Contagion by Steering with Implicit Traits
por: Chang, Maria, et al.
Publicado: (2026)
por: Chang, Maria, et al.
Publicado: (2026)
Programming Refusal with Conditional Activation Steering
por: Lee, Bruce W., et al.
Publicado: (2024)
por: Lee, Bruce W., et al.
Publicado: (2024)
Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents
por: Ngong, Ivoline, et al.
Publicado: (2025)
por: Ngong, Ivoline, et al.
Publicado: (2025)
Sparsity May Be All You Need: Sparse Random Parameter Adaptation
por: Rios, Jesus, et al.
Publicado: (2025)
por: Rios, Jesus, et al.
Publicado: (2025)
Evaluating the Prompt Steerability of Large Language Models
por: Miehling, Erik, et al.
Publicado: (2024)
por: Miehling, Erik, et al.
Publicado: (2024)
Multi-Level Explanations for Generative Language Models
por: Paes, Lucas Monteiro, et al.
Publicado: (2024)
por: Paes, Lucas Monteiro, et al.
Publicado: (2024)
Reasoning about Affordances: Causal and Compositional Reasoning in LLMs
por: Gjerde, Magnus F., et al.
Publicado: (2025)
por: Gjerde, Magnus F., et al.
Publicado: (2025)
Few-shot Policy (de)composition in Conversational Question Answering
por: Erwin, Kyle, et al.
Publicado: (2025)
por: Erwin, Kyle, et al.
Publicado: (2025)
Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models
por: Yan, Qianqi, et al.
Publicado: (2025)
por: Yan, Qianqi, et al.
Publicado: (2025)
Improved Evidence Extraction and Metrics for Document Inconsistency Detection with LLMs
por: Tan, Nelvin, et al.
Publicado: (2026)
por: Tan, Nelvin, et al.
Publicado: (2026)
Trust Regions for Explanations via Black-Box Probabilistic Certification
por: Dhurandhar, Amit, et al.
Publicado: (2024)
por: Dhurandhar, Amit, et al.
Publicado: (2024)
BIS Reasoning 1.0: The First Large-Scale Japanese Benchmark for Belief-Inconsistent Syllogistic Reasoning
por: Nguyen, Ha-Thanh, et al.
Publicado: (2025)
por: Nguyen, Ha-Thanh, et al.
Publicado: (2025)
Language Models Coupled with Metacognition Can Outperform Reasoning Models
por: Khandelwal, Vedant, et al.
Publicado: (2025)
por: Khandelwal, Vedant, et al.
Publicado: (2025)
Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?
por: Addepalli, Sravanti, et al.
Publicado: (2024)
por: Addepalli, Sravanti, et al.
Publicado: (2024)
Inconsistencies in Masked Language Models
por: Young, Tom, et al.
Publicado: (2022)
por: Young, Tom, et al.
Publicado: (2022)
Time-Reversal Provides Unsupervised Feedback to LLMs
por: Varun, Yerram, et al.
Publicado: (2024)
por: Varun, Yerram, et al.
Publicado: (2024)
Reasoning about Intent for Ambiguous Requests
por: Saparina, Irina, et al.
Publicado: (2025)
por: Saparina, Irina, et al.
Publicado: (2025)
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
por: Lei, Chao, et al.
Publicado: (2025)
por: Lei, Chao, et al.
Publicado: (2025)
SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter
por: Jung-Mok, Lee, et al.
Publicado: (2026)
por: Jung-Mok, Lee, et al.
Publicado: (2026)
Knowledge Localization in Mixture-of-Experts LLMs Using Cross-Lingual Inconsistency
por: Bandarkar, Lucas, et al.
Publicado: (2026)
por: Bandarkar, Lucas, et al.
Publicado: (2026)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
por: Liu, Xiao, et al.
Publicado: (2024)
por: Liu, Xiao, et al.
Publicado: (2024)
Reasoning Inconsistencies and How to Mitigate Them in Deep Learning
por: Arakelyan, Erik
Publicado: (2025)
por: Arakelyan, Erik
Publicado: (2025)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
por: Jiang, Ruixiang, et al.
Publicado: (2025)
por: Jiang, Ruixiang, et al.
Publicado: (2025)
Mirror-Consistency: Harnessing Inconsistency in Majority Voting
por: Huang, Siyuan, et al.
Publicado: (2024)
por: Huang, Siyuan, et al.
Publicado: (2024)
Do LLMs Encode Functional Importance of Reasoning Tokens?
por: Singh, Janvijay, et al.
Publicado: (2026)
por: Singh, Janvijay, et al.
Publicado: (2026)
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
por: Mei, Zhiting, et al.
Publicado: (2025)
por: Mei, Zhiting, et al.
Publicado: (2025)
MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
por: Wu, Juncheng, et al.
Publicado: (2025)
por: Wu, Juncheng, et al.
Publicado: (2025)
PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
por: Harary, Sapir, et al.
Publicado: (2025)
por: Harary, Sapir, et al.
Publicado: (2025)
Beyond Self-Consistency: Ensemble Reasoning Boosts Consistency and Accuracy of LLMs in Cancer Staging
por: Chang, Chia-Hsuan, et al.
Publicado: (2024)
por: Chang, Chia-Hsuan, et al.
Publicado: (2024)
Flow of Reasoning: Training LLMs for Divergent Reasoning with Minimal Examples
por: Yu, Fangxu, et al.
Publicado: (2024)
por: Yu, Fangxu, et al.
Publicado: (2024)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
por: Jin, Keyan, et al.
Publicado: (2025)
por: Jin, Keyan, et al.
Publicado: (2025)
Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives
por: Zhang, Wenqi, et al.
Publicado: (2024)
por: Zhang, Wenqi, et al.
Publicado: (2024)
Fast and Accurate Factual Inconsistency Detection Over Long Documents
por: Lattimer, Barrett Martin, et al.
Publicado: (2023)
por: Lattimer, Barrett Martin, et al.
Publicado: (2023)
Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
por: Kim, Jaemin, et al.
Publicado: (2025)
por: Kim, Jaemin, et al.
Publicado: (2025)
Dissociation of Faithful and Unfaithful Reasoning in LLMs
por: Yee, Evelyn, et al.
Publicado: (2024)
por: Yee, Evelyn, et al.
Publicado: (2024)
Are Your LLMs Capable of Stable Reasoning?
por: Liu, Junnan, et al.
Publicado: (2024)
por: Liu, Junnan, et al.
Publicado: (2024)
Ejemplares similares
-
Cross-Examiner: Evaluating Consistency of Large Language Model-Generated Explanations
por: Villa, Danielle, et al.
Publicado: (2025) -
Ranking Large Language Models without Ground Truth
por: Dhurandhar, Amit, et al.
Publicado: (2024) -
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
por: Achintalwar, Swapnaja, et al.
Publicado: (2024) -
AI Steerability 360: A Toolkit for Steering Large Language Models
por: Miehling, Erik, et al.
Publicado: (2026) -
Mitigating Misalignment Contagion by Steering with Implicit Traits
por: Chang, Maria, et al.
Publicado: (2026)