NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Srikanth, Neha, Rudinger, Rachel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How often are errors in natural language reasoning due to paraphrastic variability?
von: Srikanth, Neha, et al.
Veröffentlicht: (2024)
von: Srikanth, Neha, et al.
Veröffentlicht: (2024)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
von: Srikanth, Neha, et al.
Veröffentlicht: (2026)
von: Srikanth, Neha, et al.
Veröffentlicht: (2026)
Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat Logs
von: Sarkar, Rupak, et al.
Veröffentlicht: (2025)
von: Sarkar, Rupak, et al.
Veröffentlicht: (2025)
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
von: Srikanth, Neha, et al.
Veröffentlicht: (2023)
von: Srikanth, Neha, et al.
Veröffentlicht: (2023)
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
Atomic Inference for NLI with Generated Facts as Atoms
von: Stacey, Joe, et al.
Veröffentlicht: (2023)
von: Stacey, Joe, et al.
Veröffentlicht: (2023)
Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the U.S
von: Acquaye, Christabel, et al.
Veröffentlicht: (2024)
von: Acquaye, Christabel, et al.
Veröffentlicht: (2024)
On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models
von: Sancheti, Abhilasha, et al.
Veröffentlicht: (2024)
von: Sancheti, Abhilasha, et al.
Veröffentlicht: (2024)
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
von: Balepur, Nishant, et al.
Veröffentlicht: (2023)
von: Balepur, Nishant, et al.
Veröffentlicht: (2023)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
On the Mutual Influence of Gender and Occupation in LLM Representations
von: An, Haozhe, et al.
Veröffentlicht: (2025)
von: An, Haozhe, et al.
Veröffentlicht: (2025)
LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
von: Liu, Jiarui, et al.
Veröffentlicht: (2025)
von: Liu, Jiarui, et al.
Veröffentlicht: (2025)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
Language Models Predict Empathy Gaps Between Social In-groups and Out-groups
von: Hou, Yu, et al.
Veröffentlicht: (2025)
von: Hou, Yu, et al.
Veröffentlicht: (2025)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
von: Acquaye, Christabel, et al.
Veröffentlicht: (2026)
von: Acquaye, Christabel, et al.
Veröffentlicht: (2026)
Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race, Ethnicity, and Gender?
von: An, Haozhe, et al.
Veröffentlicht: (2024)
von: An, Haozhe, et al.
Veröffentlicht: (2024)
Multiple LLM Agents Debate for Equitable Cultural Alignment
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
Speaking the Right Language: The Impact of Expertise Alignment in User-AI Interactions
von: Palta, Shramay, et al.
Veröffentlicht: (2025)
von: Palta, Shramay, et al.
Veröffentlicht: (2025)
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
von: Palta, Shramay, et al.
Veröffentlicht: (2025)
von: Palta, Shramay, et al.
Veröffentlicht: (2025)
SQLSpace: A Representation Space for Text-to-SQL to Discover and Mitigate Robustness Gaps
von: Srikanth, Neha, et al.
Veröffentlicht: (2025)
von: Srikanth, Neha, et al.
Veröffentlicht: (2025)
Don't Fight Hallucinations, Use Them: Estimating Image Realism using NLI over Atomic Facts
von: Rykov, Elisei, et al.
Veröffentlicht: (2025)
von: Rykov, Elisei, et al.
Veröffentlicht: (2025)
Entailed Between the Lines: Incorporating Implication into NLI
von: Havaldar, Shreya, et al.
Veröffentlicht: (2025)
von: Havaldar, Shreya, et al.
Veröffentlicht: (2025)
Rethinking STS and NLI in Large Language Models
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
Exploring Continual Learning of Compositional Generalization in NLI
von: Fu, Xiyan, et al.
Veröffentlicht: (2024)
von: Fu, Xiyan, et al.
Veröffentlicht: (2024)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
From Disagreement to Understanding: The Case for Ambiguity Detection in NLI
von: Jayaweera, Chathuri, et al.
Veröffentlicht: (2025)
von: Jayaweera, Chathuri, et al.
Veröffentlicht: (2025)
For Generated Text, Is NLI-Neutral Text the Best Text?
von: Mersinias, Michail, et al.
Veröffentlicht: (2023)
von: Mersinias, Michail, et al.
Veröffentlicht: (2023)
Testing the Deliteralization Hypothesis in Human and Machine Translation
von: Marmonier, Malik, et al.
Veröffentlicht: (2026)
von: Marmonier, Malik, et al.
Veröffentlicht: (2026)
When Does Meaning Backfire? Investigating the Role of AMRs in NLI
von: Min, Junghyun, et al.
Veröffentlicht: (2025)
von: Min, Junghyun, et al.
Veröffentlicht: (2025)
SocialNLI: A Dialogue-Centric Social Inference Dataset
von: Deo, Akhil, et al.
Veröffentlicht: (2025)
von: Deo, Akhil, et al.
Veröffentlicht: (2025)
A synthetic data approach for domain generalization of NLI models
von: Hosseini, Mohammad Javad, et al.
Veröffentlicht: (2024)
von: Hosseini, Mohammad Javad, et al.
Veröffentlicht: (2024)
Exploring Factual Entailment with NLI: A News Media Study
von: Mor-Lan, Guy, et al.
Veröffentlicht: (2024)
von: Mor-Lan, Guy, et al.
Veröffentlicht: (2024)
EconNLI: Evaluating Large Language Models on Economics Reasoning
von: Guo, Yue, et al.
Veröffentlicht: (2024)
von: Guo, Yue, et al.
Veröffentlicht: (2024)
NL-Eye: Abductive NLI for Images
von: Ventura, Mor, et al.
Veröffentlicht: (2024)
von: Ventura, Mor, et al.
Veröffentlicht: (2024)
Atomic Fact Decomposition Helps Attributed Question Answering
von: Yan, Zhichao, et al.
Veröffentlicht: (2024)
von: Yan, Zhichao, et al.
Veröffentlicht: (2024)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
'Rich Dad, Poor Lad': How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
Lost in Variation? Evaluating NLI Performance in Basque and Spanish Geographical Variants
von: Bengoetxea, Jaione, et al.
Veröffentlicht: (2025)
von: Bengoetxea, Jaione, et al.
Veröffentlicht: (2025)
EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI
von: Zuo, Longfei, et al.
Veröffentlicht: (2025)
von: Zuo, Longfei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How often are errors in natural language reasoning due to paraphrastic variability?
von: Srikanth, Neha, et al.
Veröffentlicht: (2024) -
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
von: Srikanth, Neha, et al.
Veröffentlicht: (2026) -
Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat Logs
von: Sarkar, Rupak, et al.
Veröffentlicht: (2025) -
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
von: Srikanth, Neha, et al.
Veröffentlicht: (2023) -
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)