Saved in:
| Main Authors: | Srikanth, Neha, Rudinger, Rachel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2502.08080 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How often are errors in natural language reasoning due to paraphrastic variability?
by: Srikanth, Neha, et al.
Published: (2024)
by: Srikanth, Neha, et al.
Published: (2024)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
by: Srikanth, Neha, et al.
Published: (2026)
by: Srikanth, Neha, et al.
Published: (2026)
Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat Logs
by: Sarkar, Rupak, et al.
Published: (2025)
by: Sarkar, Rupak, et al.
Published: (2025)
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
by: Srikanth, Neha, et al.
Published: (2023)
by: Srikanth, Neha, et al.
Published: (2023)
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the U.S
by: Acquaye, Christabel, et al.
Published: (2024)
by: Acquaye, Christabel, et al.
Published: (2024)
On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models
by: Sancheti, Abhilasha, et al.
Published: (2024)
by: Sancheti, Abhilasha, et al.
Published: (2024)
Atomic Inference for NLI with Generated Facts as Atoms
by: Stacey, Joe, et al.
Published: (2023)
by: Stacey, Joe, et al.
Published: (2023)
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
by: Balepur, Nishant, et al.
Published: (2023)
by: Balepur, Nishant, et al.
Published: (2023)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
On the Mutual Influence of Gender and Occupation in LLM Representations
by: An, Haozhe, et al.
Published: (2025)
by: An, Haozhe, et al.
Published: (2025)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Language Models Predict Empathy Gaps Between Social In-groups and Out-groups
by: Hou, Yu, et al.
Published: (2025)
by: Hou, Yu, et al.
Published: (2025)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
by: Acquaye, Christabel, et al.
Published: (2026)
by: Acquaye, Christabel, et al.
Published: (2026)
Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race, Ethnicity, and Gender?
by: An, Haozhe, et al.
Published: (2024)
by: An, Haozhe, et al.
Published: (2024)
Multiple LLM Agents Debate for Equitable Cultural Alignment
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
Speaking the Right Language: The Impact of Expertise Alignment in User-AI Interactions
by: Palta, Shramay, et al.
Published: (2025)
by: Palta, Shramay, et al.
Published: (2025)
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
by: Palta, Shramay, et al.
Published: (2025)
by: Palta, Shramay, et al.
Published: (2025)
LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
by: Liu, Jiarui, et al.
Published: (2025)
by: Liu, Jiarui, et al.
Published: (2025)
SQLSpace: A Representation Space for Text-to-SQL to Discover and Mitigate Robustness Gaps
by: Srikanth, Neha, et al.
Published: (2025)
by: Srikanth, Neha, et al.
Published: (2025)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Don't Fight Hallucinations, Use Them: Estimating Image Realism using NLI over Atomic Facts
by: Rykov, Elisei, et al.
Published: (2025)
by: Rykov, Elisei, et al.
Published: (2025)
Entailed Between the Lines: Incorporating Implication into NLI
by: Havaldar, Shreya, et al.
Published: (2025)
by: Havaldar, Shreya, et al.
Published: (2025)
Rethinking STS and NLI in Large Language Models
by: Wang, Yuxia, et al.
Published: (2023)
by: Wang, Yuxia, et al.
Published: (2023)
Exploring Continual Learning of Compositional Generalization in NLI
by: Fu, Xiyan, et al.
Published: (2024)
by: Fu, Xiyan, et al.
Published: (2024)
'Rich Dad, Poor Lad': How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
From Disagreement to Understanding: The Case for Ambiguity Detection in NLI
by: Jayaweera, Chathuri, et al.
Published: (2025)
by: Jayaweera, Chathuri, et al.
Published: (2025)
For Generated Text, Is NLI-Neutral Text the Best Text?
by: Mersinias, Michail, et al.
Published: (2023)
by: Mersinias, Michail, et al.
Published: (2023)
NL-Eye: Abductive NLI for Images
by: Ventura, Mor, et al.
Published: (2024)
by: Ventura, Mor, et al.
Published: (2024)
FRIDA to the Rescue! Analyzing Synthetic Data Effectiveness in Object-Based Common Sense Reasoning for Disaster Response
by: Shichman, Mollie, et al.
Published: (2025)
by: Shichman, Mollie, et al.
Published: (2025)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
by: Palta, Shramay, et al.
Published: (2024)
by: Palta, Shramay, et al.
Published: (2024)
Testing the Deliteralization Hypothesis in Human and Machine Translation
by: Marmonier, Malik, et al.
Published: (2026)
by: Marmonier, Malik, et al.
Published: (2026)
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Natural Language Inference Improves Compositionality in Vision-Language Models
by: Cascante-Bonilla, Paola, et al.
Published: (2024)
by: Cascante-Bonilla, Paola, et al.
Published: (2024)
When Does Meaning Backfire? Investigating the Role of AMRs in NLI
by: Min, Junghyun, et al.
Published: (2025)
by: Min, Junghyun, et al.
Published: (2025)
SocialNLI: A Dialogue-Centric Social Inference Dataset
by: Deo, Akhil, et al.
Published: (2025)
by: Deo, Akhil, et al.
Published: (2025)
A synthetic data approach for domain generalization of NLI models
by: Hosseini, Mohammad Javad, et al.
Published: (2024)
by: Hosseini, Mohammad Javad, et al.
Published: (2024)
Exploring Factual Entailment with NLI: A News Media Study
by: Mor-Lan, Guy, et al.
Published: (2024)
by: Mor-Lan, Guy, et al.
Published: (2024)
Similar Items
-
How often are errors in natural language reasoning due to paraphrastic variability?
by: Srikanth, Neha, et al.
Published: (2024) -
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
by: Srikanth, Neha, et al.
Published: (2026) -
Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat Logs
by: Sarkar, Rupak, et al.
Published: (2025) -
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
by: Srikanth, Neha, et al.
Published: (2023) -
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
by: Balepur, Nishant, et al.
Published: (2024)