Let's CONFER: A Dataset for Evaluating Natural Language Inference Models on CONditional InFERence and Presupposition
Fuente:
arXiv
Guardado en:
| Autores principales: | Azin, Tara, Dumitrescu, Daniel, Inkpen, Diana, Singh, Raj |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Do Language Models Know Theo Has a Wife? Investigating the Proviso Problem
por: Azin, Tara, et al.
Publicado: (2026)
por: Azin, Tara, et al.
Publicado: (2026)
Presupposition and Reasoning in Conditionals: A Theory-Based Study of Humans and LLMs
por: Azin, Tara, et al.
Publicado: (2026)
por: Azin, Tara, et al.
Publicado: (2026)
SoftHateBench: Evaluating Moderation Models Against Reasoning-Driven, Policy-Compliant Hostility
por: Su, Xuanyu, et al.
Publicado: (2026)
por: Su, Xuanyu, et al.
Publicado: (2026)
Evaluating Reasoning Models for Queries with Presuppositions
por: Sathyanathan, Rose, et al.
Publicado: (2026)
por: Sathyanathan, Rose, et al.
Publicado: (2026)
uOttawa at LegalLens-2024: Transformer-based Classification Experiments
por: Meghdadi, Nima, et al.
Publicado: (2024)
por: Meghdadi, Nima, et al.
Publicado: (2024)
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
por: Zhu, Wang Bill, et al.
Publicado: (2025)
por: Zhu, Wang Bill, et al.
Publicado: (2025)
Evaluating Large Language Models for Health-related Queries with Presuppositions
por: Kaur, Navreet, et al.
Publicado: (2023)
por: Kaur, Navreet, et al.
Publicado: (2023)
Persian Abstract Meaning Representation: Annotation Guidelines and Gold Standard Dataset
por: Takhshid, Reza, et al.
Publicado: (2022)
por: Takhshid, Reza, et al.
Publicado: (2022)
Safer in Translation? Presupposition Robustness in Indic Languages
por: Palnitkar, Aadi, et al.
Publicado: (2025)
por: Palnitkar, Aadi, et al.
Publicado: (2025)
Synthetic Dataset for Evaluating Complex Compositional Knowledge for Natural Language Inference
por: Akoju, Sushma Anand, et al.
Publicado: (2023)
por: Akoju, Sushma Anand, et al.
Publicado: (2023)
The Presupposition Problem in Representation Genesis
por: Wu, Yiling
Publicado: (2026)
por: Wu, Yiling
Publicado: (2026)
HateSieve: A Contrastive Learning Framework for Detecting and Segmenting Hateful Content in Multimodal Memes
por: Su, Xuanyu, et al.
Publicado: (2024)
por: Su, Xuanyu, et al.
Publicado: (2024)
A New Benchmark Dataset and Mixture-of-Experts Language Models for Adversarial Natural Language Inference in Vietnamese
por: Van Huynh, Tin, et al.
Publicado: (2024)
por: Van Huynh, Tin, et al.
Publicado: (2024)
BNLI: A Linguistically-Refined Bengali Dataset for Natural Language Inference
por: Haque, Farah Binta, et al.
Publicado: (2025)
por: Haque, Farah Binta, et al.
Publicado: (2025)
LTD-Bench: Evaluating Large Language Models by Letting Them Draw
por: Lin, Liuhao, et al.
Publicado: (2025)
por: Lin, Liuhao, et al.
Publicado: (2025)
LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High
por: Sieker, Judith, et al.
Publicado: (2025)
por: Sieker, Judith, et al.
Publicado: (2025)
Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models
por: Sinha, Samridhi Raj, et al.
Publicado: (2025)
por: Sinha, Samridhi Raj, et al.
Publicado: (2025)
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
por: Madaan, Lovish, et al.
Publicado: (2024)
por: Madaan, Lovish, et al.
Publicado: (2024)
Evaluating Gender Bias of Pre-trained Language Models in Natural Language Inference by Considering All Labels
por: Anantaprayoon, Panatchakorn, et al.
Publicado: (2023)
por: Anantaprayoon, Panatchakorn, et al.
Publicado: (2023)
Defining and Evaluating Decision and Composite Risk in Language Models Applied to Natural Language Inference
por: Shen, Ke, et al.
Publicado: (2024)
por: Shen, Ke, et al.
Publicado: (2024)
If We May De-Presuppose: Robustly Verifying Claims through Presupposition-Free Question Decomposition
por: Dipta, Shubhashis Roy, et al.
Publicado: (2025)
por: Dipta, Shubhashis Roy, et al.
Publicado: (2025)
Multi-head attention debiasing and contrastive learning for mitigating Dataset Artifacts in Natural Language Inference
por: Sivakoti, Karthik
Publicado: (2024)
por: Sivakoti, Karthik
Publicado: (2024)
NatLogAttack: A Framework for Attacking Natural Language Inference Models with Natural Logic
por: Zheng, Zi'ou, et al.
Publicado: (2023)
por: Zheng, Zi'ou, et al.
Publicado: (2023)
On Reference (In-)Determinacy in Natural Language Inference
por: Chen, Sihao, et al.
Publicado: (2025)
por: Chen, Sihao, et al.
Publicado: (2025)
Hypothesis-only Biases in Large Language Model-Elicited Natural Language Inference
por: Proebsting, Grace, et al.
Publicado: (2024)
por: Proebsting, Grace, et al.
Publicado: (2024)
Natural Language Inference Improves Compositionality in Vision-Language Models
por: Cascante-Bonilla, Paola, et al.
Publicado: (2024)
por: Cascante-Bonilla, Paola, et al.
Publicado: (2024)
Biases in Large Language Model-Elicited Text: A Case Study in Natural Language Inference
por: Proebsting, Grace, et al.
Publicado: (2025)
por: Proebsting, Grace, et al.
Publicado: (2025)
Fine-Grained Natural Language Inference Based Faithfulness Evaluation for Diverse Summarisation Tasks
por: Zhang, Huajian, et al.
Publicado: (2024)
por: Zhang, Huajian, et al.
Publicado: (2024)
A MISMATCHED Benchmark for Scientific Natural Language Inference
por: Shaik, Firoz, et al.
Publicado: (2025)
por: Shaik, Firoz, et al.
Publicado: (2025)
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
por: Li, Baiqi, et al.
Publicado: (2024)
por: Li, Baiqi, et al.
Publicado: (2024)
Let Me Teach You: Pedagogical Foundations of Feedback for Language Models
por: Borges, Beatriz, et al.
Publicado: (2023)
por: Borges, Beatriz, et al.
Publicado: (2023)
Can Large Language Models Robustly Perform Natural Language Inference for Japanese Comparatives?
por: Mikami, Yosuke, et al.
Publicado: (2025)
por: Mikami, Yosuke, et al.
Publicado: (2025)
Natural Language Processing for Dialects of a Language: A Survey
por: Joshi, Aditya, et al.
Publicado: (2024)
por: Joshi, Aditya, et al.
Publicado: (2024)
Towards Controllable Natural Language Inference through Lexical Inference Types
por: Zhang, Yingji, et al.
Publicado: (2023)
por: Zhang, Yingji, et al.
Publicado: (2023)
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
por: Schaeffer, Rylan, et al.
Publicado: (2025)
por: Schaeffer, Rylan, et al.
Publicado: (2025)
Evaluating Large Language Models for Radiology Natural Language Processing
por: Liu, Zhengliang, et al.
Publicado: (2023)
por: Liu, Zhengliang, et al.
Publicado: (2023)
A Chinese Dataset for Evaluating the Safeguards in Large Language Models
por: Wang, Yuxia, et al.
Publicado: (2024)
por: Wang, Yuxia, et al.
Publicado: (2024)
A Systematic Evaluation of Large Language Models for Natural Language Generation Tasks
por: Ni, Xuanfan, et al.
Publicado: (2024)
por: Ni, Xuanfan, et al.
Publicado: (2024)
Unraveling the Dominance of Large Language Models Over Transformer Models for Bangla Natural Language Inference: A Comprehensive Study
por: Faria, Fatema Tuj Johora, et al.
Publicado: (2024)
por: Faria, Fatema Tuj Johora, et al.
Publicado: (2024)
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
por: Kaing, Hour, et al.
Publicado: (2025)
por: Kaing, Hour, et al.
Publicado: (2025)
Ejemplares similares
-
Do Language Models Know Theo Has a Wife? Investigating the Proviso Problem
por: Azin, Tara, et al.
Publicado: (2026) -
Presupposition and Reasoning in Conditionals: A Theory-Based Study of Humans and LLMs
por: Azin, Tara, et al.
Publicado: (2026) -
SoftHateBench: Evaluating Moderation Models Against Reasoning-Driven, Policy-Compliant Hostility
por: Su, Xuanyu, et al.
Publicado: (2026) -
Evaluating Reasoning Models for Queries with Presuppositions
por: Sathyanathan, Rose, et al.
Publicado: (2026) -
uOttawa at LegalLens-2024: Transformer-based Classification Experiments
por: Meghdadi, Nima, et al.
Publicado: (2024)