Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Stacey, Joe, Alazraki, Lisa, Ubhi, Aran, Ermis, Beyza, Mueller, Aaron, Rei, Marek |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Atomic Inference for NLI with Generated Facts as Atoms
by: Stacey, Joe, et al.
Published: (2023)
by: Stacey, Joe, et al.
Published: (2023)
Distilling Robustness into Natural Language Inference Models with Domain-Targeted Augmentation
by: Stacey, Joe, et al.
Published: (2023)
by: Stacey, Joe, et al.
Published: (2023)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
Entangled Relations: Leveraging NLI and Meta-analysis to Enhance Biomedical Relation Extraction
by: Hogan, William, et al.
Published: (2024)
by: Hogan, William, et al.
Published: (2024)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
by: Stacey, Joe, et al.
Published: (2026)
by: Stacey, Joe, et al.
Published: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
VERITAS-NLI : Validation and Extraction of Reliable Information Through Automated Scraping and Natural Language Inference
by: Shah, Arjun, et al.
Published: (2024)
by: Shah, Arjun, et al.
Published: (2024)
LUCID: LLM-Generated Utterances for Complex and Interesting Dialogues
by: Stacey, Joe, et al.
Published: (2024)
by: Stacey, Joe, et al.
Published: (2024)
Arithmetic OOD Failure Unfolds in Stages in Minimal GPTs
by: Shintani, Seine A.
Published: (2026)
by: Shintani, Seine A.
Published: (2026)
Improving LLMs with a knowledge from databases
by: Máša, Petr
Published: (2025)
by: Máša, Petr
Published: (2025)
Pareto-Optimized Open-Source LLMs for Healthcare via Context Retrieval
by: Bayarri-Planas, Jordi, et al.
Published: (2024)
by: Bayarri-Planas, Jordi, et al.
Published: (2024)
Improving QA Model Performance with Cartographic Inoculation
by: Chen, Allen, et al.
Published: (2024)
by: Chen, Allen, et al.
Published: (2024)
Budget-Xfer: Budget-Constrained Source Language Selection for Cross-Lingual Transfer to African Languages
by: Idris, Tewodros Kederalah, et al.
Published: (2026)
by: Idris, Tewodros Kederalah, et al.
Published: (2026)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
by: Nikankin, Yaniv, et al.
Published: (2024)
by: Nikankin, Yaniv, et al.
Published: (2024)
Continuous Predictive Modeling of Clinical Notes and ICD Codes in Patient Health Records
by: Caralt, Mireia Hernandez, et al.
Published: (2024)
by: Caralt, Mireia Hernandez, et al.
Published: (2024)
SEED: Enhancing Text-to-SQL Performance and Practical Usability Through Automatic Evidence Generation
by: Yun, Janghyeon, et al.
Published: (2025)
by: Yun, Janghyeon, et al.
Published: (2025)
LLMs and Memorization: On Quality and Specificity of Copyright Compliance
by: Mueller, Felix B, et al.
Published: (2024)
by: Mueller, Felix B, et al.
Published: (2024)
From Extraction to Synthesis: Entangled Heuristics for Agent-Augmented Strategic Reasoning
by: Ghisellini, Renato, et al.
Published: (2025)
by: Ghisellini, Renato, et al.
Published: (2025)
The Open Source Advantage in Large Language Models (LLMs)
by: Manchanda, Jiya, et al.
Published: (2024)
by: Manchanda, Jiya, et al.
Published: (2024)
Parsing Akkadian Verbs with Prolog
by: Macks, Aaron
Published: (2024)
by: Macks, Aaron
Published: (2024)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
When Self-Reference Fails to Close: Matrix-Level Dynamics in Large Language Models
by: Bae, Ji Ho
Published: (2026)
by: Bae, Ji Ho
Published: (2026)
LLMs as Architects and Critics for Multi-Source Opinion Summarization
by: Attri, Anuj, et al.
Published: (2025)
by: Attri, Anuj, et al.
Published: (2025)
Evaluating Relational Reasoning in LLMs with REL
by: Fesser, Lukas, et al.
Published: (2026)
by: Fesser, Lukas, et al.
Published: (2026)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
by: Simhi, Adi, et al.
Published: (2024)
by: Simhi, Adi, et al.
Published: (2024)
Whose Facts Win? LLM Source Preferences under Knowledge Conflicts
by: Schuster, Jakob, et al.
Published: (2026)
by: Schuster, Jakob, et al.
Published: (2026)
LLMs and the Human Condition
by: Wallis, Peter
Published: (2024)
by: Wallis, Peter
Published: (2024)
Evaluating LLM Metrics Through Real-World Capabilities
by: Miller, Justin K, et al.
Published: (2025)
by: Miller, Justin K, et al.
Published: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
by: Ivanov, Igor
Published: (2025)
by: Ivanov, Igor
Published: (2025)
LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
by: Cui, Hyang
Published: (2025)
by: Cui, Hyang
Published: (2025)
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
by: Sun, Xiangkun, et al.
Published: (2026)
by: Sun, Xiangkun, et al.
Published: (2026)
RAG-Optimized Tibetan Tourism LLMs: Enhancing Accuracy and Personalization
by: Qi, Jinhu, et al.
Published: (2024)
by: Qi, Jinhu, et al.
Published: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
by: Chehbouni, Khaoula, et al.
Published: (2025)
by: Chehbouni, Khaoula, et al.
Published: (2025)
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
by: Geuter, Jonathan, et al.
Published: (2025)
by: Geuter, Jonathan, et al.
Published: (2025)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
by: Bouchekif, Abdessalam, et al.
Published: (2026)
by: Bouchekif, Abdessalam, et al.
Published: (2026)
MIRIAD: Augmenting LLMs with millions of medical query-response pairs
by: Zheng, Qinyue, et al.
Published: (2025)
by: Zheng, Qinyue, et al.
Published: (2025)
Extracting Structured Insights from Financial News: An Augmented LLM Driven Approach
by: Dolphin, Rian, et al.
Published: (2024)
by: Dolphin, Rian, et al.
Published: (2024)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
by: Collado-Montañez, Jaime, et al.
Published: (2025)
by: Collado-Montañez, Jaime, et al.
Published: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
Similar Items
-
Atomic Inference for NLI with Generated Facts as Atoms
by: Stacey, Joe, et al.
Published: (2023) -
Distilling Robustness into Natural Language Inference Models with Domain-Targeted Augmentation
by: Stacey, Joe, et al.
Published: (2023) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025) -
Entangled Relations: Leveraging NLI and Meta-analysis to Enhance Biomedical Relation Extraction
by: Hogan, William, et al.
Published: (2024) -
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
by: Stacey, Joe, et al.
Published: (2026)