Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Stacey, Joe, Alazraki, Lisa, Ubhi, Aran, Ermis, Beyza, Mueller, Aaron, Rei, Marek |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Atomic Inference for NLI with Generated Facts as Atoms
von: Stacey, Joe, et al.
Veröffentlicht: (2023)
von: Stacey, Joe, et al.
Veröffentlicht: (2023)
Distilling Robustness into Natural Language Inference Models with Domain-Targeted Augmentation
von: Stacey, Joe, et al.
Veröffentlicht: (2023)
von: Stacey, Joe, et al.
Veröffentlicht: (2023)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
Entangled Relations: Leveraging NLI and Meta-analysis to Enhance Biomedical Relation Extraction
von: Hogan, William, et al.
Veröffentlicht: (2024)
von: Hogan, William, et al.
Veröffentlicht: (2024)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
von: Stacey, Joe, et al.
Veröffentlicht: (2026)
von: Stacey, Joe, et al.
Veröffentlicht: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
VERITAS-NLI : Validation and Extraction of Reliable Information Through Automated Scraping and Natural Language Inference
von: Shah, Arjun, et al.
Veröffentlicht: (2024)
von: Shah, Arjun, et al.
Veröffentlicht: (2024)
LUCID: LLM-Generated Utterances for Complex and Interesting Dialogues
von: Stacey, Joe, et al.
Veröffentlicht: (2024)
von: Stacey, Joe, et al.
Veröffentlicht: (2024)
Arithmetic OOD Failure Unfolds in Stages in Minimal GPTs
von: Shintani, Seine A.
Veröffentlicht: (2026)
von: Shintani, Seine A.
Veröffentlicht: (2026)
Improving LLMs with a knowledge from databases
von: Máša, Petr
Veröffentlicht: (2025)
von: Máša, Petr
Veröffentlicht: (2025)
Improving QA Model Performance with Cartographic Inoculation
von: Chen, Allen, et al.
Veröffentlicht: (2024)
von: Chen, Allen, et al.
Veröffentlicht: (2024)
Pareto-Optimized Open-Source LLMs for Healthcare via Context Retrieval
von: Bayarri-Planas, Jordi, et al.
Veröffentlicht: (2024)
von: Bayarri-Planas, Jordi, et al.
Veröffentlicht: (2024)
Budget-Xfer: Budget-Constrained Source Language Selection for Cross-Lingual Transfer to African Languages
von: Idris, Tewodros Kederalah, et al.
Veröffentlicht: (2026)
von: Idris, Tewodros Kederalah, et al.
Veröffentlicht: (2026)
Continuous Predictive Modeling of Clinical Notes and ICD Codes in Patient Health Records
von: Caralt, Mireia Hernandez, et al.
Veröffentlicht: (2024)
von: Caralt, Mireia Hernandez, et al.
Veröffentlicht: (2024)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2024)
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2024)
SEED: Enhancing Text-to-SQL Performance and Practical Usability Through Automatic Evidence Generation
von: Yun, Janghyeon, et al.
Veröffentlicht: (2025)
von: Yun, Janghyeon, et al.
Veröffentlicht: (2025)
LLMs and Memorization: On Quality and Specificity of Copyright Compliance
von: Mueller, Felix B, et al.
Veröffentlicht: (2024)
von: Mueller, Felix B, et al.
Veröffentlicht: (2024)
The Open Source Advantage in Large Language Models (LLMs)
von: Manchanda, Jiya, et al.
Veröffentlicht: (2024)
von: Manchanda, Jiya, et al.
Veröffentlicht: (2024)
From Extraction to Synthesis: Entangled Heuristics for Agent-Augmented Strategic Reasoning
von: Ghisellini, Renato, et al.
Veröffentlicht: (2025)
von: Ghisellini, Renato, et al.
Veröffentlicht: (2025)
Parsing Akkadian Verbs with Prolog
von: Macks, Aaron
Veröffentlicht: (2024)
von: Macks, Aaron
Veröffentlicht: (2024)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
When Self-Reference Fails to Close: Matrix-Level Dynamics in Large Language Models
von: Bae, Ji Ho
Veröffentlicht: (2026)
von: Bae, Ji Ho
Veröffentlicht: (2026)
LLMs as Architects and Critics for Multi-Source Opinion Summarization
von: Attri, Anuj, et al.
Veröffentlicht: (2025)
von: Attri, Anuj, et al.
Veröffentlicht: (2025)
Evaluating Relational Reasoning in LLMs with REL
von: Fesser, Lukas, et al.
Veröffentlicht: (2026)
von: Fesser, Lukas, et al.
Veröffentlicht: (2026)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
Whose Facts Win? LLM Source Preferences under Knowledge Conflicts
von: Schuster, Jakob, et al.
Veröffentlicht: (2026)
von: Schuster, Jakob, et al.
Veröffentlicht: (2026)
Evaluating LLM Metrics Through Real-World Capabilities
von: Miller, Justin K, et al.
Veröffentlicht: (2025)
von: Miller, Justin K, et al.
Veröffentlicht: (2025)
LLMs and the Human Condition
von: Wallis, Peter
Veröffentlicht: (2024)
von: Wallis, Peter
Veröffentlicht: (2024)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
von: Ivanov, Igor
Veröffentlicht: (2025)
von: Ivanov, Igor
Veröffentlicht: (2025)
LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
von: Cui, Hyang
Veröffentlicht: (2025)
von: Cui, Hyang
Veröffentlicht: (2025)
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
von: Sun, Xiangkun, et al.
Veröffentlicht: (2026)
von: Sun, Xiangkun, et al.
Veröffentlicht: (2026)
RAG-Optimized Tibetan Tourism LLMs: Enhancing Accuracy and Personalization
von: Qi, Jinhu, et al.
Veröffentlicht: (2024)
von: Qi, Jinhu, et al.
Veröffentlicht: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2025)
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2025)
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
von: Geuter, Jonathan, et al.
Veröffentlicht: (2025)
von: Geuter, Jonathan, et al.
Veröffentlicht: (2025)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2026)
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2026)
MIRIAD: Augmenting LLMs with millions of medical query-response pairs
von: Zheng, Qinyue, et al.
Veröffentlicht: (2025)
von: Zheng, Qinyue, et al.
Veröffentlicht: (2025)
Extracting Structured Insights from Financial News: An Augmented LLM Driven Approach
von: Dolphin, Rian, et al.
Veröffentlicht: (2024)
von: Dolphin, Rian, et al.
Veröffentlicht: (2024)
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks
von: Tahir, Munief Hassan, et al.
Veröffentlicht: (2024)
von: Tahir, Munief Hassan, et al.
Veröffentlicht: (2024)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Atomic Inference for NLI with Generated Facts as Atoms
von: Stacey, Joe, et al.
Veröffentlicht: (2023) -
Distilling Robustness into Natural Language Inference Models with Domain-Targeted Augmentation
von: Stacey, Joe, et al.
Veröffentlicht: (2023) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025) -
Entangled Relations: Leveraging NLI and Meta-analysis to Enhance Biomedical Relation Extraction
von: Hogan, William, et al.
Veröffentlicht: (2024) -
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
von: Stacey, Joe, et al.
Veröffentlicht: (2026)