How Hard is this Test Set? NLI Characterization by Exploiting Training Dynamics
Fuente:
arXiv
Saved in:
| Main Authors: | Cosma, Adrian, Ruseti, Stefan, Dascalu, Mihai, Caragea, Cornelia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training Language Models with homotokens Leads to Delayed Overfitting
by: Cosma, Adrian, et al.
Published: (2026)
by: Cosma, Adrian, et al.
Published: (2026)
The Strawberry Problem: Emergence of Character-level Understanding in Tokenized Language Models
by: Cosma, Adrian, et al.
Published: (2025)
by: Cosma, Adrian, et al.
Published: (2025)
Neural Grammatical Error Correction for Romanian
by: Cotet, Teodor-Mihai, et al.
Published: (2026)
by: Cotet, Teodor-Mihai, et al.
Published: (2026)
Value-Aware Numerical Representations for Transformer Language Models
by: Dutulescu, Andreea, et al.
Published: (2026)
by: Dutulescu, Andreea, et al.
Published: (2026)
MSciNLI: A Diverse Benchmark for Scientific Natural Language Inference
by: Sadat, Mobashir, et al.
Published: (2024)
by: Sadat, Mobashir, et al.
Published: (2024)
A Novel Cartography-Based Curriculum Learning Method Applied on RoNLI: The First Romanian Natural Language Inference Corpus
by: Poesina, Eduard, et al.
Published: (2024)
by: Poesina, Eduard, et al.
Published: (2024)
Co-training for Low Resource Scientific Natural Language Inference
by: Sadat, Mobashir, et al.
Published: (2024)
by: Sadat, Mobashir, et al.
Published: (2024)
Stanceformer: Target-Aware Transformer for Stance Detection
by: Garg, Krishna, et al.
Published: (2024)
by: Garg, Krishna, et al.
Published: (2024)
BLooP: Zero-Shot Abstractive Summarization using Large Language Models with Bigram Lookahead Promotion
by: Iyer, Varun, et al.
Published: (2026)
by: Iyer, Varun, et al.
Published: (2026)
Evaluating Large Language Models for Stance Detection on Financial Targets from SEC Filing Reports and Earnings Call Transcripts
by: Gyawali, Nikesh, et al.
Published: (2025)
by: Gyawali, Nikesh, et al.
Published: (2025)
"Înţelegi Româneşte?'' A Recipe for Romanian Vision-Language Models
by: Masala, Mihai, et al.
Published: (2026)
by: Masala, Mihai, et al.
Published: (2026)
A MISMATCHED Benchmark for Scientific Natural Language Inference
by: Shaik, Firoz, et al.
Published: (2025)
by: Shaik, Firoz, et al.
Published: (2025)
Zero-Shot Verification-guided Chain of Thoughts
by: Chowdhury, Jishnu Ray, et al.
Published: (2025)
by: Chowdhury, Jishnu Ray, et al.
Published: (2025)
MADIAVE: Multi-Agent Debate for Implicit Attribute Value Extraction
by: Huang, Wei-Chieh, et al.
Published: (2025)
by: Huang, Wei-Chieh, et al.
Published: (2025)
The Shifting Landscape of Vaccine Discourse: Insights From a Decade of Pre- to Post-COVID-19 Vaccine Posts on Social Media
by: Gyawali, Nikesh, et al.
Published: (2025)
by: Gyawali, Nikesh, et al.
Published: (2025)
DynClean: Training Dynamics-based Label Cleaning for Distantly-Supervised Named Entity Recognition
by: Zhang, Qi, et al.
Published: (2025)
by: Zhang, Qi, et al.
Published: (2025)
Language Models (Mostly) Do Not Consider Emotion Triggers When Predicting Emotion
by: Singh, Smriti, et al.
Published: (2023)
by: Singh, Smriti, et al.
Published: (2023)
An In-Vitro Study on Cross-Lingual Generalization in Language Models
by: Cosma, Adrian
Published: (2026)
by: Cosma, Adrian
Published: (2026)
Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper
by: Garg, Krishna, et al.
Published: (2025)
by: Garg, Krishna, et al.
Published: (2025)
MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text Classification
by: Sirbu, Iustin, et al.
Published: (2025)
by: Sirbu, Iustin, et al.
Published: (2025)
LLM-guided Semi-Supervised Approaches for Social Media Crisis Data Classification
by: Ativo, Jacob, et al.
Published: (2026)
by: Ativo, Jacob, et al.
Published: (2026)
Zero-Shot Keyphrase Generation: Investigating Specialized Instructions and Multi-Sample Aggregation on Large Language Models
by: Mohan, Jayanth, et al.
Published: (2025)
by: Mohan, Jayanth, et al.
Published: (2025)
MorphNLI: A Stepwise Approach to Natural Language Inference Using Text Morphing
by: Negru, Vlad Andrei, et al.
Published: (2025)
by: Negru, Vlad Andrei, et al.
Published: (2025)
RoCode: A Dataset for Measuring Code Intelligence from Problem Definitions in Romanian
by: Cosma, Adrian, et al.
Published: (2024)
by: Cosma, Adrian, et al.
Published: (2024)
What Makes a Good Doctor Response? A Study on Text-Based Telemedicine
by: Cosma, Adrian, et al.
Published: (2026)
by: Cosma, Adrian, et al.
Published: (2026)
A Retrieval-Based Approach to Medical Procedure Matching in Romanian
by: Niculae, Andrei, et al.
Published: (2025)
by: Niculae, Andrei, et al.
Published: (2025)
First Train to Generate, then Generate to Train: UnitedSynT5 for Few-Shot NLI
by: Banerjee, Sourav, et al.
Published: (2024)
by: Banerjee, Sourav, et al.
Published: (2024)
SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific Documents
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
"Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions
by: Masala, Mihai, et al.
Published: (2024)
by: Masala, Mihai, et al.
Published: (2024)
Gumbel Machine: Counterfactual Student Writing Generation via Gumbel Noise Steering
by: McNichols, Hunter, et al.
Published: (2026)
by: McNichols, Hunter, et al.
Published: (2026)
Investigating Recurrent Transformers with Dynamic Halt
by: Chowdhury, Jishnu Ray, et al.
Published: (2024)
by: Chowdhury, Jishnu Ray, et al.
Published: (2024)
Exploring Continual Learning of Compositional Generalization in NLI
by: Fu, Xiyan, et al.
Published: (2024)
by: Fu, Xiyan, et al.
Published: (2024)
Rethinking STS and NLI in Large Language Models
by: Wang, Yuxia, et al.
Published: (2023)
by: Wang, Yuxia, et al.
Published: (2023)
Entailed Between the Lines: Incorporating Implication into NLI
by: Havaldar, Shreya, et al.
Published: (2025)
by: Havaldar, Shreya, et al.
Published: (2025)
OpenLLM-Ro -- Technical Report on Open-source Romanian LLMs
by: Masala, Mihai, et al.
Published: (2024)
by: Masala, Mihai, et al.
Published: (2024)
Dr.Copilot: A Multi-Agent Prompt Optimized Assistant for Improving Patient-Doctor Communication in Romanian
by: Niculae, Andrei, et al.
Published: (2025)
by: Niculae, Andrei, et al.
Published: (2025)
RoMath: A Mathematical Reasoning Benchmark in Romanian
by: Cosma, Adrian, et al.
Published: (2024)
by: Cosma, Adrian, et al.
Published: (2024)
For Generated Text, Is NLI-Neutral Text the Best Text?
by: Mersinias, Michail, et al.
Published: (2023)
by: Mersinias, Michail, et al.
Published: (2023)
From Disagreement to Understanding: The Case for Ambiguity Detection in NLI
by: Jayaweera, Chathuri, et al.
Published: (2025)
by: Jayaweera, Chathuri, et al.
Published: (2025)
A synthetic data approach for domain generalization of NLI models
by: Hosseini, Mohammad Javad, et al.
Published: (2024)
by: Hosseini, Mohammad Javad, et al.
Published: (2024)
Similar Items
-
Training Language Models with homotokens Leads to Delayed Overfitting
by: Cosma, Adrian, et al.
Published: (2026) -
The Strawberry Problem: Emergence of Character-level Understanding in Tokenized Language Models
by: Cosma, Adrian, et al.
Published: (2025) -
Neural Grammatical Error Correction for Romanian
by: Cotet, Teodor-Mihai, et al.
Published: (2026) -
Value-Aware Numerical Representations for Transformer Language Models
by: Dutulescu, Andreea, et al.
Published: (2026) -
MSciNLI: A Diverse Benchmark for Scientific Natural Language Inference
by: Sadat, Mobashir, et al.
Published: (2024)