The Self-Contained Negation Test Set
Fuente:
arXiv
Saved in:
| Main Authors: | Kletz, David, Amsili, Pascal, Candito, Marie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probing structural constraints of negation in Pretrained Language Models
by: Kletz, David, et al.
Published: (2024)
by: Kletz, David, et al.
Published: (2024)
Auxiliary Tasks to Boost Biaffine Semantic Dependency Parsing
by: Candito, Marie
Published: (2024)
by: Candito, Marie
Published: (2024)
Injecting Wiktionary to improve token-level contextual representations using contrastive learning
by: Mosolova, Anna, et al.
Published: (2024)
by: Mosolova, Anna, et al.
Published: (2024)
In the LLM era, Word Sense Induction remains unsolved
by: Mosolova, Anna, et al.
Published: (2026)
by: Mosolova, Anna, et al.
Published: (2026)
How to Evaluate Coreference in Literary Texts?
by: Duron-Tejedor, Ana-Isabel, et al.
Published: (2023)
by: Duron-Tejedor, Ana-Isabel, et al.
Published: (2023)
Modeling the human lexicon under temperature variations: linguistic factors, diversity and typicality in LLM word associations
by: Rodriguez, Maria Andueza, et al.
Published: (2026)
by: Rodriguez, Maria Andueza, et al.
Published: (2026)
Are the LLMs Capable of Maintaining at Least the Language Genus?
by: Mitrović, Sandra, et al.
Published: (2025)
by: Mitrović, Sandra, et al.
Published: (2025)
Automatic Prompt Optimization for Dataset-Level Feature Discovery
by: Cosma, Adrian, et al.
Published: (2026)
by: Cosma, Adrian, et al.
Published: (2026)
Negated String Containment is Decidable (Technical Report)
by: Havlena, Vojtěch, et al.
Published: (2025)
by: Havlena, Vojtěch, et al.
Published: (2025)
SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative Examples
by: Fu, Deqing, et al.
Published: (2023)
by: Fu, Deqing, et al.
Published: (2023)
UltraWiki: Ultra-fine-grained Entity Set Expansion with Negative Seed Entities
by: Li, Yangning, et al.
Published: (2024)
by: Li, Yangning, et al.
Published: (2024)
Carrot and Stick: Inducing Self-Motivation with Positive & Negative Feedback
by: Sohn, Jimin, et al.
Published: (2024)
by: Sohn, Jimin, et al.
Published: (2024)
Test Set Quality in Multilingual LLM Evaluation
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
The Atomic Instruction Gap: Instruction-Tuned LLMs Struggle with Simple, Self-Contained Directives
by: Lim, Henry, et al.
Published: (2025)
by: Lim, Henry, et al.
Published: (2025)
How Hard is this Test Set? NLI Characterization by Exploiting Training Dynamics
by: Cosma, Adrian, et al.
Published: (2024)
by: Cosma, Adrian, et al.
Published: (2024)
Self-Reflective Generation at Test Time
by: Mu, Jian, et al.
Published: (2025)
by: Mu, Jian, et al.
Published: (2025)
Mitigating Negative Interference in Multilingual Sequential Knowledge Editing through Null-Space Constraints
by: Sun, Wei, et al.
Published: (2025)
by: Sun, Wei, et al.
Published: (2025)
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
by: Hou, Wenjun, et al.
Published: (2024)
by: Hou, Wenjun, et al.
Published: (2024)
Towards Self-Contained Answers: Entity-Based Answer Rewriting in Conversational Search
by: Sekulić, Ivan, et al.
Published: (2024)
by: Sekulić, Ivan, et al.
Published: (2024)
DyePack: Provably Flagging Test Set Contamination in LLMs Using Backdoors
by: Cheng, Yize, et al.
Published: (2025)
by: Cheng, Yize, et al.
Published: (2025)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
by: Schaeffer, Rylan, et al.
Published: (2026)
by: Schaeffer, Rylan, et al.
Published: (2026)
Commonsense Knowledge with Negation: A Resource to Enhance Negation Understanding
by: Wang, Zijie, et al.
Published: (2026)
by: Wang, Zijie, et al.
Published: (2026)
Pantagruel: Unified Self-Supervised Encoders for French Text and Speech
by: Le, Phuong-Hang, et al.
Published: (2026)
by: Le, Phuong-Hang, et al.
Published: (2026)
An Exploration of Self-Supervised Mutual Information Alignment for Multi-Task Settings
by: Govande, Soham V.
Published: (2024)
by: Govande, Soham V.
Published: (2024)
BenCSSmark: Making the Social Sciences Count in LLM Research
by: Chatelain, Arnault, et al.
Published: (2026)
by: Chatelain, Arnault, et al.
Published: (2026)
Confidence-Weighted Token Set Cover for Early Hypothesis Pruning in Self-Consistency
by: Sultan, Md Arafat, et al.
Published: (2025)
by: Sultan, Md Arafat, et al.
Published: (2025)
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients
by: Xu, Mingwei, et al.
Published: (2026)
by: Xu, Mingwei, et al.
Published: (2026)
StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario
by: Boito, Marcely Zanon, et al.
Published: (2026)
by: Boito, Marcely Zanon, et al.
Published: (2026)
When to Commit? Towards Variable-Size Self-Contained Blocks for Discrete Diffusion Language Models
by: Wang, Danny, et al.
Published: (2026)
by: Wang, Danny, et al.
Published: (2026)
Hypothesis Testing for Quantifying LLM-Human Misalignment in Multiple Choice Settings
by: Hong, Harbin, et al.
Published: (2025)
by: Hong, Harbin, et al.
Published: (2025)
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
by: Duan, Shitong, et al.
Published: (2024)
by: Duan, Shitong, et al.
Published: (2024)
Neural Networks Against (and For) Self-Training: Classification with Small Labeled and Large Unlabeled Sets
by: Karisani, Payam
Published: (2023)
by: Karisani, Payam
Published: (2023)
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
by: Porada, Ian, et al.
Published: (2024)
by: Porada, Ian, et al.
Published: (2024)
HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities
by: Dönmez, Esra, et al.
Published: (2026)
by: Dönmez, Esra, et al.
Published: (2026)
Negation-Induced Forgetting in LLMs
by: Capuano, Francesca, et al.
Published: (2025)
by: Capuano, Francesca, et al.
Published: (2025)
NegativePrompt: Leveraging Psychology for Large Language Models Enhancement via Negative Emotional Stimuli
by: Wang, Xu, et al.
Published: (2024)
by: Wang, Xu, et al.
Published: (2024)
Test-time Recursive Thinking: Self-Improvement without External Feedback
by: Zhuang, Yufan, et al.
Published: (2026)
by: Zhuang, Yufan, et al.
Published: (2026)
ISSR: Iterative Selection with Self-Review for Vocabulary Test Distractor Generation
by: Liu, Yu-Cheng, et al.
Published: (2025)
by: Liu, Yu-Cheng, et al.
Published: (2025)
LatentEvolve: Self-Evolving Test-Time Scaling in Latent Space
by: Zhang, Guibin, et al.
Published: (2025)
by: Zhang, Guibin, et al.
Published: (2025)
Similar Items
-
Probing structural constraints of negation in Pretrained Language Models
by: Kletz, David, et al.
Published: (2024) -
Auxiliary Tasks to Boost Biaffine Semantic Dependency Parsing
by: Candito, Marie
Published: (2024) -
Injecting Wiktionary to improve token-level contextual representations using contrastive learning
by: Mosolova, Anna, et al.
Published: (2024) -
In the LLM era, Word Sense Induction remains unsolved
by: Mosolova, Anna, et al.
Published: (2026) -
How to Evaluate Coreference in Literary Texts?
by: Duron-Tejedor, Ana-Isabel, et al.
Published: (2023)