Unpacking the Resilience of SNLI Contradiction Examples to Attacks
Fuente:
arXiv
Salvato in:
| Autori principali: | Verma, Chetan, Agarwal, Archit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Atomic-SNLI: Fine-Grained Natural Language Inference through Atomic Fact Decomposition
di: Huang, Minghui
Pubblicazione: (2026)
di: Huang, Minghui
Pubblicazione: (2026)
Securing Genomic Data Against Inference Attacks in Federated Learning Environments
di: Pathade, Chetan, et al.
Pubblicazione: (2025)
di: Pathade, Chetan, et al.
Pubblicazione: (2025)
Explanation Generation for Contradiction Reconciliation with LLMs
di: Chan, Jason, et al.
Pubblicazione: (2026)
di: Chan, Jason, et al.
Pubblicazione: (2026)
The Best Defense is Attack: Repairing Semantics in Textual Adversarial Examples
di: Yang, Heng, et al.
Pubblicazione: (2023)
di: Yang, Heng, et al.
Pubblicazione: (2023)
Attacking Misinformation Detection Using Adversarial Examples Generated by Language Models
di: Przybyła, Piotr, et al.
Pubblicazione: (2024)
di: Przybyła, Piotr, et al.
Pubblicazione: (2024)
Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack
di: Fazla, Arnisa, et al.
Pubblicazione: (2025)
di: Fazla, Arnisa, et al.
Pubblicazione: (2025)
destroR: Attacking Transfer Models with Obfuscous Examples to Discard Perplexity
di: Ahmed, Saadat Rafid, et al.
Pubblicazione: (2025)
di: Ahmed, Saadat Rafid, et al.
Pubblicazione: (2025)
A Straightforward Pipeline for Targeted Entailment and Contradiction Detection
di: Sulc, Antonin
Pubblicazione: (2025)
di: Sulc, Antonin
Pubblicazione: (2025)
ContraDoc: Understanding Self-Contradictions in Documents with Large Language Models
di: Li, Jierui, et al.
Pubblicazione: (2023)
di: Li, Jierui, et al.
Pubblicazione: (2023)
SparseCL: Sparse Contrastive Learning for Contradiction Retrieval
di: Xu, Haike, et al.
Pubblicazione: (2024)
di: Xu, Haike, et al.
Pubblicazione: (2024)
ContraSolver: Self-Alignment of Language Models by Resolving Internal Preference Contradictions
di: Zhang, Xu, et al.
Pubblicazione: (2024)
di: Zhang, Xu, et al.
Pubblicazione: (2024)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
di: Pathade, Chetan
Pubblicazione: (2025)
di: Pathade, Chetan
Pubblicazione: (2025)
HealthContradict: Evaluating Biomedical Knowledge Conflicts in Language Models
di: Zhang, Boya, et al.
Pubblicazione: (2025)
di: Zhang, Boya, et al.
Pubblicazione: (2025)
Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions
di: Hu, Zhe, et al.
Pubblicazione: (2024)
di: Hu, Zhe, et al.
Pubblicazione: (2024)
Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study
di: Surana, Mokshit, et al.
Pubblicazione: (2026)
di: Surana, Mokshit, et al.
Pubblicazione: (2026)
Large Language Models as an Indirect Reasoner: Contrapositive and Contradiction for Automated Reasoning
di: Zhang, Yanfang, et al.
Pubblicazione: (2024)
di: Zhang, Yanfang, et al.
Pubblicazione: (2024)
Unpacking Ambiguity: The Interaction of Polysemous Discourse Markers and Non-DM Signals
di: Wu, Jingni, et al.
Pubblicazione: (2025)
di: Wu, Jingni, et al.
Pubblicazione: (2025)
Homograph Attacks on Maghreb Sentiment Analyzers
di: Qachfar, Fatima Zahra, et al.
Pubblicazione: (2024)
di: Qachfar, Fatima Zahra, et al.
Pubblicazione: (2024)
When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews
di: Kumar, Sandeep, et al.
Pubblicazione: (2026)
di: Kumar, Sandeep, et al.
Pubblicazione: (2026)
LLM-based Extraction of Contradictions from Patents
di: Trapp, Stefan, et al.
Pubblicazione: (2024)
di: Trapp, Stefan, et al.
Pubblicazione: (2024)
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
di: Ivison, Hamish, et al.
Pubblicazione: (2024)
di: Ivison, Hamish, et al.
Pubblicazione: (2024)
Framing Social Movements on Social Media: Unpacking Diagnostic, Prognostic, and Motivational Strategies
di: Mendelsohn, Julia, et al.
Pubblicazione: (2024)
di: Mendelsohn, Julia, et al.
Pubblicazione: (2024)
Arithmetics-Based Decomposition of Numeral Words -- Arithmetic Conditions give the Unpacking Strategy
di: Maier, Isidor Konrad, et al.
Pubblicazione: (2023)
di: Maier, Isidor Konrad, et al.
Pubblicazione: (2023)
Reasoning Trajectories for Socratic Debugging of Student Code: From Misconceptions to Contradictions and Updated Beliefs
di: Al-Hossami, Erfan, et al.
Pubblicazione: (2025)
di: Al-Hossami, Erfan, et al.
Pubblicazione: (2025)
Contradiction Detection in RAG Systems: Evaluating LLMs as Context Validators for Improved Information Consistency
di: Gokul, Vignesh, et al.
Pubblicazione: (2025)
di: Gokul, Vignesh, et al.
Pubblicazione: (2025)
Not All Demonstration Examples are Equally Beneficial: Reweighting Demonstration Examples for In-Context Learning
di: Yang, Zhe, et al.
Pubblicazione: (2023)
di: Yang, Zhe, et al.
Pubblicazione: (2023)
Contradiction to Consensus: Dual Perspective, Multi Source Retrieval Based Claim Verification with Source Level Disagreement using LLM
di: Biswas, Md Badsha, et al.
Pubblicazione: (2026)
di: Biswas, Md Badsha, et al.
Pubblicazione: (2026)
Unpacking Robustness in Inflectional Languages: Adversarial Evaluation and Mechanistic Insights
di: Walkowiak, Paweł, et al.
Pubblicazione: (2025)
di: Walkowiak, Paweł, et al.
Pubblicazione: (2025)
Tailored Truths: Optimizing LLM Persuasion with Personalization and Fabricated Statistics
di: Timm, Jasper, et al.
Pubblicazione: (2025)
di: Timm, Jasper, et al.
Pubblicazione: (2025)
On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency
di: Wang, Yiming, et al.
Pubblicazione: (2026)
di: Wang, Yiming, et al.
Pubblicazione: (2026)
Residual Drift Dominates Contradiction in Multi-Turn Constraint Reasoning
di: Kawada, Sebastien
Pubblicazione: (2026)
di: Kawada, Sebastien
Pubblicazione: (2026)
Model Editing with Canonical Examples
di: Hewitt, John, et al.
Pubblicazione: (2024)
di: Hewitt, John, et al.
Pubblicazione: (2024)
Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison
di: Shen, Judy Hanwen, et al.
Pubblicazione: (2024)
di: Shen, Judy Hanwen, et al.
Pubblicazione: (2024)
IM-Unpack: Training and Inference with Arbitrarily Low Precision Integers
di: Zeng, Zhanpeng, et al.
Pubblicazione: (2024)
di: Zeng, Zhanpeng, et al.
Pubblicazione: (2024)
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
di: Petrova, Nora, et al.
Pubblicazione: (2026)
di: Petrova, Nora, et al.
Pubblicazione: (2026)
GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples
di: Zhang, Harry, et al.
Pubblicazione: (2025)
di: Zhang, Harry, et al.
Pubblicazione: (2025)
Not the Example, but the Process: How Self-Generated Examples Enhance LLM Reasoning
di: Gwak, Daehoon, et al.
Pubblicazione: (2026)
di: Gwak, Daehoon, et al.
Pubblicazione: (2026)
Identification of Entailment and Contradiction Relations between Natural Language Sentences: A Neurosymbolic Approach
di: Feng, Xuyao, et al.
Pubblicazione: (2024)
di: Feng, Xuyao, et al.
Pubblicazione: (2024)
Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond
di: Fan, Chongyu, et al.
Pubblicazione: (2025)
di: Fan, Chongyu, et al.
Pubblicazione: (2025)
In-Context Example Ordering Guided by Label Distributions
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Atomic-SNLI: Fine-Grained Natural Language Inference through Atomic Fact Decomposition
di: Huang, Minghui
Pubblicazione: (2026) -
Securing Genomic Data Against Inference Attacks in Federated Learning Environments
di: Pathade, Chetan, et al.
Pubblicazione: (2025) -
Explanation Generation for Contradiction Reconciliation with LLMs
di: Chan, Jason, et al.
Pubblicazione: (2026) -
The Best Defense is Attack: Repairing Semantics in Textual Adversarial Examples
di: Yang, Heng, et al.
Pubblicazione: (2023) -
Attacking Misinformation Detection Using Adversarial Examples Generated by Language Models
di: Przybyła, Piotr, et al.
Pubblicazione: (2024)