Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questions
Fuente:
arXiv
Guardado en:
| Autores principales: | Sachdeva, Rachneet, Hazra, Rima, Gurevych, Iryna |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Localizing and Mitigating Errors in Long-form Question Answering
por: Sachdeva, Rachneet, et al.
Publicado: (2024)
por: Sachdeva, Rachneet, et al.
Publicado: (2024)
CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration
por: Sachdeva, Rachneet, et al.
Publicado: (2023)
por: Sachdeva, Rachneet, et al.
Publicado: (2023)
Are Emergent Abilities in Large Language Models just In-Context Learning?
por: Lu, Sheng, et al.
Publicado: (2023)
por: Lu, Sheng, et al.
Publicado: (2023)
Dive into the Chasm: Probing the Gap between In- and Cross-Topic Generalization
por: Waldis, Andreas, et al.
Publicado: (2024)
por: Waldis, Andreas, et al.
Publicado: (2024)
MAGneT: Coordinated Multi-Agent Generation of Synthetic Multi-Turn Mental Health Counseling Sessions
por: Mandal, Aishik, et al.
Publicado: (2025)
por: Mandal, Aishik, et al.
Publicado: (2025)
DARA: Decomposition-Alignment-Reasoning Autonomous Language Agent for Question Answering over Knowledge Graphs
por: Fang, Haishuo, et al.
Publicado: (2024)
por: Fang, Haishuo, et al.
Publicado: (2024)
Aligned Probing: Relating Toxic Behavior and Model Internals
por: Waldis, Andreas, et al.
Publicado: (2025)
por: Waldis, Andreas, et al.
Publicado: (2025)
PeerQA: A Scientific Question Answering Dataset from Peer Reviews
por: Baumgärtner, Tim, et al.
Publicado: (2025)
por: Baumgärtner, Tim, et al.
Publicado: (2025)
Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
por: Dycke, Nils, et al.
Publicado: (2025)
por: Dycke, Nils, et al.
Publicado: (2025)
Citation Failure: Definition, Analysis and Efficient Mitigation
por: Buchmann, Jan, et al.
Publicado: (2025)
por: Buchmann, Jan, et al.
Publicado: (2025)
Like a Good Nearest Neighbor: Practical Content Moderation and Text Classification
por: Bates, Luke, et al.
Publicado: (2023)
por: Bates, Luke, et al.
Publicado: (2023)
Enhancing Depression Detection via Question-wise Modality Fusion
por: Mandal, Aishik, et al.
Publicado: (2025)
por: Mandal, Aishik, et al.
Publicado: (2025)
Reward Modeling for Scientific Writing Evaluation
por: Şahinuç, Furkan, et al.
Publicado: (2026)
por: Şahinuç, Furkan, et al.
Publicado: (2026)
Token Weighting for Long-Range Language Modeling
por: Helm, Falko, et al.
Publicado: (2025)
por: Helm, Falko, et al.
Publicado: (2025)
Attribute or Abstain: Large Language Models as Long Document Assistants
por: Buchmann, Jan, et al.
Publicado: (2024)
por: Buchmann, Jan, et al.
Publicado: (2024)
Robust Utility-Preserving Text Anonymization Based on Large Language Models
por: Yang, Tianyu, et al.
Publicado: (2024)
por: Yang, Tianyu, et al.
Publicado: (2024)
Re3: A Holistic Framework and Dataset for Modeling Collaborative Document Revision
por: Ruan, Qian, et al.
Publicado: (2024)
por: Ruan, Qian, et al.
Publicado: (2024)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
por: Paul, Indraneil, et al.
Publicado: (2024)
por: Paul, Indraneil, et al.
Publicado: (2024)
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
por: Baumgärtner, Tim, et al.
Publicado: (2026)
por: Baumgärtner, Tim, et al.
Publicado: (2026)
Are Large Language Models Good Classifiers? A Study on Edit Intent Classification in Scientific Document Revisions
por: Ruan, Qian, et al.
Publicado: (2024)
por: Ruan, Qian, et al.
Publicado: (2024)
Cultural Learning-Based Culture Adaptation of Language Models
por: Liu, Chen Cecilia, et al.
Publicado: (2025)
por: Liu, Chen Cecilia, et al.
Publicado: (2025)
M2QA: Multi-domain Multilingual Question Answering
por: Engländer, Leon, et al.
Publicado: (2024)
por: Engländer, Leon, et al.
Publicado: (2024)
Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities
por: Mandal, Aishik, et al.
Publicado: (2025)
por: Mandal, Aishik, et al.
Publicado: (2025)
NeoQA: Evidence-based Question Answering with Generated News Events
por: Glockner, Max, et al.
Publicado: (2025)
por: Glockner, Max, et al.
Publicado: (2025)
Commitment Checklist: Auditing Author Commitments in Peer Review
por: Chen, Chung-Chi, et al.
Publicado: (2026)
por: Chen, Chung-Chi, et al.
Publicado: (2026)
Identifying Aspects in Peer Reviews
por: Lu, Sheng, et al.
Publicado: (2025)
por: Lu, Sheng, et al.
Publicado: (2025)
COVE: COntext and VEracity prediction for out-of-context images
por: Tonglet, Jonathan, et al.
Publicado: (2025)
por: Tonglet, Jonathan, et al.
Publicado: (2025)
Expert Preference-based Evaluation of Automated Related Work Generation
por: Şahinuç, Furkan, et al.
Publicado: (2025)
por: Şahinuç, Furkan, et al.
Publicado: (2025)
M4FC: a Multimodal, Multilingual, Multicultural, Multitask Real-World Fact-Checking Dataset
por: Geng, Jiahui, et al.
Publicado: (2025)
por: Geng, Jiahui, et al.
Publicado: (2025)
How to Handle Different Types of Out-of-Distribution Scenarios in Computational Argumentation? A Comprehensive and Fine-Grained Field Study
por: Waldis, Andreas, et al.
Publicado: (2023)
por: Waldis, Andreas, et al.
Publicado: (2023)
Overview of PerpectiveArg2024: The First Shared Task on Perspective Argument Retrieval
por: Falk, Neele, et al.
Publicado: (2024)
por: Falk, Neele, et al.
Publicado: (2024)
LLM Roleplay: Simulating Human-Chatbot Interaction
por: Tamoyan, Hovhannes, et al.
Publicado: (2024)
por: Tamoyan, Hovhannes, et al.
Publicado: (2024)
How are Prompts Different in Terms of Sensitivity?
por: Lu, Sheng, et al.
Publicado: (2023)
por: Lu, Sheng, et al.
Publicado: (2023)
Patches of Nonlinearity: Instruction Vectors in Large Language Models
por: Bigoulaeva, Irina, et al.
Publicado: (2026)
por: Bigoulaeva, Irina, et al.
Publicado: (2026)
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
por: Tamoyan, Hovhannes, et al.
Publicado: (2025)
por: Tamoyan, Hovhannes, et al.
Publicado: (2025)
ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
por: Bates, Luke, et al.
Publicado: (2025)
por: Bates, Luke, et al.
Publicado: (2025)
Bridging the Multilingual Safety Divide: Efficient, Culturally-Aware Alignment for Global South Languages
por: Banerjee, Somnath, et al.
Publicado: (2026)
por: Banerjee, Somnath, et al.
Publicado: (2026)
The Inherent Limits of Pretrained LLMs: The Unexpected Convergence of Instruction Tuning and In-Context Learning Capabilities
por: Bigoulaeva, Irina, et al.
Publicado: (2025)
por: Bigoulaeva, Irina, et al.
Publicado: (2025)
"Image, Tell me your story!" Predicting the original meta-context of visual misinformation
por: Tonglet, Jonathan, et al.
Publicado: (2024)
por: Tonglet, Jonathan, et al.
Publicado: (2024)
Culturally Aware and Adapted NLP: A Taxonomy and a Survey of the State of the Art
por: Liu, Chen Cecilia, et al.
Publicado: (2024)
por: Liu, Chen Cecilia, et al.
Publicado: (2024)
Ejemplares similares
-
Localizing and Mitigating Errors in Long-form Question Answering
por: Sachdeva, Rachneet, et al.
Publicado: (2024) -
CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration
por: Sachdeva, Rachneet, et al.
Publicado: (2023) -
Are Emergent Abilities in Large Language Models just In-Context Learning?
por: Lu, Sheng, et al.
Publicado: (2023) -
Dive into the Chasm: Probing the Gap between In- and Cross-Topic Generalization
por: Waldis, Andreas, et al.
Publicado: (2024) -
MAGneT: Coordinated Multi-Agent Generation of Synthetic Multi-Turn Mental Health Counseling Sessions
por: Mandal, Aishik, et al.
Publicado: (2025)