Contextual Breach: Assessing the Robustness of Transformer-based QA Models
Fuente:
arXiv
Saved in:
| Main Authors: | Saadat, Asir, Asad, Nahian Ibn |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Not to Answer: Evaluating Prompts on GPT Models for Effective Abstention in Unanswerable Math Word Problems
by: Saadat, Asir, et al.
Published: (2024)
by: Saadat, Asir, et al.
Published: (2024)
Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs
by: Le, Chenqian, et al.
Published: (2025)
by: Le, Chenqian, et al.
Published: (2025)
JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
by: Cao, Zhihan, et al.
Published: (2025)
by: Cao, Zhihan, et al.
Published: (2025)
HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA
by: Chen, Xinyue, et al.
Published: (2024)
by: Chen, Xinyue, et al.
Published: (2024)
UETQuintet at BioCreative IX -- MedHopQA: Enhancing Biomedical QA with Selective Multi-hop Reasoning and Contextual Retrieval
by: Nguyen, Quoc-An, et al.
Published: (2026)
by: Nguyen, Quoc-An, et al.
Published: (2026)
Assessing The Potential Of Mid-Sized Language Models For Clinical QA
by: Bolton, Elliot, et al.
Published: (2024)
by: Bolton, Elliot, et al.
Published: (2024)
A Transformer and Prototype-based Interpretable Model for Contextual Sarcasm Detection
by: Wen, Ximing, et al.
Published: (2025)
by: Wen, Ximing, et al.
Published: (2025)
MasonTigers at SemEval-2024 Task 8: Performance Analysis of Transformer-based Models on Machine-Generated Text Detection
by: Puspo, Sadiya Sayara Chowdhury, et al.
Published: (2024)
by: Puspo, Sadiya Sayara Chowdhury, et al.
Published: (2024)
From Model to Breach: Towards Actionable LLM-Generated Vulnerabilities Reporting
by: Vallez, Cyril, et al.
Published: (2025)
by: Vallez, Cyril, et al.
Published: (2025)
DiversityMedQA: Assessing Demographic Biases in Medical Diagnosis using Large Language Models
by: Rawat, Rajat, et al.
Published: (2024)
by: Rawat, Rajat, et al.
Published: (2024)
Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation
by: Adib, Shefayat E Shams, et al.
Published: (2026)
by: Adib, Shefayat E Shams, et al.
Published: (2026)
DisasterQA: A Benchmark for Assessing the performance of LLMs in Disaster Response
by: Rawat, Rajat
Published: (2024)
by: Rawat, Rajat
Published: (2024)
Understanding QA generation: Extracting Parametric and Contextual Knowledge with CQA for Low Resource Bangla Language
by: Azmary, Umme Abira, et al.
Published: (2026)
by: Azmary, Umme Abira, et al.
Published: (2026)
A Morphologically-Aware Dictionary-based Data Augmentation Technique for Machine Translation of Under-Represented Languages
by: Alam, Md Mahfuz Ibn, et al.
Published: (2024)
by: Alam, Md Mahfuz Ibn, et al.
Published: (2024)
Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models
by: Dai, Runpeng, et al.
Published: (2025)
by: Dai, Runpeng, et al.
Published: (2025)
Clinical Reading Comprehension with Encoder-Decoder Models Enhanced by Direct Preference Optimization
by: Nahian, Md Sultan Al, et al.
Published: (2024)
by: Nahian, Md Sultan Al, et al.
Published: (2024)
Hidden Data Privacy Breaches in Federated Learning
by: Gong, Xueluan, et al.
Published: (2024)
by: Gong, Xueluan, et al.
Published: (2024)
TCE at Qur'an QA 2023 Shared Task: Low Resource Enhanced Transformer-based Ensemble Approach for Qur'anic QA
by: Elkomy, Mohammed Alaa, et al.
Published: (2024)
by: Elkomy, Mohammed Alaa, et al.
Published: (2024)
COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
by: Govil, Priyanshul, et al.
Published: (2024)
by: Govil, Priyanshul, et al.
Published: (2024)
Harnessing Collective Intelligence of LLMs for Robust Biomedical QA: A Multi-Model Approach
by: Panou, Dimitra, et al.
Published: (2025)
by: Panou, Dimitra, et al.
Published: (2025)
Certainty robustness: Evaluating LLM stability under self-challenging prompts
by: Saadat, Mohammadreza, et al.
Published: (2026)
by: Saadat, Mohammadreza, et al.
Published: (2026)
Evidence Contextualization and Counterfactual Attribution for Conversational QA over Heterogeneous Data with RAG Systems
by: Roy, Rishiraj Saha, et al.
Published: (2024)
by: Roy, Rishiraj Saha, et al.
Published: (2024)
Semantic Reformulation Entropy for Robust Hallucination Detection in QA Tasks
by: Tong, Chaodong, et al.
Published: (2025)
by: Tong, Chaodong, et al.
Published: (2025)
A Case Study on Filtering for End-to-End Speech Translation
by: Alam, Md Mahfuz Ibn, et al.
Published: (2024)
by: Alam, Md Mahfuz Ibn, et al.
Published: (2024)
How Robust are the Tabular QA Models for Scientific Tables? A Study using Customized Dataset
by: Ghosh, Akash, et al.
Published: (2024)
by: Ghosh, Akash, et al.
Published: (2024)
Transformers, Contextualism, and Polysemy
by: Grindrod, Jumbly
Published: (2024)
by: Grindrod, Jumbly
Published: (2024)
MasonTigers at SemEval-2024 Task 10: Emotion Discovery and Flip Reasoning in Conversation with Ensemble of Transformers and Prompting
by: Emran, Al Nahian Bin, et al.
Published: (2024)
by: Emran, Al Nahian Bin, et al.
Published: (2024)
Do LLMs Surpass Encoders for Biomedical NER?
by: Obeidat, Motasem S, et al.
Published: (2025)
by: Obeidat, Motasem S, et al.
Published: (2025)
RetinaQA: A Robust Knowledge Base Question Answering Model for both Answerable and Unanswerable Questions
by: Faldu, Prayushi, et al.
Published: (2024)
by: Faldu, Prayushi, et al.
Published: (2024)
RJUA-QA: A Comprehensive QA Dataset for Urology
by: Lyu, Shiwei, et al.
Published: (2023)
by: Lyu, Shiwei, et al.
Published: (2023)
RetrievalQA: Assessing Adaptive Retrieval-Augmented Generation for Short-form Open-Domain Question Answering
by: Zhang, Zihan, et al.
Published: (2024)
by: Zhang, Zihan, et al.
Published: (2024)
Assessing "Implicit" Retrieval Robustness of Large Language Models
by: Shen, Xiaoyu, et al.
Published: (2024)
by: Shen, Xiaoyu, et al.
Published: (2024)
Aspect-Based Opinion Summarization with Argumentation Schemes
by: Zhou, Wendi, et al.
Published: (2025)
by: Zhou, Wendi, et al.
Published: (2025)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
by: Lai, Viet Dac, et al.
Published: (2024)
by: Lai, Viet Dac, et al.
Published: (2024)
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
by: Dineen, Jacob, et al.
Published: (2025)
by: Dineen, Jacob, et al.
Published: (2025)
MasonPerplexity at Multimodal Hate Speech Event Detection 2024: Hate Speech and Target Detection Using Transformer Ensembles
by: Ganguly, Amrita, et al.
Published: (2024)
by: Ganguly, Amrita, et al.
Published: (2024)
Beyond QA Pairs: Assessing Parameter-Efficient Fine-Tuning for Fact Embedding in LLMs
by: Ratnakar, Shivam, et al.
Published: (2025)
by: Ratnakar, Shivam, et al.
Published: (2025)
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
by: Zong, Qing, et al.
Published: (2024)
by: Zong, Qing, et al.
Published: (2024)
CODET: A Benchmark for Contrastive Dialectal Evaluation of Machine Translation
by: Alam, Md Mahfuz Ibn, et al.
Published: (2023)
by: Alam, Md Mahfuz Ibn, et al.
Published: (2023)
Towards Better Question Generation in QA-based Event Extraction
by: Hong, Zijin, et al.
Published: (2024)
by: Hong, Zijin, et al.
Published: (2024)
Similar Items
-
When Not to Answer: Evaluating Prompts on GPT Models for Effective Abstention in Unanswerable Math Word Problems
by: Saadat, Asir, et al.
Published: (2024) -
Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs
by: Le, Chenqian, et al.
Published: (2025) -
JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
by: Cao, Zhihan, et al.
Published: (2025) -
HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA
by: Chen, Xinyue, et al.
Published: (2024) -
UETQuintet at BioCreative IX -- MedHopQA: Enhancing Biomedical QA with Selective Multi-hop Reasoning and Contextual Retrieval
by: Nguyen, Quoc-An, et al.
Published: (2026)