ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ghosh, Shubhra, Borah, Abhilekh, Guru, Aditya Kumar, Ghosh, Kripabandhu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
von: Khanna, Danush, et al.
Veröffentlicht: (2025)
von: Khanna, Danush, et al.
Veröffentlicht: (2025)
Calibrated Confidence Estimation for Tabular Question Answering
von: Voss, Lukas
Veröffentlicht: (2026)
von: Voss, Lukas
Veröffentlicht: (2026)
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering
von: Muller, Sacha, et al.
Veröffentlicht: (2024)
von: Muller, Sacha, et al.
Veröffentlicht: (2024)
Revisiting Word Embeddings in the LLM Era
von: Mahajan, Yash, et al.
Veröffentlicht: (2024)
von: Mahajan, Yash, et al.
Veröffentlicht: (2024)
LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
von: Nguyen, Huyen, et al.
Veröffentlicht: (2026)
von: Nguyen, Huyen, et al.
Veröffentlicht: (2026)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
von: Dejl, Adam, et al.
Veröffentlicht: (2025)
von: Dejl, Adam, et al.
Veröffentlicht: (2025)
CR-LT-KGQA: A Knowledge Graph Question Answering Dataset Requiring Commonsense Reasoning and Long-Tail Knowledge
von: Guo, Willis, et al.
Veröffentlicht: (2024)
von: Guo, Willis, et al.
Veröffentlicht: (2024)
Generator-Guided Crowd Reaction Assessment
von: Ghosh, Sohom, et al.
Veröffentlicht: (2024)
von: Ghosh, Sohom, et al.
Veröffentlicht: (2024)
AskSport: Web Application for Sports Question-Answering
von: Onofre, Enzo B, et al.
Veröffentlicht: (2025)
von: Onofre, Enzo B, et al.
Veröffentlicht: (2025)
Question Answering Over Spatio-Temporal Knowledge Graph
von: Dai, Xinbang, et al.
Veröffentlicht: (2024)
von: Dai, Xinbang, et al.
Veröffentlicht: (2024)
Exploration of Augmentation Strategies in Multi-modal Retrieval-Augmented Generation for the Biomedical Domain: A Case Study Evaluating Question Answering in Glycobiology
von: Kocbek, Primož, et al.
Veröffentlicht: (2025)
von: Kocbek, Primož, et al.
Veröffentlicht: (2025)
MIX : a Multi-task Learning Approach to Solve Open-Domain Question Answering
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2020)
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2020)
EfficientQA : a RoBERTa Based Phrase-Indexed Question-Answering System
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2021)
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2021)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
BERTologyNavigator: Advanced Question Answering with BERT-based Semantics
von: Rajpal, Shreya, et al.
Veröffentlicht: (2024)
von: Rajpal, Shreya, et al.
Veröffentlicht: (2024)
Right for Right Reasons: Large Language Models for Verifiable Commonsense Knowledge Graph Question Answering
von: Toroghi, Armin, et al.
Veröffentlicht: (2024)
von: Toroghi, Armin, et al.
Veröffentlicht: (2024)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
von: Wang, Yuxia, et al.
Veröffentlicht: (2024)
von: Wang, Yuxia, et al.
Veröffentlicht: (2024)
Enhancing Software-Related Information Extraction via Single-Choice Question Answering with Large Language Models
von: Otto, Wolfgang, et al.
Veröffentlicht: (2024)
von: Otto, Wolfgang, et al.
Veröffentlicht: (2024)
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
von: Iqbal, Hasan, et al.
Veröffentlicht: (2024)
von: Iqbal, Hasan, et al.
Veröffentlicht: (2024)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
von: Palit, Sayon, et al.
Veröffentlicht: (2025)
von: Palit, Sayon, et al.
Veröffentlicht: (2025)
MedFuzz: Exploring the Robustness of Large Language Models in Medical Question Answering
von: Ness, Robert Osazuwa, et al.
Veröffentlicht: (2024)
von: Ness, Robert Osazuwa, et al.
Veröffentlicht: (2024)
How Do Large Language Models Acquire Factual Knowledge During Pretraining?
von: Chang, Hoyeon, et al.
Veröffentlicht: (2024)
von: Chang, Hoyeon, et al.
Veröffentlicht: (2024)
Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
von: Dayarathne, Ranul, et al.
Veröffentlicht: (2025)
von: Dayarathne, Ranul, et al.
Veröffentlicht: (2025)
Graph Guided Question Answer Generation for Procedural Question-Answering
von: Pham, Hai X., et al.
Veröffentlicht: (2024)
von: Pham, Hai X., et al.
Veröffentlicht: (2024)
DQA: Diagnostic Question Answering for IT Support
von: Kapoor, Vishaal, et al.
Veröffentlicht: (2026)
von: Kapoor, Vishaal, et al.
Veröffentlicht: (2026)
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
von: Hashemi, Helia, et al.
Veröffentlicht: (2024)
von: Hashemi, Helia, et al.
Veröffentlicht: (2024)
A Graph-based RAG for Energy Efficiency Question Answering
von: Campi, Riccardo, et al.
Veröffentlicht: (2025)
von: Campi, Riccardo, et al.
Veröffentlicht: (2025)
QuAnTS: Question Answering on Time Series
von: Divo, Felix, et al.
Veröffentlicht: (2025)
von: Divo, Felix, et al.
Veröffentlicht: (2025)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
von: Karinshak, Elise, et al.
Veröffentlicht: (2024)
von: Karinshak, Elise, et al.
Veröffentlicht: (2024)
Identifying Fairness Issues in Automatically Generated Testing Content
von: Stowe, Kevin, et al.
Veröffentlicht: (2024)
von: Stowe, Kevin, et al.
Veröffentlicht: (2024)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
von: Souza, Débora, et al.
Veröffentlicht: (2026)
von: Souza, Débora, et al.
Veröffentlicht: (2026)
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
von: Badshah, Sher, et al.
Veröffentlicht: (2025)
von: Badshah, Sher, et al.
Veröffentlicht: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
DWFS-Obfuscation: Dynamic Weighted Feature Selection for Robust Malware Familial Classification under Obfuscation
von: Wei, Xingyuan, et al.
Veröffentlicht: (2025)
von: Wei, Xingyuan, et al.
Veröffentlicht: (2025)
Robustness of Large Language Models to Perturbations in Text
von: Singh, Ayush, et al.
Veröffentlicht: (2024)
von: Singh, Ayush, et al.
Veröffentlicht: (2024)
Meta-Evaluation of Translation Evaluation Methods: a systematic up-to-date overview
von: Han, Lifeng, et al.
Veröffentlicht: (2016)
von: Han, Lifeng, et al.
Veröffentlicht: (2016)
Temporal Knowledge Question Answering via Abstract Reasoning Induction
von: Chen, Ziyang, et al.
Veröffentlicht: (2023)
von: Chen, Ziyang, et al.
Veröffentlicht: (2023)
TrustAI at SemEval-2024 Task 8: A Comprehensive Analysis of Multi-domain Machine Generated Text Detection Techniques
von: Urlana, Ashok, et al.
Veröffentlicht: (2024)
von: Urlana, Ashok, et al.
Veröffentlicht: (2024)
ALISON: Fast and Effective Stylometric Authorship Obfuscation
von: Xing, Eric, et al.
Veröffentlicht: (2024)
von: Xing, Eric, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
von: Khanna, Danush, et al.
Veröffentlicht: (2025) -
Calibrated Confidence Estimation for Tabular Question Answering
von: Voss, Lukas
Veröffentlicht: (2026) -
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering
von: Muller, Sacha, et al.
Veröffentlicht: (2024) -
Revisiting Word Embeddings in the LLM Era
von: Mahajan, Yash, et al.
Veröffentlicht: (2024) -
LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
von: Nguyen, Huyen, et al.
Veröffentlicht: (2026)