MultiQ&A: An Analysis in Measuring Robustness via Automated Crowdsourcing of Question Perturbations and Answers
Fuente:
arXiv
Saved in:
| Main Authors: | Cho, Nicole, Watson, William |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Is There No Such Thing as a Bad Question? H4R: HalluciBot For Ratiocination, Rewriting, Ranking, and Routing
by: Watson, William, et al.
Published: (2024)
by: Watson, William, et al.
Published: (2024)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
by: Yao, Siyang, et al.
Published: (2026)
by: Yao, Siyang, et al.
Published: (2026)
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
by: Holtermann, Carolin, et al.
Published: (2024)
by: Holtermann, Carolin, et al.
Published: (2024)
QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
by: Cho, Nicole, et al.
Published: (2025)
by: Cho, Nicole, et al.
Published: (2025)
No One Size Fits All: QueryBandits for Hallucination Mitigation
by: Cho, Nicole, et al.
Published: (2026)
by: Cho, Nicole, et al.
Published: (2026)
FISHNET: Financial Intelligence from Sub-querying, Harmonizing, Neural-Conditioning, Expert Swarms, and Task Planning
by: Cho, Nicole, et al.
Published: (2024)
by: Cho, Nicole, et al.
Published: (2024)
The Challenge of Achieving Attributability in Multilingual Table-to-Text Generation with Question-Answer Blueprints
by: Haussmann, Aden
Published: (2025)
by: Haussmann, Aden
Published: (2025)
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions
by: Li, Ruizhe, et al.
Published: (2024)
by: Li, Ruizhe, et al.
Published: (2024)
Explicit Diversity Conditions for Effective Question Answer Generation with Large Language Models
by: Yadav, Vikas, et al.
Published: (2024)
by: Yadav, Vikas, et al.
Published: (2024)
TASER: Table Agents for Schema-guided Extraction and Recommendation
by: Cho, Nicole, et al.
Published: (2025)
by: Cho, Nicole, et al.
Published: (2025)
Measuring and Reducing LLM Hallucination without Gold-Standard Answers
by: Wei, Jiaheng, et al.
Published: (2024)
by: Wei, Jiaheng, et al.
Published: (2024)
MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters
by: Dada, Amin, et al.
Published: (2025)
by: Dada, Amin, et al.
Published: (2025)
Multilingual Non-Factoid Question Answering with Answer Paragraph Selection
by: Mishra, Ritwik, et al.
Published: (2024)
by: Mishra, Ritwik, et al.
Published: (2024)
QUIS: Question-guided Insights Generation for Automated Exploratory Data Analysis
by: Manatkar, Abhijit, et al.
Published: (2024)
by: Manatkar, Abhijit, et al.
Published: (2024)
Can GPT Improve the State of Prior Authorization via Guideline Based Automated Question Answering?
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
SEMQA: Semi-Extractive Multi-Source Question Answering
by: Schuster, Tal, et al.
Published: (2023)
by: Schuster, Tal, et al.
Published: (2023)
LinkQ: An LLM-Assisted Visual Interface for Knowledge Graph Question-Answering
by: Li, Harry, et al.
Published: (2024)
by: Li, Harry, et al.
Published: (2024)
Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
by: Wang, Zhilin, et al.
Published: (2025)
by: Wang, Zhilin, et al.
Published: (2025)
Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
by: Wang, Dongwei, et al.
Published: (2024)
by: Wang, Dongwei, et al.
Published: (2024)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
by: Fu, Tingchen, et al.
Published: (2025)
by: Fu, Tingchen, et al.
Published: (2025)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
by: Tian, Yijun, et al.
Published: (2024)
by: Tian, Yijun, et al.
Published: (2024)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
Evidence-Focused Fact Summarization for Knowledge-Augmented Zero-Shot Question Answering
by: Ko, Sungho, et al.
Published: (2024)
by: Ko, Sungho, et al.
Published: (2024)
Statistical Comparative Analysis of Semantic Similarities and Model Transferability Across Datasets for Short Answer Grading
by: Bonthu, Sridevi, et al.
Published: (2025)
by: Bonthu, Sridevi, et al.
Published: (2025)
From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
by: Li, Tianle, et al.
Published: (2024)
by: Li, Tianle, et al.
Published: (2024)
ASAG2024: A Combined Benchmark for Short Answer Grading
by: Meyer, Gérôme, et al.
Published: (2024)
by: Meyer, Gérôme, et al.
Published: (2024)
Proving that Cryptic Crossword Clue Answers are Correct
by: Andrews, Martin, et al.
Published: (2024)
by: Andrews, Martin, et al.
Published: (2024)
Verif.ai: Towards an Open-Source Scientific Generative Question-Answering System with Referenced and Verifiable Answers
by: Košprdić, Miloš, et al.
Published: (2024)
by: Košprdić, Miloš, et al.
Published: (2024)
Remote Labor Index: Measuring AI Automation of Remote Work
by: Mazeika, Mantas, et al.
Published: (2025)
by: Mazeika, Mantas, et al.
Published: (2025)
Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle
by: Wang, Zihan, et al.
Published: (2026)
by: Wang, Zihan, et al.
Published: (2026)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance
by: Watson, William, et al.
Published: (2026)
by: Watson, William, et al.
Published: (2026)
Towards A Unified View of Answer Calibration for Multi-Step Reasoning
by: Deng, Shumin, et al.
Published: (2023)
by: Deng, Shumin, et al.
Published: (2023)
Jailbreak-as-a-Service++: Unveiling Distributed AI-Driven Malicious Information Campaigns Powered by LLM Crowdsourcing
by: Yan, Yu, et al.
Published: (2025)
by: Yan, Yu, et al.
Published: (2025)
Multi-hop Question Answering under Temporal Knowledge Editing
by: Cheng, Keyuan, et al.
Published: (2024)
by: Cheng, Keyuan, et al.
Published: (2024)
AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning
by: Zhang, Jiayi, et al.
Published: (2025)
by: Zhang, Jiayi, et al.
Published: (2025)
Towards Automated Patent Workflows: AI-Orchestrated Multi-Agent Framework for Intellectual Property Management and Analysis
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
CRACQ: A Multi-Dimensional Approach To Automated Document Assessment
by: Soltani, Ishak, et al.
Published: (2025)
by: Soltani, Ishak, et al.
Published: (2025)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Similar Items
-
Is There No Such Thing as a Bad Question? H4R: HalluciBot For Ratiocination, Rewriting, Ranking, and Routing
by: Watson, William, et al.
Published: (2024) -
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
by: Yao, Siyang, et al.
Published: (2026) -
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
by: Holtermann, Carolin, et al.
Published: (2024) -
QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
by: Cho, Nicole, et al.
Published: (2025) -
No One Size Fits All: QueryBandits for Hallucination Mitigation
by: Cho, Nicole, et al.
Published: (2026)