MultiHoax: A Dataset of Multi-hop False-Premise Questions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shafiei, Mohammadamin, Saffari, Hamidreza, Moosavi, Nafise Sadat |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language
von: Saffari, Hamidreza, et al.
Veröffentlicht: (2024)
von: Saffari, Hamidreza, et al.
Veröffentlicht: (2024)
Not All Jokes Land: Evaluating Large Language Models Understanding of Workplace Humor
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
von: Permadi, Vynska Amalia, et al.
Veröffentlicht: (2026)
von: Permadi, Vynska Amalia, et al.
Veröffentlicht: (2026)
Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
von: Pandya, Mugdha, et al.
Veröffentlicht: (2024)
von: Pandya, Mugdha, et al.
Veröffentlicht: (2024)
How to Leverage Digit Embeddings to Represent Numbers?
von: Sivakumar, Jasivan Alex, et al.
Veröffentlicht: (2024)
von: Sivakumar, Jasivan Alex, et al.
Veröffentlicht: (2024)
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
von: Kennedy, Ian W., et al.
Veröffentlicht: (2026)
von: Kennedy, Ian W., et al.
Veröffentlicht: (2026)
From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
von: Mi, Maggie, et al.
Veröffentlicht: (2025)
von: Mi, Maggie, et al.
Veröffentlicht: (2025)
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
von: Liu, Yiqi, et al.
Veröffentlicht: (2023)
von: Liu, Yiqi, et al.
Veröffentlicht: (2023)
Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
von: Mi, Maggie, et al.
Veröffentlicht: (2024)
von: Mi, Maggie, et al.
Veröffentlicht: (2024)
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection
von: Pastorino, Valeria, et al.
Veröffentlicht: (2024)
von: Pastorino, Valeria, et al.
Veröffentlicht: (2024)
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
von: Xue, Huiyin, et al.
Veröffentlicht: (2025)
von: Xue, Huiyin, et al.
Veröffentlicht: (2025)
LLMs Do Not See Age: Assessing Demographic Bias in Automated Systematic Review Synthesis
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2025)
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2025)
Faithful Summarisation under Disagreement via Belief-Level Aggregation
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2026)
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2026)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
von: James, Joseph, et al.
Veröffentlicht: (2026)
von: James, Joseph, et al.
Veröffentlicht: (2026)
Judge Before Answer: Can MLLM Discern the False Premise in Question?
von: Li, Jidong, et al.
Veröffentlicht: (2025)
von: Li, Jidong, et al.
Veröffentlicht: (2025)
Exploring Gender Disparities in Automatic Speech Recognition Technology
von: ElGhazaly, Hend, et al.
Veröffentlicht: (2025)
von: ElGhazaly, Hend, et al.
Veröffentlicht: (2025)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
von: Stacey, Joe, et al.
Veröffentlicht: (2026)
von: Stacey, Joe, et al.
Veröffentlicht: (2026)
MultiCube-RAG for Multi-hop Question Answering
von: Shi, Jimeng, et al.
Veröffentlicht: (2026)
von: Shi, Jimeng, et al.
Veröffentlicht: (2026)
KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions
von: Zhu, Yanxu, et al.
Veröffentlicht: (2024)
von: Zhu, Yanxu, et al.
Veröffentlicht: (2024)
Multi-hop Question Answering
von: Mavi, Vaibhav, et al.
Veröffentlicht: (2022)
von: Mavi, Vaibhav, et al.
Veröffentlicht: (2022)
Hoaxpedia: A Unified Wikipedia Hoax Articles Dataset
von: Borkakoty, Hsuvas, et al.
Veröffentlicht: (2024)
von: Borkakoty, Hsuvas, et al.
Veröffentlicht: (2024)
Can I introduce my boyfriend to my grandmother? Evaluating Large Language Models Capabilities on Iranian Social Norm Classification
von: Saffari, Hamidreza, et al.
Veröffentlicht: (2024)
von: Saffari, Hamidreza, et al.
Veröffentlicht: (2024)
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering
von: Astaraki, Mahdi, et al.
Veröffentlicht: (2026)
von: Astaraki, Mahdi, et al.
Veröffentlicht: (2026)
KCS: Diversify Multi-hop Question Generation with Knowledge Composition Sampling
von: Wang, Yangfan, et al.
Veröffentlicht: (2025)
von: Wang, Yangfan, et al.
Veröffentlicht: (2025)
KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
von: Lee, Jihyung, et al.
Veröffentlicht: (2025)
von: Lee, Jihyung, et al.
Veröffentlicht: (2025)
Self-Critique Guided Iterative Reasoning for Multi-hop Question Answering
von: Chu, Zheng, et al.
Veröffentlicht: (2025)
von: Chu, Zheng, et al.
Veröffentlicht: (2025)
PokeMQA: Programmable knowledge editing for Multi-hop Question Answering
von: Gu, Hengrui, et al.
Veröffentlicht: (2023)
von: Gu, Hengrui, et al.
Veröffentlicht: (2023)
Explainable Multi-hop Question Generation: An End-to-End Approach without Intermediate Question Labeling
von: Hwang, Seonjeong, et al.
Veröffentlicht: (2024)
von: Hwang, Seonjeong, et al.
Veröffentlicht: (2024)
GenDec: A robust generative Question-decomposition method for Multi-hop reasoning
von: Wu, Jian, et al.
Veröffentlicht: (2024)
von: Wu, Jian, et al.
Veröffentlicht: (2024)
TRACE: An Experiential Framework for Coherent Multi-hop Knowledge Graph Question Answering
von: Wang, Yingxu, et al.
Veröffentlicht: (2026)
von: Wang, Yingxu, et al.
Veröffentlicht: (2026)
VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions
von: Su, Hung-Ting, et al.
Veröffentlicht: (2026)
von: Su, Hung-Ting, et al.
Veröffentlicht: (2026)
DEEPAMBIGQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
von: Ji, Jiabao, et al.
Veröffentlicht: (2025)
von: Ji, Jiabao, et al.
Veröffentlicht: (2025)
Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering
von: Shi, Zhengliang, et al.
Veröffentlicht: (2024)
von: Shi, Zhengliang, et al.
Veröffentlicht: (2024)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
RISE: Reasoning Enhancement via Iterative Self-Exploration in Multi-hop Question Answering
von: He, Bolei, et al.
Veröffentlicht: (2025)
von: He, Bolei, et al.
Veröffentlicht: (2025)
HOLMES: Hyper-Relational Knowledge Graphs for Multi-hop Question Answering using LLMs
von: Panda, Pranoy, et al.
Veröffentlicht: (2024)
von: Panda, Pranoy, et al.
Veröffentlicht: (2024)
MQA-KEAL: Multi-hop Question Answering under Knowledge Editing for Arabic Language
von: Ali, Muhammad Asif, et al.
Veröffentlicht: (2024)
von: Ali, Muhammad Asif, et al.
Veröffentlicht: (2024)
Multi-hop Question Answering under Temporal Knowledge Editing
von: Cheng, Keyuan, et al.
Veröffentlicht: (2024)
von: Cheng, Keyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025) -
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language
von: Saffari, Hamidreza, et al.
Veröffentlicht: (2024) -
Not All Jokes Land: Evaluating Large Language Models Understanding of Workplace Humor
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025) -
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
von: Permadi, Vynska Amalia, et al.
Veröffentlicht: (2026) -
Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
von: Pandya, Mugdha, et al.
Veröffentlicht: (2024)