A Fano-Style Accuracy Upper Bound for LLM Single-Pass Reasoning in Multi-Hop QA
Fuente:
arXiv
Saved in:
| Main Authors: | Wan, Kaiyang, Gao, Lang, Mu, Honglin, Nakov, Preslav, Wang, Yuxia, Chen, Xiuying |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI Security Beyond Core Domains: Resume Screening as a Case Study of Adversarial Vulnerabilities in Specialized LLM Applications
by: Mu, Honglin, et al.
Published: (2025)
by: Mu, Honglin, et al.
Published: (2025)
A Cognitive Writing Perspective for Constrained Long-Form Text Generation
by: Wan, Kaiyang, et al.
Published: (2025)
by: Wan, Kaiyang, et al.
Published: (2025)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
by: Mekky, Ali, et al.
Published: (2025)
by: Mekky, Ali, et al.
Published: (2025)
The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations
by: Elbadry, Rania, et al.
Published: (2026)
by: Elbadry, Rania, et al.
Published: (2026)
ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA
by: Zhao, Xinjie, et al.
Published: (2025)
by: Zhao, Xinjie, et al.
Published: (2025)
Fast or Better? Balancing Accuracy and Cost in Retrieval-Augmented Generation with Flexible User Control
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
VSCBench: Bridging the Gap in Vision-Language Model Safety Calibration
by: Geng, Jiahui, et al.
Published: (2025)
by: Geng, Jiahui, et al.
Published: (2025)
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
by: Islamaj, Rezarta, et al.
Published: (2026)
by: Islamaj, Rezarta, et al.
Published: (2026)
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection
by: Gao, Lang, et al.
Published: (2025)
by: Gao, Lang, et al.
Published: (2025)
RELOOP: Recursive Retrieval with Multi-Hop Reasoner and Planners for Heterogeneous QA
by: Yang, Ruiyi, et al.
Published: (2025)
by: Yang, Ruiyi, et al.
Published: (2025)
A Survey of Confidence Estimation and Calibration in Large Language Models
by: Geng, Jiahui, et al.
Published: (2023)
by: Geng, Jiahui, et al.
Published: (2023)
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
by: Sundriyal, Megha, et al.
Published: (2023)
by: Sundriyal, Megha, et al.
Published: (2023)
DenoiseRank: Learning to Rank by Diffusion Models
by: Wang, Ying, et al.
Published: (2026)
by: Wang, Ying, et al.
Published: (2026)
Adapting Fake News Detection to the Era of Large Language Models
by: Su, Jinyan, et al.
Published: (2023)
by: Su, Jinyan, et al.
Published: (2023)
DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QA
by: Wang, Changhao, et al.
Published: (2025)
by: Wang, Changhao, et al.
Published: (2025)
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
by: Gao, Lang, et al.
Published: (2025)
by: Gao, Lang, et al.
Published: (2025)
FinReflectKG -- MultiHop: Financial QA Benchmark for Reasoning with Knowledge Graph Evidence
by: Arun, Abhinav, et al.
Published: (2025)
by: Arun, Abhinav, et al.
Published: (2025)
MALMM: Multi-Agent Large Language Models for Zero-Shot Robotics Manipulation
by: Singh, Harsh, et al.
Published: (2024)
by: Singh, Harsh, et al.
Published: (2024)
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
by: Iqbal, Hasan, et al.
Published: (2024)
by: Iqbal, Hasan, et al.
Published: (2024)
Factuality of Large Language Models: A Survey
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
Multimodal Large Language Models to Support Real-World Fact-Checking
by: Geng, Jiahui, et al.
Published: (2024)
by: Geng, Jiahui, et al.
Published: (2024)
MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
by: Elozeiri, Kareem, et al.
Published: (2025)
by: Elozeiri, Kareem, et al.
Published: (2025)
Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck
by: Zhang, Meiru, et al.
Published: (2026)
by: Zhang, Meiru, et al.
Published: (2026)
A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models
by: Geng, Jiahui, et al.
Published: (2025)
by: Geng, Jiahui, et al.
Published: (2025)
HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning
by: Wang, Shenzhi, et al.
Published: (2026)
by: Wang, Shenzhi, et al.
Published: (2026)
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
by: Gao, Lang, et al.
Published: (2024)
by: Gao, Lang, et al.
Published: (2024)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
by: Almheiri, Saeed, et al.
Published: (2025)
by: Almheiri, Saeed, et al.
Published: (2025)
Beyond Generation: Multi-Hop Reasoning for Factual Accuracy in Vision-Language Models
by: Hossain, Shamima
Published: (2025)
by: Hossain, Shamima
Published: (2025)
COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations
by: Xing, Rui, et al.
Published: (2025)
by: Xing, Rui, et al.
Published: (2025)
Detecting Propaganda Techniques in Code-Switched Social Media Text
by: Salman, Muhammad Umar, et al.
Published: (2023)
by: Salman, Muhammad Umar, et al.
Published: (2023)
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
by: Thareja, Rushil, et al.
Published: (2025)
by: Thareja, Rushil, et al.
Published: (2025)
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
by: Rubashevskii, Aleksandr, et al.
Published: (2026)
by: Rubashevskii, Aleksandr, et al.
Published: (2026)
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
by: Orel, Daniil, et al.
Published: (2025)
by: Orel, Daniil, et al.
Published: (2025)
Under the Shadow of Babel: How Language Shapes Reasoning in LLMs
by: Wang, Chenxi, et al.
Published: (2025)
by: Wang, Chenxi, et al.
Published: (2025)
From Single to Multi-Agent Reasoning: Advancing GeneGPT for Genomics QA
by: Abedini, Kimia, et al.
Published: (2026)
by: Abedini, Kimia, et al.
Published: (2026)
How Does Prefix Matter in Reasoning Model Tuning?
by: Tomar, Raj Vardhan, et al.
Published: (2026)
by: Tomar, Raj Vardhan, et al.
Published: (2026)
FIRE: Fact-checking with Iterative Retrieval and Verification
by: Xie, Zhuohan, et al.
Published: (2024)
by: Xie, Zhuohan, et al.
Published: (2024)
UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
by: Tomar, Raj Vardhan, et al.
Published: (2025)
by: Tomar, Raj Vardhan, et al.
Published: (2025)
Multi-Sourced, Multi-Agent Evidence Retrieval for Fact-Checking
by: Gong, Shuzhi, et al.
Published: (2026)
by: Gong, Shuzhi, et al.
Published: (2026)
Similar Items
-
AI Security Beyond Core Domains: Resume Screening as a Case Study of Adversarial Vulnerabilities in Specialized LLM Applications
by: Mu, Honglin, et al.
Published: (2025) -
A Cognitive Writing Perspective for Constrained Long-Form Text Generation
by: Wan, Kaiyang, et al.
Published: (2025) -
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
by: Mekky, Ali, et al.
Published: (2025) -
The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations
by: Elbadry, Rania, et al.
Published: (2026) -
ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA
by: Zhao, Xinjie, et al.
Published: (2025)