Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Wenyu, Vougiouklis, Pavlos, Lapata, Mirella, Pan, Jeff Z. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Less is More: Making Smaller Language Models Competent Subgraph Retrievers for Multi-hop KGQA
by: Huang, Wenyu, et al.
Published: (2024)
by: Huang, Wenyu, et al.
Published: (2024)
Prompting Large Language Models with Knowledge Graphs for Question Answering Involving Long-tail Facts
by: Huang, Wenyu, et al.
Published: (2024)
by: Huang, Wenyu, et al.
Published: (2024)
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
by: Zheng, Danna, et al.
Published: (2024)
by: Zheng, Danna, et al.
Published: (2024)
Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning
by: Zheng, Danna, et al.
Published: (2024)
by: Zheng, Danna, et al.
Published: (2024)
Long-Form Information Alignment Evaluation Beyond Atomic Facts
by: Zheng, Danna, et al.
Published: (2025)
by: Zheng, Danna, et al.
Published: (2025)
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
by: Zheng, Danna, et al.
Published: (2024)
by: Zheng, Danna, et al.
Published: (2024)
Context-Aware Hierarchical Merging for Long Document Summarization
by: Ou, Litu, et al.
Published: (2025)
by: Ou, Litu, et al.
Published: (2025)
A Usage-centric Take on Intent Understanding in E-Commerce
by: Zhou, Wendi, et al.
Published: (2024)
by: Zhou, Wendi, et al.
Published: (2024)
Integrating Large Language Models with Graph-based Reasoning for Conversational Question Answering
by: Jain, Parag, et al.
Published: (2024)
by: Jain, Parag, et al.
Published: (2024)
Millions of $\text{GeAR}$-s: Extending GraphRAG to Millions of Documents
by: Shen, Zhili, et al.
Published: (2025)
by: Shen, Zhili, et al.
Published: (2025)
Improving Generalization in Semantic Parsing by Increasing Natural Language Variation
by: Saparina, Irina, et al.
Published: (2024)
by: Saparina, Irina, et al.
Published: (2024)
Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books
by: Papoudakis, Argyrios, et al.
Published: (2026)
by: Papoudakis, Argyrios, et al.
Published: (2026)
Rethinking Memory in LLM based Agents: Representations, Operations, and Emerging Topics
by: Du, Yiming, et al.
Published: (2025)
by: Du, Yiming, et al.
Published: (2025)
Improving Retrieval-augmented Text-to-SQL with AST-based Ranking and Schema Pruning
by: Shen, Zhili, et al.
Published: (2024)
by: Shen, Zhili, et al.
Published: (2024)
Learning to Reason for Long-Form Story Generation
by: Gurung, Alexander, et al.
Published: (2025)
by: Gurung, Alexander, et al.
Published: (2025)
A Modular Approach for Multimodal Summarization of TV Shows
by: Mahon, Louis, et al.
Published: (2024)
by: Mahon, Louis, et al.
Published: (2024)
AMBROSIA: A Benchmark for Parsing Ambiguous Questions into Database Queries
by: Saparina, Irina, et al.
Published: (2024)
by: Saparina, Irina, et al.
Published: (2024)
CHIRON: Rich Character Representations in Long-Form Narratives
by: Gurung, Alexander, et al.
Published: (2024)
by: Gurung, Alexander, et al.
Published: (2024)
OpenSIR: Open-Ended Self-Improving Reasoner
by: Kwan, Wai-Chung, et al.
Published: (2025)
by: Kwan, Wai-Chung, et al.
Published: (2025)
Reasoning about Intent for Ambiguous Requests
by: Saparina, Irina, et al.
Published: (2025)
by: Saparina, Irina, et al.
Published: (2025)
Debating for Better Reasoning: An Unsupervised Multimodal Approach
by: Adhikari, Ashutosh, et al.
Published: (2025)
by: Adhikari, Ashutosh, et al.
Published: (2025)
Disambiguate First, Parse Later: Generating Interpretations for Ambiguity Resolution in Semantic Parsing
by: Saparina, Irina, et al.
Published: (2025)
by: Saparina, Irina, et al.
Published: (2025)
Uncertainty Quantification in Retrieval Augmented Question Answering
by: Perez-Beltrachini, Laura, et al.
Published: (2025)
by: Perez-Beltrachini, Laura, et al.
Published: (2025)
SimLM: Can Language Models Infer Parameters of Physical Systems?
by: Memery, Sean, et al.
Published: (2023)
by: Memery, Sean, et al.
Published: (2023)
An Extensive Evaluation of PDDL Capabilities in off-the-shelf LLMs
by: Vyas, Kaustubh, et al.
Published: (2025)
by: Vyas, Kaustubh, et al.
Published: (2025)
Lightweight Latent Reasoning for Narrative Tasks
by: Gurung, Alexander, et al.
Published: (2025)
by: Gurung, Alexander, et al.
Published: (2025)
PixT3: Pixel-based Table-To-Text Generation
by: Alonso, Iñigo, et al.
Published: (2023)
by: Alonso, Iñigo, et al.
Published: (2023)
Hierarchical Indexing for Retrieval-Augmented Opinion Summarization
by: Hosking, Tom, et al.
Published: (2024)
by: Hosking, Tom, et al.
Published: (2024)
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
by: Li, Miao, et al.
Published: (2026)
by: Li, Miao, et al.
Published: (2026)
BookWorm: A Dataset for Character Description and Analysis
by: Papoudakis, Argyrios, et al.
Published: (2024)
by: Papoudakis, Argyrios, et al.
Published: (2024)
MoreHopQA: More Than Multi-hop Reasoning
by: Schnitzler, Julian, et al.
Published: (2024)
by: Schnitzler, Julian, et al.
Published: (2024)
Little Red Riding Hood Goes Around the Globe:Crosslingual Story Planning and Generation with Large Language Models
by: Razumovskaia, Evgeniia, et al.
Published: (2022)
by: Razumovskaia, Evgeniia, et al.
Published: (2022)
Generating Visual Stories with Grounded and Coreferent Characters
by: Liu, Danyang, et al.
Published: (2024)
by: Liu, Danyang, et al.
Published: (2024)
Automatic Inter-document Multi-hop Scientific QA Generation
by: Lee, Seungmin, et al.
Published: (2026)
by: Lee, Seungmin, et al.
Published: (2026)
Compositional Generalisation for Explainable Hate Speech Detection
by: Calabrese, Agostina, et al.
Published: (2025)
by: Calabrese, Agostina, et al.
Published: (2025)
Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback
by: Rashkin, Hannah, et al.
Published: (2025)
by: Rashkin, Hannah, et al.
Published: (2025)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
by: Gupta, Akash, et al.
Published: (2025)
by: Gupta, Akash, et al.
Published: (2025)
A Controllable Examination for Long-Context Language Models
by: Yang, Yijun, et al.
Published: (2025)
by: Yang, Yijun, et al.
Published: (2025)
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
by: Eisenstein, Jacob, et al.
Published: (2026)
by: Eisenstein, Jacob, et al.
Published: (2026)
Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark
by: Wu, Jian, et al.
Published: (2024)
by: Wu, Jian, et al.
Published: (2024)
Similar Items
-
Less is More: Making Smaller Language Models Competent Subgraph Retrievers for Multi-hop KGQA
by: Huang, Wenyu, et al.
Published: (2024) -
Prompting Large Language Models with Knowledge Graphs for Question Answering Involving Long-tail Facts
by: Huang, Wenyu, et al.
Published: (2024) -
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
by: Zheng, Danna, et al.
Published: (2024) -
Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning
by: Zheng, Danna, et al.
Published: (2024) -
Long-Form Information Alignment Evaluation Beyond Atomic Facts
by: Zheng, Danna, et al.
Published: (2025)