MoreHopQA: More Than Multi-hop Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schnitzler, Julian, Ho, Xanh, Huang, Jiahao, Boudin, Florian, Sugawara, Saku, Aizawa, Akiko |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
Unsupervised Domain Adaptation for Keyphrase Generation using Citation Contexts
von: Boudin, Florian, et al.
Veröffentlicht: (2024)
von: Boudin, Florian, et al.
Veröffentlicht: (2024)
Table-Text Alignment: Explaining Claim Verification Against Tables in Scientific Papers
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
von: Ho, Xanh, et al.
Veröffentlicht: (2025)
Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
von: Kumar, Sunisth, et al.
Veröffentlicht: (2026)
von: Kumar, Sunisth, et al.
Veröffentlicht: (2026)
An Analysis of Datasets, Metrics and Models in Keyphrase Generation
von: Boudin, Florian, et al.
Veröffentlicht: (2025)
von: Boudin, Florian, et al.
Veröffentlicht: (2025)
Preface to the Special Issue of the TAL Journal on Scholarly Document Processing
von: Boudin, Florian, et al.
Veröffentlicht: (2025)
von: Boudin, Florian, et al.
Veröffentlicht: (2025)
Automatically Suggesting Diverse Example Sentences for L2 Japanese Learners Using Pre-Trained Language Models
von: Benedetti, Enrico, et al.
Veröffentlicht: (2025)
von: Benedetti, Enrico, et al.
Veröffentlicht: (2025)
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
von: Ho, Xanh, et al.
Veröffentlicht: (2026)
von: Ho, Xanh, et al.
Veröffentlicht: (2026)
A Survey of Pre-trained Language Models for Processing Scientific Text
von: Ho, Xanh, et al.
Veröffentlicht: (2024)
von: Ho, Xanh, et al.
Veröffentlicht: (2024)
Self-Compositional Data Augmentation for Scientific Keyphrase Generation
von: Houbre, Mael, et al.
Veröffentlicht: (2024)
von: Houbre, Mael, et al.
Veröffentlicht: (2024)
ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction
von: Jourdan, Léane, et al.
Veröffentlicht: (2025)
von: Jourdan, Léane, et al.
Veröffentlicht: (2025)
UETQuintet at BioCreative IX -- MedHopQA: Enhancing Biomedical QA with Selective Multi-hop Reasoning and Contextual Retrieval
von: Nguyen, Quoc-An, et al.
Veröffentlicht: (2026)
von: Nguyen, Quoc-An, et al.
Veröffentlicht: (2026)
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models
von: Jiang, Junfeng, et al.
Veröffentlicht: (2024)
von: Jiang, Junfeng, et al.
Veröffentlicht: (2024)
Rationale-Aware Answer Verification by Pairwise Self-Evaluation
von: Kawabata, Akira, et al.
Veröffentlicht: (2024)
von: Kawabata, Akira, et al.
Veröffentlicht: (2024)
What Makes Language Models Good-enough?
von: Asami, Daiki, et al.
Veröffentlicht: (2024)
von: Asami, Daiki, et al.
Veröffentlicht: (2024)
A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
von: Emura, Rei, et al.
Veröffentlicht: (2026)
von: Emura, Rei, et al.
Veröffentlicht: (2026)
CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
von: Oba, Miyu, et al.
Veröffentlicht: (2026)
von: Oba, Miyu, et al.
Veröffentlicht: (2026)
Specification-Aware Machine Translation and Evaluation for Purpose Alignment
von: Kayano, Yoko, et al.
Veröffentlicht: (2025)
von: Kayano, Yoko, et al.
Veröffentlicht: (2025)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
von: Kim, Kon Woo, et al.
Veröffentlicht: (2025)
von: Kim, Kon Woo, et al.
Veröffentlicht: (2025)
TactfulToM: Do LLMs Have the Theory of Mind Ability to Understand White Lies?
von: Liu, Yiwei, et al.
Veröffentlicht: (2025)
von: Liu, Yiwei, et al.
Veröffentlicht: (2025)
Long Is More Important Than Difficult for Training Reasoning Models
von: Shen, Si, et al.
Veröffentlicht: (2025)
von: Shen, Si, et al.
Veröffentlicht: (2025)
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
von: Kawabata, Akira, et al.
Veröffentlicht: (2026)
von: Kawabata, Akira, et al.
Veröffentlicht: (2026)
DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections
von: Park, Jiwon, et al.
Veröffentlicht: (2025)
von: Park, Jiwon, et al.
Veröffentlicht: (2025)
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
von: Gupta, Abhay, et al.
Veröffentlicht: (2025)
von: Gupta, Abhay, et al.
Veröffentlicht: (2025)
Less is More: Making Smaller Language Models Competent Subgraph Retrievers for Multi-hop KGQA
von: Huang, Wenyu, et al.
Veröffentlicht: (2024)
von: Huang, Wenyu, et al.
Veröffentlicht: (2024)
FrugalRAG: Less is More in RL Finetuning for Multi-Hop Question Answering
von: Java, Abhinav, et al.
Veröffentlicht: (2025)
von: Java, Abhinav, et al.
Veröffentlicht: (2025)
RELOOP: Recursive Retrieval with Multi-Hop Reasoner and Planners for Heterogeneous QA
von: Yang, Ruiyi, et al.
Veröffentlicht: (2025)
von: Yang, Ruiyi, et al.
Veröffentlicht: (2025)
DeepRAG: Integrating Hierarchical Reasoning and Process Supervision for Biomedical Multi-Hop QA
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems
von: Junqueras, Juan, et al.
Veröffentlicht: (2026)
von: Junqueras, Juan, et al.
Veröffentlicht: (2026)
Are Emotions Arranged in a Circle? Geometric Analysis of Emotion Representations via Hyperspherical Contrastive Learning
von: Yamauchi, Yusuke, et al.
Veröffentlicht: (2026)
von: Yamauchi, Yusuke, et al.
Veröffentlicht: (2026)
Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
Robust Knowledge Editing via Explicit Reasoning Chains for Distractor-Resilient Multi-Hop QA
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
Automatic Inter-document Multi-hop Scientific QA Generation
von: Lee, Seungmin, et al.
Veröffentlicht: (2026)
von: Lee, Seungmin, et al.
Veröffentlicht: (2026)
Tokenization Is More Than Compression
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
Trivial Vocabulary Bans Improve LLM Reasoning More Than Deep Linguistic Constraints
von: Jehu-Appiah, Rodney
Veröffentlicht: (2026)
von: Jehu-Appiah, Rodney
Veröffentlicht: (2026)
Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation
von: Huang, Wenyu, et al.
Veröffentlicht: (2025)
von: Huang, Wenyu, et al.
Veröffentlicht: (2025)
Should LLM Safety Be More Than Refusing Harmful Instructions?
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
Confidence Should Be Calibrated More Than One Turn Deep
von: Zhang, Zhaohan, et al.
Veröffentlicht: (2026)
von: Zhang, Zhaohan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
von: Ho, Xanh, et al.
Veröffentlicht: (2025) -
Unsupervised Domain Adaptation for Keyphrase Generation using Citation Contexts
von: Boudin, Florian, et al.
Veröffentlicht: (2024) -
Table-Text Alignment: Explaining Claim Verification Against Tables in Scientific Papers
von: Ho, Xanh, et al.
Veröffentlicht: (2025) -
Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts
von: Ho, Xanh, et al.
Veröffentlicht: (2025) -
Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
von: Kumar, Sunisth, et al.
Veröffentlicht: (2026)