KnowShiftQA: How Robust are RAG Systems when Textbook Knowledge Shifts in K-12 Education?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Tianshi, Li, Weihan, Bai, Jiaxin, Wang, Weiqi, Song, Yangqiu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
von: Zong, Qing, et al.
Veröffentlicht: (2024)
von: Zong, Qing, et al.
Veröffentlicht: (2024)
The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
von: Xu, Baixuan, et al.
Veröffentlicht: (2025)
von: Xu, Baixuan, et al.
Veröffentlicht: (2025)
Advancing Abductive Reasoning in Knowledge Graphs through Complex Logical Hypothesis Generation
von: Bai, Jiaxin, et al.
Veröffentlicht: (2023)
von: Bai, Jiaxin, et al.
Veröffentlicht: (2023)
From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery
von: Zheng, Tianshi, et al.
Veröffentlicht: (2025)
von: Zheng, Tianshi, et al.
Veröffentlicht: (2025)
CLR-Fact: Evaluating the Complex Logical Reasoning Capability of Large Language Models over Factual Knowledge
von: Zheng, Tianshi, et al.
Veröffentlicht: (2024)
von: Zheng, Tianshi, et al.
Veröffentlicht: (2024)
Structuring the Unstructured: A Systematic Review of Text-to-Structure Generation for Agentic AI with a Universal Evaluation Framework
von: Deng, Zheye, et al.
Veröffentlicht: (2025)
von: Deng, Zheye, et al.
Veröffentlicht: (2025)
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
von: Wang, Weiqi, et al.
Veröffentlicht: (2024)
von: Wang, Weiqi, et al.
Veröffentlicht: (2024)
NAACL: Noise-AwAre Verbal Confidence Calibration for Robust LLMs in RAG Systems
von: Liu, Jiayu, et al.
Veröffentlicht: (2026)
von: Liu, Jiayu, et al.
Veröffentlicht: (2026)
EcomEdit: An Automated E-commerce Knowledge Editing Framework for Enhanced Product and Purchase Intention Understanding
von: Lau, Ching Ming Samuel, et al.
Veröffentlicht: (2024)
von: Lau, Ching Ming Samuel, et al.
Veröffentlicht: (2024)
INFERENCEDYNAMICS: Efficient Routing Across LLMs through Structured Capability and Knowledge Profiling
von: Shi, Haochen, et al.
Veröffentlicht: (2025)
von: Shi, Haochen, et al.
Veröffentlicht: (2025)
AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
von: Tsang, Hong Ting, et al.
Veröffentlicht: (2025)
von: Tsang, Hong Ting, et al.
Veröffentlicht: (2025)
ConKE: Conceptualization-Augmented Knowledge Editing in Large Language Models for Commonsense Reasoning
von: Zhang, Liyu, et al.
Veröffentlicht: (2024)
von: Zhang, Liyu, et al.
Veröffentlicht: (2024)
Acquiring and Modelling Abstract Commonsense Knowledge via Conceptualization
von: He, Mutian, et al.
Veröffentlicht: (2022)
von: He, Mutian, et al.
Veröffentlicht: (2022)
Enhancing Transformers for Generalizable First-Order Logical Entailment
von: Zheng, Tianshi, et al.
Veröffentlicht: (2025)
von: Zheng, Tianshi, et al.
Veröffentlicht: (2025)
Towards Multi-Agent Reasoning Systems for Collaborative Expertise Delegation: An Exploratory Design Study
von: Xu, Baixuan, et al.
Veröffentlicht: (2025)
von: Xu, Baixuan, et al.
Veröffentlicht: (2025)
Patterns Over Principles: The Fragility of Inductive Reasoning in LLMs under Noisy Observations
von: Li, Chunyang, et al.
Veröffentlicht: (2025)
von: Li, Chunyang, et al.
Veröffentlicht: (2025)
A Textbook Remedy for Domain Shifts: Knowledge Priors for Medical Image Analysis
von: Yang, Yue, et al.
Veröffentlicht: (2024)
von: Yang, Yue, et al.
Veröffentlicht: (2024)
Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple Extraction
von: Deng, Zheye, et al.
Veröffentlicht: (2024)
von: Deng, Zheye, et al.
Veröffentlicht: (2024)
IntentionQA: A Benchmark for Evaluating Purchase Intention Comprehension Abilities of Language Models in E-commerce
von: Ding, Wenxuan, et al.
Veröffentlicht: (2024)
von: Ding, Wenxuan, et al.
Veröffentlicht: (2024)
LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning
von: Zheng, Tianshi, et al.
Veröffentlicht: (2025)
von: Zheng, Tianshi, et al.
Veröffentlicht: (2025)
Understanding Inter-Session Intentions via Complex Logical Reasoning
von: Bai, Jiaxin, et al.
Veröffentlicht: (2023)
von: Bai, Jiaxin, et al.
Veröffentlicht: (2023)
Intention Knowledge Graph Construction for User Intention Relation Modeling
von: Bai, Jiaxin, et al.
Veröffentlicht: (2024)
von: Bai, Jiaxin, et al.
Veröffentlicht: (2024)
CKBP v2: Better Annotation and Reasoning for Commonsense Knowledge Base Population
von: Fang, Tianqing, et al.
Veröffentlicht: (2023)
von: Fang, Tianqing, et al.
Veröffentlicht: (2023)
Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty?
von: Liu, Jiayu, et al.
Veröffentlicht: (2025)
von: Liu, Jiayu, et al.
Veröffentlicht: (2025)
Legal Rule Induction: Towards Generalizable Principle Discovery from Analogous Judicial Precedents
von: Fan, Wei, et al.
Veröffentlicht: (2025)
von: Fan, Wei, et al.
Veröffentlicht: (2025)
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
von: Zong, Qing, et al.
Veröffentlicht: (2025)
von: Zong, Qing, et al.
Veröffentlicht: (2025)
LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
von: Liang, Fangzhou, et al.
Veröffentlicht: (2025)
von: Liang, Fangzhou, et al.
Veröffentlicht: (2025)
EntailE: Introducing Textual Entailment in Commonsense Knowledge Graph Completion
von: Su, Ying, et al.
Veröffentlicht: (2024)
von: Su, Ying, et al.
Veröffentlicht: (2024)
AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
von: Huang, Haoyu, et al.
Veröffentlicht: (2025)
von: Huang, Haoyu, et al.
Veröffentlicht: (2025)
Towards Subgraph Isomorphism Counting with Graph Kernels
von: Liu, Xin, et al.
Veröffentlicht: (2024)
von: Liu, Xin, et al.
Veröffentlicht: (2024)
GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory
von: Fan, Wei, et al.
Veröffentlicht: (2024)
von: Fan, Wei, et al.
Veröffentlicht: (2024)
UniD$^3$: A Knowledge Graph-Enhanced RAG Framework for Drug-Disease Discovery and Reasoning
von: Wang, Qing, et al.
Veröffentlicht: (2026)
von: Wang, Qing, et al.
Veröffentlicht: (2026)
arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
von: Wang, Weiqi, et al.
Veröffentlicht: (2025)
von: Wang, Weiqi, et al.
Veröffentlicht: (2025)
On the Role of Entity and Event Level Conceptualization in Generalizable Reasoning: A Survey of Tasks, Methods, Applications, and Future Directions
von: Wang, Weiqi, et al.
Veröffentlicht: (2024)
von: Wang, Weiqi, et al.
Veröffentlicht: (2024)
Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
von: Tan, Chenchen, et al.
Veröffentlicht: (2025)
von: Tan, Chenchen, et al.
Veröffentlicht: (2025)
HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA
von: Chen, Xinyue, et al.
Veröffentlicht: (2024)
von: Chen, Xinyue, et al.
Veröffentlicht: (2024)
Transformers for Complex Query Answering over Knowledge Hypergraphs
von: Tsang, Hong Ting, et al.
Veröffentlicht: (2025)
von: Tsang, Hong Ting, et al.
Veröffentlicht: (2025)
QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
Controllable Logical Hypothesis Generation for Abductive Reasoning in Knowledge Graphs
von: Gao, Yisen, et al.
Veröffentlicht: (2025)
von: Gao, Yisen, et al.
Veröffentlicht: (2025)
DeepRefine: Agent-Compiled Knowledge Refinement via Reinforcement Learning
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
von: Zong, Qing, et al.
Veröffentlicht: (2024) -
The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
von: Xu, Baixuan, et al.
Veröffentlicht: (2025) -
Advancing Abductive Reasoning in Knowledge Graphs through Complex Logical Hypothesis Generation
von: Bai, Jiaxin, et al.
Veröffentlicht: (2023) -
From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery
von: Zheng, Tianshi, et al.
Veröffentlicht: (2025) -
CLR-Fact: Evaluating the Complex Logical Reasoning Capability of Large Language Models over Factual Knowledge
von: Zheng, Tianshi, et al.
Veröffentlicht: (2024)