SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zhiqiang, Niu, Enpei, Hua, Yin, Sun, Mengshu, Liang, Lei, Chen, Huajun, Zhang, Wen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
K-ON: Stacking Knowledge On the Head Layer of Large Language Model
von: Guo, Lingbing, et al.
Veröffentlicht: (2025)
von: Guo, Lingbing, et al.
Veröffentlicht: (2025)
Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis
von: Yuan, Lin, et al.
Veröffentlicht: (2025)
von: Yuan, Lin, et al.
Veröffentlicht: (2025)
Retrieve, Summarize, Plan: Advancing Multi-hop Question Answering with an Iterative Approach
von: Jiang, Zhouyu, et al.
Veröffentlicht: (2024)
von: Jiang, Zhouyu, et al.
Veröffentlicht: (2024)
Have We Designed Generalizable Structural Knowledge Promptings? Systematic Evaluation and Rethinking
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
OneGen: Efficient One-Pass Unified Generation and Retrieval for LLMs
von: Zhang, Jintian, et al.
Veröffentlicht: (2024)
von: Zhang, Jintian, et al.
Veröffentlicht: (2024)
RTQA : Recursive Thinking for Complex Temporal Knowledge Graph Question Answering with Large Language Models
von: Gong, Zhaoyan, et al.
Veröffentlicht: (2025)
von: Gong, Zhaoyan, et al.
Veröffentlicht: (2025)
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
von: Li, Youquan, et al.
Veröffentlicht: (2024)
von: Li, Youquan, et al.
Veröffentlicht: (2024)
Collaboration of Fusion and Independence: Hypercomplex-driven Robust Multi-Modal Knowledge Graph Completion
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2025)
KAG: Boosting LLMs in Professional Domains via Knowledge Augmented Generation
von: Liang, Lei, et al.
Veröffentlicht: (2024)
von: Liang, Lei, et al.
Veröffentlicht: (2024)
UniHR: Hierarchical Representation Learning for Unified Knowledge Graph Link Prediction
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2024)
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2024)
Bi'an: A Bilingual Benchmark and Model for Hallucination Detection in Retrieval-Augmented Generation
von: Jiang, Zhouyu, et al.
Veröffentlicht: (2025)
von: Jiang, Zhouyu, et al.
Veröffentlicht: (2025)
Knowledgeable Preference Alignment for LLMs in Domain-specific Question Answering
von: Zhang, Yichi, et al.
Veröffentlicht: (2023)
von: Zhang, Yichi, et al.
Veröffentlicht: (2023)
SciCUEval: A Comprehensive Dataset for Evaluating Scientific Context Understanding in Large Language Models
von: Yu, Jing, et al.
Veröffentlicht: (2025)
von: Yu, Jing, et al.
Veröffentlicht: (2025)
Beyond Completion: A Foundation Model for General Knowledge Graph Reasoning
von: Hua, Yin, et al.
Veröffentlicht: (2025)
von: Hua, Yin, et al.
Veröffentlicht: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
Multi-domain Knowledge Graph Collaborative Pre-training and Prompt Tuning for Diverse Downstream Tasks
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
von: Hua, Tianyu, et al.
Veröffentlicht: (2025)
von: Hua, Tianyu, et al.
Veröffentlicht: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
Self-Correction Distillation for Structured Data Question Answering
von: Zhu, Yushan, et al.
Veröffentlicht: (2025)
von: Zhu, Yushan, et al.
Veröffentlicht: (2025)
MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
von: Bai, Ge, et al.
Veröffentlicht: (2024)
von: Bai, Ge, et al.
Veröffentlicht: (2024)
TrustUQA: A Trustful Framework for Unified Structured Data Question Answering
von: Zhang, Wen, et al.
Veröffentlicht: (2024)
von: Zhang, Wen, et al.
Veröffentlicht: (2024)
MoralBench: Moral Evaluation of LLMs
von: Ji, Jianchao, et al.
Veröffentlicht: (2024)
von: Ji, Jianchao, et al.
Veröffentlicht: (2024)
MAQInstruct: Instruction-based Unified Event Relation Extraction
von: Xu, Jun, et al.
Veröffentlicht: (2025)
von: Xu, Jun, et al.
Veröffentlicht: (2025)
ChatUIE: Exploring Chat-based Unified Information Extraction using Large Language Models
von: Xu, Jun, et al.
Veröffentlicht: (2024)
von: Xu, Jun, et al.
Veröffentlicht: (2024)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
von: Zhu, Hengchuan, et al.
Veröffentlicht: (2025)
von: Zhu, Hengchuan, et al.
Veröffentlicht: (2025)
Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning
von: Li, Songze, et al.
Veröffentlicht: (2025)
von: Li, Songze, et al.
Veröffentlicht: (2025)
IEPile: Unearthing Large-Scale Schema-Based Information Extraction Corpus
von: Gui, Honghao, et al.
Veröffentlicht: (2024)
von: Gui, Honghao, et al.
Veröffentlicht: (2024)
SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models
von: Lai, Peichao, et al.
Veröffentlicht: (2025)
von: Lai, Peichao, et al.
Veröffentlicht: (2025)
MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
von: Niu, Yadong, et al.
Veröffentlicht: (2025)
von: Niu, Yadong, et al.
Veröffentlicht: (2025)
OneKE: A Dockerized Schema-Guided LLM Agent-based Knowledge Extraction System
von: Luo, Yujie, et al.
Veröffentlicht: (2024)
von: Luo, Yujie, et al.
Veröffentlicht: (2024)
Large Knowledge Model: Perspectives and Challenges
von: Chen, Huajun
Veröffentlicht: (2023)
von: Chen, Huajun
Veröffentlicht: (2023)
Unleashing the Power of Imbalanced Modality Information for Multi-modal Knowledge Graph Completion
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
Efficient Knowledge Infusion via KG-LLM Alignment
von: Jiang, Zhouyu, et al.
Veröffentlicht: (2024)
von: Jiang, Zhouyu, et al.
Veröffentlicht: (2024)
RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
von: Chen, Yelin, et al.
Veröffentlicht: (2026)
von: Chen, Yelin, et al.
Veröffentlicht: (2026)
OneEdit: A Neural-Symbolic Collaboratively Knowledge Editing System
von: Zhang, Ningyu, et al.
Veröffentlicht: (2024)
von: Zhang, Ningyu, et al.
Veröffentlicht: (2024)
KAG-Thinker: Interactive Thinking and Deep Reasoning in LLMs via Knowledge-Augmented Generation
von: Zhang, Dalong, et al.
Veröffentlicht: (2025)
von: Zhang, Dalong, et al.
Veröffentlicht: (2025)
AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans
von: Xie, Wei, et al.
Veröffentlicht: (2025)
von: Xie, Wei, et al.
Veröffentlicht: (2025)
FineMath: A Fine-Grained Mathematical Evaluation Benchmark for Chinese Large Language Models
von: Liu, Yan, et al.
Veröffentlicht: (2024)
von: Liu, Yan, et al.
Veröffentlicht: (2024)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
K-ON: Stacking Knowledge On the Head Layer of Large Language Model
von: Guo, Lingbing, et al.
Veröffentlicht: (2025) -
Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis
von: Yuan, Lin, et al.
Veröffentlicht: (2025) -
Retrieve, Summarize, Plan: Advancing Multi-hop Question Answering with an Iterative Approach
von: Jiang, Zhouyu, et al.
Veröffentlicht: (2024) -
Have We Designed Generalizable Structural Knowledge Promptings? Systematic Evaluation and Rethinking
von: Zhang, Yichi, et al.
Veröffentlicht: (2024) -
OneGen: Efficient One-Pass Unified Generation and Retrieval for LLMs
von: Zhang, Jintian, et al.
Veröffentlicht: (2024)