Characterizing Knowledge Graph Tasks in LLM Benchmarks Using Cognitive Complexity Frameworks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Todorovikj, Sara, Meyer, Lars-Peter, Martin, Michael |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LLM-assisted Knowledge Graph Engineering: Experiments with ChatGPT
par: Meyer, Lars-Peter, et autres
Publié: (2023)
par: Meyer, Lars-Peter, et autres
Publié: (2023)
How do Scaling Laws Apply to Knowledge Graph Engineering Tasks? The Impact of Model Size on Large Language Model Performance
par: Heim, Desiree, et autres
Publié: (2025)
par: Heim, Desiree, et autres
Publié: (2025)
A Fine-Tuning Approach for T5 Using Knowledge Graphs to Address Complex Tasks
par: Liao, Xiaoxuan, et autres
Publié: (2025)
par: Liao, Xiaoxuan, et autres
Publié: (2025)
ARUQULA -- An LLM based Text2SPARQL Approach using ReAct and Knowledge Graph Exploration Utilities
par: Brei, Felix, et autres
Publié: (2025)
par: Brei, Felix, et autres
Publié: (2025)
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
par: Wang, Zhenting, et autres
Publié: (2025)
par: Wang, Zhenting, et autres
Publié: (2025)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
par: Song, Yuanyi, et autres
Publié: (2025)
par: Song, Yuanyi, et autres
Publié: (2025)
ARK-V1: An LLM-Agent for Knowledge Graph Question Answering Requiring Commonsense Reasoning
par: Klein, Jan-Felix, et autres
Publié: (2025)
par: Klein, Jan-Felix, et autres
Publié: (2025)
The Semantic Ladder: A Framework for Progressive Formalization of Natural Language Content for Knowledge Graphs and AI Systems
par: Vogt, Lars
Publié: (2026)
par: Vogt, Lars
Publié: (2026)
KGHaluBench: A Knowledge Graph-Based Hallucination Benchmark for Evaluating the Breadth and Depth of LLM Knowledge
par: Robertson, Alex, et autres
Publié: (2026)
par: Robertson, Alex, et autres
Publié: (2026)
Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge Graphs
par: Hu, Nan, et autres
Publié: (2024)
par: Hu, Nan, et autres
Publié: (2024)
Spider4SPARQL: A Complex Benchmark for Evaluating Knowledge Graph Question Answering Systems
par: Kosten, Catherine, et autres
Publié: (2023)
par: Kosten, Catherine, et autres
Publié: (2023)
Temporal Dynamics of Emotion and Cognition in Human Translation: Integrating the Task Segment Framework and the HOF Taxonomy
par: Carl, Michael
Publié: (2024)
par: Carl, Michael
Publié: (2024)
ROG: Retrieval-Augmented LLM Reasoning for Complex First-Order Queries over Knowledge Graphs
par: Zhang, Ziyan, et autres
Publié: (2026)
par: Zhang, Ziyan, et autres
Publié: (2026)
KG-Agent: An Efficient Autonomous Agent Framework for Complex Reasoning over Knowledge Graph
par: Jiang, Jinhao, et autres
Publié: (2024)
par: Jiang, Jinhao, et autres
Publié: (2024)
LegalRikai: Open Benchmark -- Benchmark for Complex Japanese Corporate Legal Tasks
par: Fujita, Shogo, et autres
Publié: (2025)
par: Fujita, Shogo, et autres
Publié: (2025)
Can LLM Reasoning Be Trusted? A Comparative Study: Using Human Benchmarking on Statistical Tasks
par: Nagarkar, Crish, et autres
Publié: (2026)
par: Nagarkar, Crish, et autres
Publié: (2026)
Grounding LLM Reasoning with Knowledge Graphs
par: Amayuelas, Alfonso, et autres
Publié: (2025)
par: Amayuelas, Alfonso, et autres
Publié: (2025)
Dual Reasoning: A GNN-LLM Collaborative Framework for Knowledge Graph Question Answering
par: Liu, Guangyi, et autres
Publié: (2024)
par: Liu, Guangyi, et autres
Publié: (2024)
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests
par: Momentè, Filippo, et autres
Publié: (2025)
par: Momentè, Filippo, et autres
Publié: (2025)
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
par: Kranti, Chalamalasetti, et autres
Publié: (2025)
par: Kranti, Chalamalasetti, et autres
Publié: (2025)
A Training-free LLM Framework with Interaction between Contextually Related Subtasks in Solving Complex Tasks
par: Liu, Hongjia, et autres
Publié: (2025)
par: Liu, Hongjia, et autres
Publié: (2025)
TaskComplexity: A Dataset for Task Complexity Classification with In-Context Learning, FLAN-T5 and GPT-4o Benchmarks
par: Rasheed, Areeg Fahad, et autres
Publié: (2024)
par: Rasheed, Areeg Fahad, et autres
Publié: (2024)
EMERGE: A Benchmark for Updating Knowledge Graphs with Emerging Textual Knowledge
par: Zaporojets, Klim, et autres
Publié: (2025)
par: Zaporojets, Klim, et autres
Publié: (2025)
Evaluating Cultural Knowledge Processing in Large Language Models: A Cognitive Benchmarking Framework Integrating Retrieval-Augmented Generation
par: Lee, Hung-Shin, et autres
Publié: (2025)
par: Lee, Hung-Shin, et autres
Publié: (2025)
LLM-Guided Knowledge Distillation for Temporal Knowledge Graph Reasoning
par: Xing, Wang, et autres
Publié: (2026)
par: Xing, Wang, et autres
Publié: (2026)
COKE: A Cognitive Knowledge Graph for Machine Theory of Mind
par: Wu, Jincenzi, et autres
Publié: (2023)
par: Wu, Jincenzi, et autres
Publié: (2023)
LLM-KG-Bench 3.0: A Compass for SemanticTechnology Capabilities in the Ocean of LLMs
par: Meyer, Lars-Peter, et autres
Publié: (2025)
par: Meyer, Lars-Peter, et autres
Publié: (2025)
An Senegalese Legal Texts Structuration Using LLM-augmented Knowledge Graph
par: Kane, Oumar, et autres
Publié: (2025)
par: Kane, Oumar, et autres
Publié: (2025)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
par: Long, Xiang, et autres
Publié: (2026)
par: Long, Xiang, et autres
Publié: (2026)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
par: Markowitz, Elan, et autres
Publié: (2025)
par: Markowitz, Elan, et autres
Publié: (2025)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
par: Sansford, Hannah, et autres
Publié: (2024)
par: Sansford, Hannah, et autres
Publié: (2024)
LLM for Complex Reasoning Task: An Exploratory Study in Fermi Problems
par: Liu, Zishuo, et autres
Publié: (2025)
par: Liu, Zishuo, et autres
Publié: (2025)
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
par: Xu, Frank F., et autres
Publié: (2024)
par: Xu, Frank F., et autres
Publié: (2024)
Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching
par: Li, Songze, et autres
Publié: (2025)
par: Li, Songze, et autres
Publié: (2025)
OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
par: Chen, Yongrui, et autres
Publié: (2025)
par: Chen, Yongrui, et autres
Publié: (2025)
AGENTiGraph: A Multi-Agent Knowledge Graph Framework for Interactive, Domain-Specific LLM Chatbots
par: Zhao, Xinjie, et autres
Publié: (2025)
par: Zhao, Xinjie, et autres
Publié: (2025)
A LLM Benchmark based on the Minecraft Builder Dialog Agent Task
par: Madge, Chris, et autres
Publié: (2024)
par: Madge, Chris, et autres
Publié: (2024)
Characterizing Large Language Models as Rationalizers of Knowledge-intensive Tasks
par: Mishra, Aditi, et autres
Publié: (2023)
par: Mishra, Aditi, et autres
Publié: (2023)
TOD-ProcBench: Benchmarking Complex Instruction-Following in Task-Oriented Dialogues
par: Ghazarian, Sarik, et autres
Publié: (2025)
par: Ghazarian, Sarik, et autres
Publié: (2025)
Hierarchical Deconstruction of LLM Reasoning: A Graph-Based Framework for Analyzing Knowledge Utilization
par: Ko, Miyoung, et autres
Publié: (2024)
par: Ko, Miyoung, et autres
Publié: (2024)
Documents similaires
-
LLM-assisted Knowledge Graph Engineering: Experiments with ChatGPT
par: Meyer, Lars-Peter, et autres
Publié: (2023) -
How do Scaling Laws Apply to Knowledge Graph Engineering Tasks? The Impact of Model Size on Large Language Model Performance
par: Heim, Desiree, et autres
Publié: (2025) -
A Fine-Tuning Approach for T5 Using Knowledge Graphs to Address Complex Tasks
par: Liao, Xiaoxuan, et autres
Publié: (2025) -
ARUQULA -- An LLM based Text2SPARQL Approach using ReAct and Knowledge Graph Exploration Utilities
par: Brei, Felix, et autres
Publié: (2025) -
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
par: Wang, Zhenting, et autres
Publié: (2025)