KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Kaijing, Du, Xinrun, Wang, Yunran, Zhang, Haoran, Wen, Zhoufutu, Qu, Xingwei, Yang, Jian, Liu, Jiaheng, Liu, Minghao, Yue, Xiang, Huang, Wenhao, Zhang, Ge |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
V3DB: Audit-on-Demand Zero-Knowledge Proofs for Verifiable Vector Search over Committed Snapshots
by: Qiu, Zipeng, et al.
Published: (2026)
by: Qiu, Zipeng, et al.
Published: (2026)
OxyEcomBench: Benchmarking Multimodal Foundation Models across E-Commerce Ecosystems
by: Liu, Yong, et al.
Published: (2026)
by: Liu, Yong, et al.
Published: (2026)
DriftBench: Defining and Generating Data and Query Workload Drift for Benchmarking
by: Liu, Guanli, et al.
Published: (2025)
by: Liu, Guanli, et al.
Published: (2025)
DW-Bench: Benchmarking LLMs on Data Warehouse Graph Topology Reasoning
by: Ahmed, Ahmed G. A. H, et al.
Published: (2026)
by: Ahmed, Ahmed G. A. H, et al.
Published: (2026)
A Categorical Unification for Multi-Model Data: Part II Categorical Algebra and Calculus
by: Lu, Jiaheng
Published: (2025)
by: Lu, Jiaheng
Published: (2025)
A Categorical Unification for Multi-Model Data: Part I Categorical Model and Normal Forms
by: Lu, Jiaheng
Published: (2025)
by: Lu, Jiaheng
Published: (2025)
NeurBench: A Benchmark Suite for Learned Database Components with Drift Modeling
by: Zhao, Zhanhao, et al.
Published: (2025)
by: Zhao, Zhanhao, et al.
Published: (2025)
MMTS-BENCH: A Comprehensive Benchmark for Time Series Understanding and Reasoning
by: Yin, Yao, et al.
Published: (2026)
by: Yin, Yao, et al.
Published: (2026)
LST-Bench: Benchmarking Log-Structured Tables in the Cloud
by: Camacho-Rodríguez, Jesús, et al.
Published: (2023)
by: Camacho-Rodríguez, Jesús, et al.
Published: (2023)
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
by: Tang, Zirui, et al.
Published: (2026)
by: Tang, Zirui, et al.
Published: (2026)
PGB: Benchmarking Differentially Private Synthetic Graph Generation Algorithms
by: Liu, Shang, et al.
Published: (2024)
by: Liu, Shang, et al.
Published: (2024)
UniDataBench: Evaluating Data Analytics Agents Across Structured and Unstructured Data
by: Weng, Han, et al.
Published: (2025)
by: Weng, Han, et al.
Published: (2025)
PandasBench: A Benchmark for the Pandas API
by: Broihier, Alex, et al.
Published: (2025)
by: Broihier, Alex, et al.
Published: (2025)
KARPA: A Training-free Method of Adapting Knowledge Graph as References for Large Language Model's Reasoning Path Aggregation
by: Fang, Siyuan, et al.
Published: (2024)
by: Fang, Siyuan, et al.
Published: (2024)
MMTU: A Massive Multi-Task Table Understanding and Reasoning Benchmark
by: Xing, Junjie, et al.
Published: (2025)
by: Xing, Junjie, et al.
Published: (2025)
Towards Temporal Knowledge Graph Alignment in the Wild
by: Zhao, Runhao, et al.
Published: (2025)
by: Zhao, Runhao, et al.
Published: (2025)
Conflict Detection for Temporal Knowledge Graphs:A Fast Constraint Mining Algorithm and New Benchmarks
by: Chen, Jianhao, et al.
Published: (2023)
by: Chen, Jianhao, et al.
Published: (2023)
SemBench: A Benchmark for Semantic Query Processing Engines
by: Lao, Jiale, et al.
Published: (2025)
by: Lao, Jiale, et al.
Published: (2025)
RelBench: A Benchmark for Deep Learning on Relational Databases
by: Robinson, Joshua, et al.
Published: (2024)
by: Robinson, Joshua, et al.
Published: (2024)
EvoRAG: Making Knowledge Graph-based RAG Automatically Evolve through Feedback-driven Backpropagation
by: Fu, Zhenbo, et al.
Published: (2026)
by: Fu, Zhenbo, et al.
Published: (2026)
EpiCastBench: Datasets and Benchmarks for Multivariate Epidemic Forecasting
by: Panja, Madhurima, et al.
Published: (2026)
by: Panja, Madhurima, et al.
Published: (2026)
KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree Search
by: Luo, Haoran, et al.
Published: (2025)
by: Luo, Haoran, et al.
Published: (2025)
Benchmarking Large Language Models for Knowledge Graph Validation
by: Shami, Farzad, et al.
Published: (2026)
by: Shami, Farzad, et al.
Published: (2026)
DP-Bench: A Benchmark for Evaluating Data Product Creation Systems
by: Chowdhury, Faisal, et al.
Published: (2025)
by: Chowdhury, Faisal, et al.
Published: (2025)
CardBench: A Benchmark for Learned Cardinality Estimation in Relational Databases
by: Chronis, Yannis, et al.
Published: (2024)
by: Chronis, Yannis, et al.
Published: (2024)
Benchmarking Time Series Databases with IoTDB-Benchmark for IoT Scenarios
by: Liu, Rui, et al.
Published: (2019)
by: Liu, Rui, et al.
Published: (2019)
Quantum Information-Theoretical Size Bounds for Conjunctive Queries with Functional Dependencies
by: Uotila, Valter, et al.
Published: (2025)
by: Uotila, Valter, et al.
Published: (2025)
PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking
by: Zhou, Yan, et al.
Published: (2025)
by: Zhou, Yan, et al.
Published: (2025)
ResBench: A Comprehensive Framework for Evaluating Database Resilience
by: Hu, Puyun, et al.
Published: (2025)
by: Hu, Puyun, et al.
Published: (2025)
ELT-Bench-Verified: Benchmark Quality Issues Underestimate AI Agent Capabilities
by: Zanoli, Christopher, et al.
Published: (2026)
by: Zanoli, Christopher, et al.
Published: (2026)
ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
by: Jin, Tengjun, et al.
Published: (2025)
by: Jin, Tengjun, et al.
Published: (2025)
From Stimuli to Minds: Enhancing Psychological Reasoning in LLMs via Bilateral Reinforcement Learning
by: Feng, Yichao, et al.
Published: (2025)
by: Feng, Yichao, et al.
Published: (2025)
DBMS-LLM Integration Strategies in Industrial and Business Applications: Current Status and Future Challenges
by: Yan, Zhengtong, et al.
Published: (2025)
by: Yan, Zhengtong, et al.
Published: (2025)
Outback: Fast and Communication-efficient Index for Key-Value Store on Disaggregated Memory
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Zero-Knowledge Verifiable Graph Query Evaluation via Expansion-Centric Operator Decomposition
by: Wu, Hao, et al.
Published: (2025)
by: Wu, Hao, et al.
Published: (2025)
OpenGLT: A Comprehensive Benchmark of Graph Neural Networks for Graph-Level Tasks
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and Solution
by: Wei, Zixin, et al.
Published: (2025)
by: Wei, Zixin, et al.
Published: (2025)
DTBench: A Synthetic Benchmark for Document-to-Table Extraction
by: Guo, Yuxiang, et al.
Published: (2026)
by: Guo, Yuxiang, et al.
Published: (2026)
Categorical Calculus and Algebra for Multi-Model Data
by: Lu, Jiaheng
Published: (2026)
by: Lu, Jiaheng
Published: (2026)
Online Marketplace: A Benchmark for Data Management in Microservices
by: Laigner, Rodrigo, et al.
Published: (2024)
by: Laigner, Rodrigo, et al.
Published: (2024)
Similar Items
-
V3DB: Audit-on-Demand Zero-Knowledge Proofs for Verifiable Vector Search over Committed Snapshots
by: Qiu, Zipeng, et al.
Published: (2026) -
OxyEcomBench: Benchmarking Multimodal Foundation Models across E-Commerce Ecosystems
by: Liu, Yong, et al.
Published: (2026) -
DriftBench: Defining and Generating Data and Query Workload Drift for Benchmarking
by: Liu, Guanli, et al.
Published: (2025) -
DW-Bench: Benchmarking LLMs on Data Warehouse Graph Topology Reasoning
by: Ahmed, Ahmed G. A. H, et al.
Published: (2026) -
A Categorical Unification for Multi-Model Data: Part II Categorical Algebra and Calculus
by: Lu, Jiaheng
Published: (2025)