ScholarBench: A Bilingual Benchmark for Abstraction, Comprehension, and Reasoning Evaluation in Academic Contexts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Noh, Dongwon, Koh, Donghyeok, Yuk, Junghun, Kim, Gyuwan, Lee, Jaeyong, Lim, Kyungtae, Park, Cheoneum |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PPA-Plan: Proactive Pitfall Avoidance for Reliable Planning in Long-Context LLM Reasoning
von: Kim, Byeongjin, et al.
Veröffentlicht: (2026)
von: Kim, Byeongjin, et al.
Veröffentlicht: (2026)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
von: Lim, Hyeonseok, et al.
Veröffentlicht: (2024)
von: Lim, Hyeonseok, et al.
Veröffentlicht: (2024)
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
von: Kim, Minsang, et al.
Veröffentlicht: (2024)
von: Kim, Minsang, et al.
Veröffentlicht: (2024)
BOK-VQA: Bilingual outside Knowledge-Based Visual Question Answering via Graph Representation Pretraining
von: Kim, Minjun, et al.
Veröffentlicht: (2024)
von: Kim, Minjun, et al.
Veröffentlicht: (2024)
RDB2G-Bench: A Comprehensive Benchmark for Automatic Graph Modeling of Relational Databases
von: Choi, Dongwon, et al.
Veröffentlicht: (2025)
von: Choi, Dongwon, et al.
Veröffentlicht: (2025)
The finitude of tamely ramified pro-$p$ extensions of number fields with cyclic $p$-class groups
von: Lee, Yoonjin, et al.
Veröffentlicht: (2024)
von: Lee, Yoonjin, et al.
Veröffentlicht: (2024)
Pedagogy-R1: Pedagogically-Aligned Reasoning Model with Balanced Educational Benchmark
von: Lee, Unggi, et al.
Veröffentlicht: (2025)
von: Lee, Unggi, et al.
Veröffentlicht: (2025)
CORRECT: Context- and Reference-Augmented Reasoning and Prompting for Fact-Checking
von: Zhang, Delvin Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Delvin Ce, et al.
Veröffentlicht: (2025)
Abductive Symbolic Solver on Abstraction and Reasoning Corpus
von: Lim, Mintaek, et al.
Veröffentlicht: (2024)
von: Lim, Mintaek, et al.
Veröffentlicht: (2024)
X-LLaVA: Optimizing Bilingual Large Vision-Language Alignment
von: Shin, Dongjae, et al.
Veröffentlicht: (2024)
von: Shin, Dongjae, et al.
Veröffentlicht: (2024)
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
von: Bai, Yushi, et al.
Veröffentlicht: (2023)
von: Bai, Yushi, et al.
Veröffentlicht: (2023)
MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG
von: Lim, Woosang, et al.
Veröffentlicht: (2025)
von: Lim, Woosang, et al.
Veröffentlicht: (2025)
On Krull-Schmidt decompositions of unit groups of number fields
von: Kumon, Asuka, et al.
Veröffentlicht: (2024)
von: Kumon, Asuka, et al.
Veröffentlicht: (2024)
On the Galois structure of units in totally real $p$-rational number fields
von: Bouazzaoui, Zakariae, et al.
Veröffentlicht: (2023)
von: Bouazzaoui, Zakariae, et al.
Veröffentlicht: (2023)
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
Motivated Reasoning and Information Aggregation
von: Acharya, Avidit, et al.
Veröffentlicht: (2025)
von: Acharya, Avidit, et al.
Veröffentlicht: (2025)
A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs
von: Jang, Hongsun, et al.
Veröffentlicht: (2025)
von: Jang, Hongsun, et al.
Veröffentlicht: (2025)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
von: Kim, Haechan, et al.
Veröffentlicht: (2026)
von: Kim, Haechan, et al.
Veröffentlicht: (2026)
KORMo: Korean Open Reasoning Model for Everyone
von: Kim, Minjun, et al.
Veröffentlicht: (2025)
von: Kim, Minjun, et al.
Veröffentlicht: (2025)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
Pedagogical Alignment for Vision-Language-Action Models: A Comprehensive Framework for Data, Architecture, and Evaluation in Education
von: Lee, Unggi, et al.
Veröffentlicht: (2026)
von: Lee, Unggi, et al.
Veröffentlicht: (2026)
Reasoning Abilities of Large Language Models: In-Depth Analysis on the Abstraction and Reasoning Corpus
von: Lee, Seungpil, et al.
Veröffentlicht: (2024)
von: Lee, Seungpil, et al.
Veröffentlicht: (2024)
KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context
von: Wang, Xiaonan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaonan, et al.
Veröffentlicht: (2024)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
von: Jeon, YoungHoon, et al.
Veröffentlicht: (2026)
von: Jeon, YoungHoon, et al.
Veröffentlicht: (2026)
Generative AI Policies under the Microscope: How CS Conferences Are Navigating the New Frontier in Scholarly Writing
von: Nahar, Mahjabin, et al.
Veröffentlicht: (2024)
von: Nahar, Mahjabin, et al.
Veröffentlicht: (2024)
Shift-Share Designs in Political Science
von: Park, Peter Kyungtae
Veröffentlicht: (2026)
von: Park, Peter Kyungtae
Veröffentlicht: (2026)
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
von: Park, Jinho, et al.
Veröffentlicht: (2026)
von: Park, Jinho, et al.
Veröffentlicht: (2026)
Eigenstructure inference for high-dimensional covariance with generalized shrinkage inverse-Wishart prior
von: Kim, Seongmin, et al.
Veröffentlicht: (2025)
von: Kim, Seongmin, et al.
Veröffentlicht: (2025)
Bayesian Analysis of Spiked Covariance Models: Correcting Eigenvalue Bias and Determining the Number of Spikes
von: Lee, Kwangmin, et al.
Veröffentlicht: (2024)
von: Lee, Kwangmin, et al.
Veröffentlicht: (2024)
Design Strategies Based on Electronic Interactions for Effective Catalysts in Lithium–Sulfur Batteries
von: Donghyeok Son, et al.
Veröffentlicht: (2025)
von: Donghyeok Son, et al.
Veröffentlicht: (2025)
Design Strategies Based on Electronic Interactions for Effective Catalysts in Lithium–Sulfur Batteries
von: Donghyeok Son, et al.
Veröffentlicht: (2025)
von: Donghyeok Son, et al.
Veröffentlicht: (2025)
MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following
von: Lee, Jaeyun, et al.
Veröffentlicht: (2026)
von: Lee, Jaeyun, et al.
Veröffentlicht: (2026)
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
von: Lee, Donggyu, et al.
Veröffentlicht: (2025)
von: Lee, Donggyu, et al.
Veröffentlicht: (2025)
ARCLE: The Abstraction and Reasoning Corpus Learning Environment for Reinforcement Learning
von: Lee, Hosung, et al.
Veröffentlicht: (2024)
von: Lee, Hosung, et al.
Veröffentlicht: (2024)
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
von: Patel, Liana, et al.
Veröffentlicht: (2025)
von: Patel, Liana, et al.
Veröffentlicht: (2025)
Enhancing Analogical Reasoning in the Abstraction and Reasoning Corpus via Model-Based RL
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
Latent-Space Mean-Field Theory for Deep BitNet-like Training: Constrained Gradient Flows with Smooth Quantization and STE Limits
von: Kim, Dongwon, et al.
Veröffentlicht: (2025)
von: Kim, Dongwon, et al.
Veröffentlicht: (2025)
Personalized Federated Learning for Gradient Alignment
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
Stochastic resonance in Schmitt trigger and its application towards weak signal detection
von: Kim, Yoonkang, et al.
Veröffentlicht: (2025)
von: Kim, Yoonkang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PPA-Plan: Proactive Pitfall Avoidance for Reliable Planning in Long-Context LLM Reasoning
von: Kim, Byeongjin, et al.
Veröffentlicht: (2026) -
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
von: Lim, Hyeonseok, et al.
Veröffentlicht: (2024) -
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
von: Chen, Ziyang, et al.
Veröffentlicht: (2026) -
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
von: Kim, Minsang, et al.
Veröffentlicht: (2024) -
BOK-VQA: Bilingual outside Knowledge-Based Visual Question Answering via Graph Representation Pretraining
von: Kim, Minjun, et al.
Veröffentlicht: (2024)