CJEval: A Benchmark for Assessing Large Language Models Using Chinese Junior High School Exam Data
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Qian-Wen, Wang, Haochen, Li, Fang, An, Siyu, Qiao, Lingfeng, Gao, Liangcai, Yin, Di, Sun, Xing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SNFinLLM: Systematic and Nuanced Financial Domain Adaptation of Chinese Large Language Models
by: Zhao, Shujuan, et al.
Published: (2024)
by: Zhao, Shujuan, et al.
Published: (2024)
FactGuard: Leveraging Multi-Agent Systems to Generate Answerable and Unanswerable Questions for Enhanced Long-Context LLM Extraction
by: Zhang, Qian-Wen, et al.
Published: (2025)
by: Zhang, Qian-Wen, et al.
Published: (2025)
DocTabQA: Answering Questions from Long Documents Using Tables
by: Wang, Haochen, et al.
Published: (2024)
by: Wang, Haochen, et al.
Published: (2024)
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long Contexts
by: Yu, Yifei, et al.
Published: (2025)
by: Yu, Yifei, et al.
Published: (2025)
EnviroExam: Benchmarking Environmental Science Knowledge of Large Language Models
by: Huang, Yu, et al.
Published: (2024)
by: Huang, Yu, et al.
Published: (2024)
CPsyExam: A Chinese Benchmark for Evaluating Psychology using Examinations
by: Zhao, Jiahao, et al.
Published: (2024)
by: Zhao, Jiahao, et al.
Published: (2024)
Assessing LLMs' Performance: Insights from the Chinese Pharmacist Exam
by: Wang, Xinran, et al.
Published: (2025)
by: Wang, Xinran, et al.
Published: (2025)
Magazines for Junior and Senior High Schools.
by: Waltzer, Margaret Allen
Published: (1987)
by: Waltzer, Margaret Allen
Published: (1987)
LAiW: A Chinese Legal Large Language Models Benchmark
by: Dai, Yongfu, et al.
Published: (2023)
by: Dai, Yongfu, et al.
Published: (2023)
MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models
by: Peng, Shuai, et al.
Published: (2024)
by: Peng, Shuai, et al.
Published: (2024)
CONGRA: Benchmarking Automatic Conflict Resolution
by: Zhang, Qingyu, et al.
Published: (2024)
by: Zhang, Qingyu, et al.
Published: (2024)
Junior High Curriculum Guide to Language Arts of the Mehlville School District.
Published: (1983)
Published: (1983)
Mathematics Library, Elementary and Junior High School.
by: Hardgrove, Clarence Ethel, et al.
Published: (1968)
by: Hardgrove, Clarence Ethel, et al.
Published: (1968)
Reading Interests of Junior High School Students.
by: Goetze, Henry J.
Published: (1972)
by: Goetze, Henry J.
Published: (1972)
Indian Literature for Junior and Senior High Schools.
by: Buck, June M.
Published: (1968)
by: Buck, June M.
Published: (1968)
The Junior High School: A Survey of Grades 7-8-9 in Junior and Junior-Senior High Schools, 1959-60. Bulletin, 1963, No. 32. OE-20046
by: Wright, Grace S., et al.
Published: (1963)
by: Wright, Grace S., et al.
Published: (1963)
Summer Program for Junior High School and Intermediate School Pupils.
by: Fox, David J., et al.
Published: (1969)
by: Fox, David J., et al.
Published: (1969)
Manual for New School Construction, Junior and Senior High Schools.
by: Van Hoose, Richard
Published: (1964)
by: Van Hoose, Richard
Published: (1964)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
by: Fang, Meng, et al.
Published: (2024)
by: Fang, Meng, et al.
Published: (2024)
GenExam: A Multidisciplinary Text-to-Image Exam
by: Wang, Zhaokai, et al.
Published: (2025)
by: Wang, Zhaokai, et al.
Published: (2025)
SegHist: A General Segmentation-based Framework for Chinese Historical Document Text Line Detection
by: Hu, Xingjian, et al.
Published: (2024)
by: Hu, Xingjian, et al.
Published: (2024)
Data Visualization of the Brazilian National High School Exam: VisDadosEnem
by: Robson Rodrigues LEMOS
Published: (2018)
by: Robson Rodrigues LEMOS
Published: (2018)
School Nurses’ Experiences of Suicide Prevention Work in Junior High School
by: Rikard Wärdig, et al.
Published: (2026)
by: Rikard Wärdig, et al.
Published: (2026)
Using Oral Exams in Physics and Astronomy Courses
by: Zanger, Brian DiGiorgio
Published: (2025)
by: Zanger, Brian DiGiorgio
Published: (2025)
ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models
by: Zhang, Hengxiang, et al.
Published: (2024)
by: Zhang, Hengxiang, et al.
Published: (2024)
Let's Be Self-generated via Step by Step: A Curriculum Learning Approach to Automated Reasoning with Large Language Models
by: Luo, Kangyang, et al.
Published: (2024)
by: Luo, Kangyang, et al.
Published: (2024)
AlignBench: Benchmarking Chinese Alignment of Large Language Models
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
FinVerse: An Autonomous Agent System for Versatile Financial Analysis
by: An, Siyu, et al.
Published: (2024)
by: An, Siyu, et al.
Published: (2024)
Native Library Resources for Elementary, Junior and Senior High Schools.
Published: (1987)
Published: (1987)
Junior High School Study Skills Program. First Draft.
Published: (1980)
Published: (1980)
Study of Reading In Indiana Middle, Junior, and Senior High Schools
by: Holland, Earlene L., et al.
Published: (2004)
by: Holland, Earlene L., et al.
Published: (2004)
Mathematics Library, Elementary and Junior High School. Fourth Edition.
by: Wheeler, Margariete Montague, et al.
Published: (1978)
by: Wheeler, Margariete Montague, et al.
Published: (1978)
Mathematics Library, Elementary and Junior High School. Second Edition.
by: Hardgrove, Clarence Ethel, et al.
Published: (1973)
by: Hardgrove, Clarence Ethel, et al.
Published: (1973)
Mathematics Library. Elementary and Junior High School. Fifth Edition.
by: Wheeler, Margariete Montague
Published: (1986)
by: Wheeler, Margariete Montague
Published: (1986)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
by: Dinh, Tu Anh, et al.
Published: (2024)
by: Dinh, Tu Anh, et al.
Published: (2024)
Size Variation in Flower Petals of Chinese Animal‐Pollinated Plants in Response to Climatic and Altitudinal Gradients
by: Siyu Chen, et al.
Published: (2025)
by: Siyu Chen, et al.
Published: (2025)
Evaluating Large‐Scale and Lightweight Large Language Models for Traditional Chinese Medicine Exam Questions: A Comparative Study
by: Yizhen Li, et al.
Published: (2026)
by: Yizhen Li, et al.
Published: (2026)
Benchmarking quantized LLaMa-based models on the Brazilian Secondary School Exam
by: Santos, Matheus L. O., et al.
Published: (2023)
by: Santos, Matheus L. O., et al.
Published: (2023)
Benchmarking Robustness of Endoscopic Depth Estimation with Synthetically Corrupted Data
by: Wang, An, et al.
Published: (2024)
by: Wang, An, et al.
Published: (2024)
Similar Items
-
SNFinLLM: Systematic and Nuanced Financial Domain Adaptation of Chinese Large Language Models
by: Zhao, Shujuan, et al.
Published: (2024) -
FactGuard: Leveraging Multi-Agent Systems to Generate Answerable and Unanswerable Questions for Enhanced Long-Context LLM Extraction
by: Zhang, Qian-Wen, et al.
Published: (2025) -
DocTabQA: Answering Questions from Long Documents Using Tables
by: Wang, Haochen, et al.
Published: (2024) -
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
by: Wang, Haochen, et al.
Published: (2025) -
Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long Contexts
by: Yu, Yifei, et al.
Published: (2025)