FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Wei, Ma, Ren, Wu, Jiang, Gu, Chenya, Peng, Jiahui, Len, Jinyang, Zhang, Songyang, Yan, Hang, Lin, Dahua, He, Conghui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
von: Sun, Jiaxing, et al.
Veröffentlicht: (2024)
von: Sun, Jiaxing, et al.
Veröffentlicht: (2024)
Evaluating Large Language Model with Knowledge Oriented Language Specific Simple Question Answering
von: Jiang, Bowen, et al.
Veröffentlicht: (2025)
von: Jiang, Bowen, et al.
Veröffentlicht: (2025)
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
von: Zhou, Baichuan, et al.
Veröffentlicht: (2024)
von: Zhou, Baichuan, et al.
Veröffentlicht: (2024)
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
von: Que, Haoran, et al.
Veröffentlicht: (2024)
von: Que, Haoran, et al.
Veröffentlicht: (2024)
PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
Unearthing Large Scale Domain-Specific Knowledge from Public Corpora
von: Fei, Zhaoye, et al.
Veröffentlicht: (2024)
von: Fei, Zhaoye, et al.
Veröffentlicht: (2024)
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
von: Chen, Zehui, et al.
Veröffentlicht: (2023)
von: Chen, Zehui, et al.
Veröffentlicht: (2023)
Who posts the advertisement: The influence of advertising authorship on in‐feed advertising effectiveness
von: Chenya Ma, et al.
Veröffentlicht: (2024)
von: Chenya Ma, et al.
Veröffentlicht: (2024)
NewsBench: A Systematic Evaluation Framework for Assessing Editorial Capabilities of Large Language Models in Chinese Journalism
von: Li, Miao, et al.
Veröffentlicht: (2024)
von: Li, Miao, et al.
Veröffentlicht: (2024)
Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training
von: Peng, Jiahui, et al.
Veröffentlicht: (2025)
von: Peng, Jiahui, et al.
Veröffentlicht: (2025)
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
von: Huang, Xu, et al.
Veröffentlicht: (2025)
von: Huang, Xu, et al.
Veröffentlicht: (2025)
CityBench: Evaluating the Capabilities of Large Language Models for Urban Tasks
von: Feng, Jie, et al.
Veröffentlicht: (2024)
von: Feng, Jie, et al.
Veröffentlicht: (2024)
Utilize the Flow before Stepping into the Same River Twice: Certainty Represented Knowledge Flow for Refusal-Aware Instruction Tuning
von: Zhu, Runchuan, et al.
Veröffentlicht: (2024)
von: Zhu, Runchuan, et al.
Veröffentlicht: (2024)
Evaluating the Generation Capabilities of Large Chinese Language Models
von: Zeng, Hui, et al.
Veröffentlicht: (2023)
von: Zeng, Hui, et al.
Veröffentlicht: (2023)
DSDL: Data Set Description Language for Bridging Modalities and Tasks in AI Data
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
DebugBench: Evaluating Debugging Capability of Large Language Models
von: Tian, Runchu, et al.
Veröffentlicht: (2024)
von: Tian, Runchu, et al.
Veröffentlicht: (2024)
TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models
von: Zhang, Quanjun, et al.
Veröffentlicht: (2024)
von: Zhang, Quanjun, et al.
Veröffentlicht: (2024)
InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language Models
von: Li, Linyi, et al.
Veröffentlicht: (2024)
von: Li, Linyi, et al.
Veröffentlicht: (2024)
Large Language Models Meet Symbolic Provers for Logical Reasoning Evaluation
von: Qi, Chengwen, et al.
Veröffentlicht: (2025)
von: Qi, Chengwen, et al.
Veröffentlicht: (2025)
From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
von: Li, Rongjie, et al.
Veröffentlicht: (2024)
von: Li, Rongjie, et al.
Veröffentlicht: (2024)
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
von: Guo, Jiawei, et al.
Veröffentlicht: (2024)
von: Guo, Jiawei, et al.
Veröffentlicht: (2024)
LongWanjuan: Towards Systematic Measurement for Long Text Quality
von: Lv, Kai, et al.
Veröffentlicht: (2024)
von: Lv, Kai, et al.
Veröffentlicht: (2024)
SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models
von: Lai, Peichao, et al.
Veröffentlicht: (2025)
von: Lai, Peichao, et al.
Veröffentlicht: (2025)
Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law
von: Ge, Qiming, et al.
Veröffentlicht: (2025)
von: Ge, Qiming, et al.
Veröffentlicht: (2025)
InternLM-Law: An Open Source Chinese Legal Large Language Model
von: Fei, Zhiwei, et al.
Veröffentlicht: (2024)
von: Fei, Zhiwei, et al.
Veröffentlicht: (2024)
OpenHuEval: Evaluating Large Language Model on Hungarian Specifics
von: Yang, Haote, et al.
Veröffentlicht: (2025)
von: Yang, Haote, et al.
Veröffentlicht: (2025)
F-Eval: Assessing Fundamental Abilities with Refined Evaluation Methods
von: Sun, Yu, et al.
Veröffentlicht: (2024)
von: Sun, Yu, et al.
Veröffentlicht: (2024)
ImagineBench: Evaluating Reinforcement Learning with Large Language Model Rollouts
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2025)
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2025)
RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues
von: Kuo, Tzu-Lin, et al.
Veröffentlicht: (2024)
von: Kuo, Tzu-Lin, et al.
Veröffentlicht: (2024)
Data-free Weight Compress and Denoise for Large Language Models
von: Peng, Runyu, et al.
Veröffentlicht: (2024)
von: Peng, Runyu, et al.
Veröffentlicht: (2024)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models
von: Zeng, Bo, et al.
Veröffentlicht: (2025)
von: Zeng, Bo, et al.
Veröffentlicht: (2025)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models
von: LI, Yizhi, et al.
Veröffentlicht: (2024)
von: LI, Yizhi, et al.
Veröffentlicht: (2024)
OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
von: Zhang, Min, et al.
Veröffentlicht: (2025)
von: Zhang, Min, et al.
Veröffentlicht: (2025)
LHMKE: A Large-scale Holistic Multi-subject Knowledge Evaluation Benchmark for Chinese Large Language Models
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
LR^2Bench: Evaluating Long-chain Reflective Reasoning Capabilities of Large Language Models via Constraint Satisfaction Problems
von: Chen, Jianghao, et al.
Veröffentlicht: (2025)
von: Chen, Jianghao, et al.
Veröffentlicht: (2025)
VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models
von: Liu, Yu, et al.
Veröffentlicht: (2024)
von: Liu, Yu, et al.
Veröffentlicht: (2024)
AlignBench: Benchmarking Chinese Alignment of Large Language Models
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
von: Sun, Jiaxing, et al.
Veröffentlicht: (2024) -
Evaluating Large Language Model with Knowledge Oriented Language Specific Simple Question Answering
von: Jiang, Bowen, et al.
Veröffentlicht: (2025) -
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
von: Zhou, Baichuan, et al.
Veröffentlicht: (2024) -
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
von: Que, Haoran, et al.
Veröffentlicht: (2024) -
PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)