GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?
Fuente:
arXiv
Salvato in:
| Autori principali: | Lei, Zhikai, Liang, Tianyi, Hu, Hanglei, Zhang, Jin, Zhou, Yunhua, Shao, Yunfan, Li, Linyang, Li, Chenchui, Wang, Changbo, Yan, Hang, Guo, Qipeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
di: Zhang, Xiaotian, et al.
Pubblicazione: (2023)
di: Zhang, Xiaotian, et al.
Pubblicazione: (2023)
Unearthing Large Scale Domain-Specific Knowledge from Public Corpora
di: Fei, Zhaoye, et al.
Pubblicazione: (2024)
di: Fei, Zhaoye, et al.
Pubblicazione: (2024)
UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance
di: Ma, Yichuan, et al.
Pubblicazione: (2025)
di: Ma, Yichuan, et al.
Pubblicazione: (2025)
FastMCTS: A Simple Sampling Strategy for Data Synthesis
di: Li, Peiji, et al.
Pubblicazione: (2025)
di: Li, Peiji, et al.
Pubblicazione: (2025)
Balanced Data Sampling for Language Model Training with Clustering
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
Does Porphyromonas gingivalis truly inhibit the oral carcinogenesis?
di: Chen‐xi Li, et al.
Pubblicazione: (2025)
di: Chen‐xi Li, et al.
Pubblicazione: (2025)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2024)
Does insect trapping truly measure insect populations?
di: Luca Rossini, et al.
Pubblicazione: (2025)
di: Luca Rossini, et al.
Pubblicazione: (2025)
Case2Code: Scalable Synthetic Data for Code Generation
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
di: Shao, Yunfan, et al.
Pubblicazione: (2024)
Interactive visual query of density maps on latent space via flow‐based models
di: Ning Li, et al.
Pubblicazione: (2024)
di: Ning Li, et al.
Pubblicazione: (2024)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
di: Wang, Bo, et al.
Pubblicazione: (2025)
di: Wang, Bo, et al.
Pubblicazione: (2025)
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
Breve história do ensino superior brasileiro e da formação de professores para a escola secundária
di: Núria Hanglei Cacete
Pubblicazione: (2014)
di: Núria Hanglei Cacete
Pubblicazione: (2014)
A EVOLUÇÃO DO ENSINO SUPERIOR BRASILEIRO E A FORMAÇÃO DE PROFESSORES DE GEOGRAFIA
di: Núria Hanglei Cacete
Pubblicazione: (2011)
di: Núria Hanglei Cacete
Pubblicazione: (2011)
Unified Active Retrieval for Retrieval Augmented Generation
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation
di: Zong, Yi, et al.
Pubblicazione: (2024)
di: Zong, Yi, et al.
Pubblicazione: (2024)
Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation Learning
di: Hu, Jiayun, et al.
Pubblicazione: (2025)
di: Hu, Jiayun, et al.
Pubblicazione: (2025)
MPJudge: Towards Perceptual Assessment of Music-Induced Paintings
di: Jiang, Shiqi, et al.
Pubblicazione: (2025)
di: Jiang, Shiqi, et al.
Pubblicazione: (2025)
F-Eval: Assessing Fundamental Abilities with Refined Evaluation Methods
di: Sun, Yu, et al.
Pubblicazione: (2024)
di: Sun, Yu, et al.
Pubblicazione: (2024)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
di: Xie, Yuxuan, et al.
Pubblicazione: (2024)
di: Xie, Yuxuan, et al.
Pubblicazione: (2024)
Combined grey prediction fuzzy control law with application to road tunnel ventilation system
di: Li Yunhua
Pubblicazione: (2015)
di: Li Yunhua
Pubblicazione: (2015)
Data-free Weight Compress and Denoise for Large Language Models
di: Peng, Runyu, et al.
Pubblicazione: (2024)
di: Peng, Runyu, et al.
Pubblicazione: (2024)
Code Needs Comments: Enhancing Code LLMs with Comment Augmentation
di: Song, Demin, et al.
Pubblicazione: (2024)
di: Song, Demin, et al.
Pubblicazione: (2024)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
di: Peng, Runyu, et al.
Pubblicazione: (2026)
di: Peng, Runyu, et al.
Pubblicazione: (2026)
Analysis of instruction-based LLMs' capabilities to score and judge text-input problems in an academic setting
di: Ramirez-Garcia, Valeria, et al.
Pubblicazione: (2025)
di: Ramirez-Garcia, Valeria, et al.
Pubblicazione: (2025)
Does the survival and sudden death of quadripartite steering in curved spacetime truly depend on multi-directionality?
di: Liu, Xiaobao, et al.
Pubblicazione: (2025)
di: Liu, Xiaobao, et al.
Pubblicazione: (2025)
How to Set the Learning Rate for Large-Scale Pre-training?
di: Zhou, Yunhua, et al.
Pubblicazione: (2026)
di: Zhou, Yunhua, et al.
Pubblicazione: (2026)
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
di: Saon, George, et al.
Pubblicazione: (2025)
di: Saon, George, et al.
Pubblicazione: (2025)
Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
di: Ma, Yichuan, et al.
Pubblicazione: (2026)
di: Ma, Yichuan, et al.
Pubblicazione: (2026)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
di: Shu, Lei, et al.
Pubblicazione: (2023)
di: Shu, Lei, et al.
Pubblicazione: (2023)
Gain on ground state of quantum system for truly $\mathcal{PT}$ symmetry
di: Liu, Bing-Bing, et al.
Pubblicazione: (2025)
di: Liu, Bing-Bing, et al.
Pubblicazione: (2025)
AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters
di: Luo, Hanjun, et al.
Pubblicazione: (2026)
di: Luo, Hanjun, et al.
Pubblicazione: (2026)
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
di: Cheng, Zhili, et al.
Pubblicazione: (2025)
di: Cheng, Zhili, et al.
Pubblicazione: (2025)
Change-Point Detection for Object-valued Time Series
di: Zhang, Yi, et al.
Pubblicazione: (2026)
di: Zhang, Yi, et al.
Pubblicazione: (2026)
Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of Go
di: Ma, Yichuan, et al.
Pubblicazione: (2026)
di: Ma, Yichuan, et al.
Pubblicazione: (2026)
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian
di: Syromiatnikov, Mykyta, et al.
Pubblicazione: (2025)
di: Syromiatnikov, Mykyta, et al.
Pubblicazione: (2025)
Evolutionary turnover of sRNA target interaction in Enterobacteriaceae
di: Yunfan, Jin
Pubblicazione: (2026)
di: Yunfan, Jin
Pubblicazione: (2026)
TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation
di: Liang, Tianyi, et al.
Pubblicazione: (2024)
di: Liang, Tianyi, et al.
Pubblicazione: (2024)
KG4RecEval: Does Knowledge Graph Really Matter for Recommender Systems?
di: Zhang, Haonan, et al.
Pubblicazione: (2024)
di: Zhang, Haonan, et al.
Pubblicazione: (2024)
Are UX evaluation methods truly accessible
di: Fuentes-Cortázar, Andrés Eduardo, et al.
Pubblicazione: (2025)
di: Fuentes-Cortázar, Andrés Eduardo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
di: Zhang, Xiaotian, et al.
Pubblicazione: (2023) -
Unearthing Large Scale Domain-Specific Knowledge from Public Corpora
di: Fei, Zhaoye, et al.
Pubblicazione: (2024) -
UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance
di: Ma, Yichuan, et al.
Pubblicazione: (2025) -
FastMCTS: A Simple Sampling Strategy for Data Synthesis
di: Li, Peiji, et al.
Pubblicazione: (2025) -
Balanced Data Sampling for Language Model Training with Clustering
di: Shao, Yunfan, et al.
Pubblicazione: (2024)