Unveiling the Competitive Dynamics: A Comparative Evaluation of American and Chinese LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Zhenhui, Li, Jiaxin, Liu, Yang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
di: Balunović, Mislav, et al.
Pubblicazione: (2025)
di: Balunović, Mislav, et al.
Pubblicazione: (2025)
CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
di: Wu, Chengwei, et al.
Pubblicazione: (2025)
di: Wu, Chengwei, et al.
Pubblicazione: (2025)
HARBOR: Exploring Persona Dynamics in Multi-Agent Competition
di: Jiang, Kenan, et al.
Pubblicazione: (2025)
di: Jiang, Kenan, et al.
Pubblicazione: (2025)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
di: Zhang, Jiaxin, et al.
Pubblicazione: (2024)
di: Zhang, Jiaxin, et al.
Pubblicazione: (2024)
Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs
di: Zeng, Liang, et al.
Pubblicazione: (2025)
di: Zeng, Liang, et al.
Pubblicazione: (2025)
MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine
di: Kong, Shufeng, et al.
Pubblicazione: (2025)
di: Kong, Shufeng, et al.
Pubblicazione: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
di: Jiang, Botian, et al.
Pubblicazione: (2024)
di: Jiang, Botian, et al.
Pubblicazione: (2024)
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
di: Wu, Xinwei, et al.
Pubblicazione: (2025)
di: Wu, Xinwei, et al.
Pubblicazione: (2025)
Encyclo-K: Evaluating LLMs with Dynamically Composed Knowledge Statements
di: Liang, Yiming, et al.
Pubblicazione: (2025)
di: Liang, Yiming, et al.
Pubblicazione: (2025)
Unveiling LLMs: The Evolution of Latent Representations in a Dynamic Knowledge Graph
di: Bronzini, Marco, et al.
Pubblicazione: (2024)
di: Bronzini, Marco, et al.
Pubblicazione: (2024)
Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon
di: Ma, Xuchen, et al.
Pubblicazione: (2025)
di: Ma, Xuchen, et al.
Pubblicazione: (2025)
CTourLLM: Enhancing LLMs with Chinese Tourism Knowledge
di: Wei, Qikai, et al.
Pubblicazione: (2024)
di: Wei, Qikai, et al.
Pubblicazione: (2024)
Flames: Benchmarking Value Alignment of LLMs in Chinese
di: Huang, Kexin, et al.
Pubblicazione: (2023)
di: Huang, Kexin, et al.
Pubblicazione: (2023)
Competition-Level Problems are Effective LLM Evaluators
di: Huang, Yiming, et al.
Pubblicazione: (2023)
di: Huang, Yiming, et al.
Pubblicazione: (2023)
AutoCode: LLMs as Problem Setters for Competitive Programming
di: Zhou, Shang, et al.
Pubblicazione: (2025)
di: Zhou, Shang, et al.
Pubblicazione: (2025)
Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark
di: Choi, Minje, et al.
Pubblicazione: (2023)
di: Choi, Minje, et al.
Pubblicazione: (2023)
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
di: Lin, Tianwei, et al.
Pubblicazione: (2024)
di: Lin, Tianwei, et al.
Pubblicazione: (2024)
InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles
di: Li, Zizhen, et al.
Pubblicazione: (2025)
di: Li, Zizhen, et al.
Pubblicazione: (2025)
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
di: Wu, Qinzhuo, et al.
Pubblicazione: (2026)
di: Wu, Qinzhuo, et al.
Pubblicazione: (2026)
Unveiling the Merits and Defects of LLMs in Automatic Review Generation for Scientific Papers
di: Li, Ruochi, et al.
Pubblicazione: (2025)
di: Li, Ruochi, et al.
Pubblicazione: (2025)
Discourse-Driven Evaluation: Unveiling Factual Inconsistency in Long Document Summarization
di: Zhong, Yang, et al.
Pubblicazione: (2025)
di: Zhong, Yang, et al.
Pubblicazione: (2025)
Unveiling Divergent Inductive Biases of LLMs on Temporal Data
di: Kishore, Sindhu, et al.
Pubblicazione: (2024)
di: Kishore, Sindhu, et al.
Pubblicazione: (2024)
Safety Evaluation of DeepSeek Models in Chinese Contexts
di: Zhang, Wenjing, et al.
Pubblicazione: (2025)
di: Zhang, Wenjing, et al.
Pubblicazione: (2025)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
di: Yang, Gao, et al.
Pubblicazione: (2025)
di: Yang, Gao, et al.
Pubblicazione: (2025)
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
di: Guo, Ruiling, et al.
Pubblicazione: (2025)
di: Guo, Ruiling, et al.
Pubblicazione: (2025)
PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations
di: Li, Jiatong, et al.
Pubblicazione: (2024)
di: Li, Jiatong, et al.
Pubblicazione: (2024)
Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese
di: Xu, Yunqi, et al.
Pubblicazione: (2024)
di: Xu, Yunqi, et al.
Pubblicazione: (2024)
Why Does New Knowledge Create Messy Ripple Effects in LLMs?
di: Qin, Jiaxin, et al.
Pubblicazione: (2024)
di: Qin, Jiaxin, et al.
Pubblicazione: (2024)
Evaluation Ethics of LLMs in Legal Domain
di: Zhang, Ruizhe, et al.
Pubblicazione: (2024)
di: Zhang, Ruizhe, et al.
Pubblicazione: (2024)
Unveiling the Lexical Sensitivity of LLMs: Combinatorial Optimization for Prompt Enhancement
di: Zhan, Pengwei, et al.
Pubblicazione: (2024)
di: Zhan, Pengwei, et al.
Pubblicazione: (2024)
ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation
di: Gao, Jing, et al.
Pubblicazione: (2025)
di: Gao, Jing, et al.
Pubblicazione: (2025)
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
di: Wang, Shanshan, et al.
Pubblicazione: (2025)
di: Wang, Shanshan, et al.
Pubblicazione: (2025)
Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
di: Zhang, Yichi, et al.
Pubblicazione: (2025)
di: Zhang, Yichi, et al.
Pubblicazione: (2025)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
di: Lu, Junyu, et al.
Pubblicazione: (2025)
di: Lu, Junyu, et al.
Pubblicazione: (2025)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
di: Li, Wei, et al.
Pubblicazione: (2024)
di: Li, Wei, et al.
Pubblicazione: (2024)
Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation
di: Luo, Kangcheng, et al.
Pubblicazione: (2025)
di: Luo, Kangcheng, et al.
Pubblicazione: (2025)
Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based Perspective
di: Xu, Chengyin, et al.
Pubblicazione: (2025)
di: Xu, Chengyin, et al.
Pubblicazione: (2025)
Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
di: Wang, Jun, et al.
Pubblicazione: (2025)
di: Wang, Jun, et al.
Pubblicazione: (2025)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
di: Yu, Linhao, et al.
Pubblicazione: (2024)
di: Yu, Linhao, et al.
Pubblicazione: (2024)
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
di: Zhao, Sihang, et al.
Pubblicazione: (2024)
di: Zhao, Sihang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
di: Balunović, Mislav, et al.
Pubblicazione: (2025) -
CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
di: Wu, Chengwei, et al.
Pubblicazione: (2025) -
HARBOR: Exploring Persona Dynamics in Multi-Agent Competition
di: Jiang, Kenan, et al.
Pubblicazione: (2025) -
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
di: Zhang, Jiaxin, et al.
Pubblicazione: (2024) -
Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs
di: Zeng, Liang, et al.
Pubblicazione: (2025)