Unveiling the Competitive Dynamics: A Comparative Evaluation of American and Chinese LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Zhenhui, Li, Jiaxin, Liu, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
by: Wu, Chengwei, et al.
Published: (2025)
by: Wu, Chengwei, et al.
Published: (2025)
HARBOR: Exploring Persona Dynamics in Multi-Agent Competition
by: Jiang, Kenan, et al.
Published: (2025)
by: Jiang, Kenan, et al.
Published: (2025)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs
by: Zeng, Liang, et al.
Published: (2025)
by: Zeng, Liang, et al.
Published: (2025)
MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine
by: Kong, Shufeng, et al.
Published: (2025)
by: Kong, Shufeng, et al.
Published: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
by: Jiang, Botian, et al.
Published: (2024)
by: Jiang, Botian, et al.
Published: (2024)
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
by: Wu, Xinwei, et al.
Published: (2025)
by: Wu, Xinwei, et al.
Published: (2025)
Encyclo-K: Evaluating LLMs with Dynamically Composed Knowledge Statements
by: Liang, Yiming, et al.
Published: (2025)
by: Liang, Yiming, et al.
Published: (2025)
Unveiling LLMs: The Evolution of Latent Representations in a Dynamic Knowledge Graph
by: Bronzini, Marco, et al.
Published: (2024)
by: Bronzini, Marco, et al.
Published: (2024)
Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon
by: Ma, Xuchen, et al.
Published: (2025)
by: Ma, Xuchen, et al.
Published: (2025)
CTourLLM: Enhancing LLMs with Chinese Tourism Knowledge
by: Wei, Qikai, et al.
Published: (2024)
by: Wei, Qikai, et al.
Published: (2024)
Flames: Benchmarking Value Alignment of LLMs in Chinese
by: Huang, Kexin, et al.
Published: (2023)
by: Huang, Kexin, et al.
Published: (2023)
Competition-Level Problems are Effective LLM Evaluators
by: Huang, Yiming, et al.
Published: (2023)
by: Huang, Yiming, et al.
Published: (2023)
AutoCode: LLMs as Problem Setters for Competitive Programming
by: Zhou, Shang, et al.
Published: (2025)
by: Zhou, Shang, et al.
Published: (2025)
Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark
by: Choi, Minje, et al.
Published: (2023)
by: Choi, Minje, et al.
Published: (2023)
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
by: Lin, Tianwei, et al.
Published: (2024)
by: Lin, Tianwei, et al.
Published: (2024)
InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles
by: Li, Zizhen, et al.
Published: (2025)
by: Li, Zizhen, et al.
Published: (2025)
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
by: Wu, Qinzhuo, et al.
Published: (2026)
by: Wu, Qinzhuo, et al.
Published: (2026)
Unveiling the Merits and Defects of LLMs in Automatic Review Generation for Scientific Papers
by: Li, Ruochi, et al.
Published: (2025)
by: Li, Ruochi, et al.
Published: (2025)
Discourse-Driven Evaluation: Unveiling Factual Inconsistency in Long Document Summarization
by: Zhong, Yang, et al.
Published: (2025)
by: Zhong, Yang, et al.
Published: (2025)
Unveiling Divergent Inductive Biases of LLMs on Temporal Data
by: Kishore, Sindhu, et al.
Published: (2024)
by: Kishore, Sindhu, et al.
Published: (2024)
Safety Evaluation of DeepSeek Models in Chinese Contexts
by: Zhang, Wenjing, et al.
Published: (2025)
by: Zhang, Wenjing, et al.
Published: (2025)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
by: Yang, Gao, et al.
Published: (2025)
by: Yang, Gao, et al.
Published: (2025)
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
by: Guo, Ruiling, et al.
Published: (2025)
by: Guo, Ruiling, et al.
Published: (2025)
PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations
by: Li, Jiatong, et al.
Published: (2024)
by: Li, Jiatong, et al.
Published: (2024)
Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese
by: Xu, Yunqi, et al.
Published: (2024)
by: Xu, Yunqi, et al.
Published: (2024)
Why Does New Knowledge Create Messy Ripple Effects in LLMs?
by: Qin, Jiaxin, et al.
Published: (2024)
by: Qin, Jiaxin, et al.
Published: (2024)
Evaluation Ethics of LLMs in Legal Domain
by: Zhang, Ruizhe, et al.
Published: (2024)
by: Zhang, Ruizhe, et al.
Published: (2024)
Unveiling the Lexical Sensitivity of LLMs: Combinatorial Optimization for Prompt Enhancement
by: Zhan, Pengwei, et al.
Published: (2024)
by: Zhan, Pengwei, et al.
Published: (2024)
ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation
by: Gao, Jing, et al.
Published: (2025)
by: Gao, Jing, et al.
Published: (2025)
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
by: Wang, Shanshan, et al.
Published: (2025)
by: Wang, Shanshan, et al.
Published: (2025)
Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
by: Lu, Junyu, et al.
Published: (2025)
by: Lu, Junyu, et al.
Published: (2025)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation
by: Luo, Kangcheng, et al.
Published: (2025)
by: Luo, Kangcheng, et al.
Published: (2025)
Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based Perspective
by: Xu, Chengyin, et al.
Published: (2025)
by: Xu, Chengyin, et al.
Published: (2025)
Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
by: Wang, Jun, et al.
Published: (2025)
by: Wang, Jun, et al.
Published: (2025)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
by: Yu, Linhao, et al.
Published: (2024)
by: Yu, Linhao, et al.
Published: (2024)
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
by: Zhao, Sihang, et al.
Published: (2024)
by: Zhao, Sihang, et al.
Published: (2024)
Similar Items
-
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
by: Balunović, Mislav, et al.
Published: (2025) -
CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
by: Wu, Chengwei, et al.
Published: (2025) -
HARBOR: Exploring Persona Dynamics in Multi-Agent Competition
by: Jiang, Kenan, et al.
Published: (2025) -
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
by: Zhang, Jiaxin, et al.
Published: (2024) -
Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs
by: Zeng, Liang, et al.
Published: (2025)