Unveiling the Competitive Dynamics: A Comparative Evaluation of American and Chinese LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Jiang, Zhenhui, Li, Jiaxin, Liu, Yang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
por: Balunović, Mislav, et al.
Publicado: (2025)
por: Balunović, Mislav, et al.
Publicado: (2025)
CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
por: Wu, Chengwei, et al.
Publicado: (2025)
por: Wu, Chengwei, et al.
Publicado: (2025)
HARBOR: Exploring Persona Dynamics in Multi-Agent Competition
por: Jiang, Kenan, et al.
Publicado: (2025)
por: Jiang, Kenan, et al.
Publicado: (2025)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
por: Zhang, Jiaxin, et al.
Publicado: (2024)
por: Zhang, Jiaxin, et al.
Publicado: (2024)
Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs
por: Zeng, Liang, et al.
Publicado: (2025)
por: Zeng, Liang, et al.
Publicado: (2025)
MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine
por: Kong, Shufeng, et al.
Publicado: (2025)
por: Kong, Shufeng, et al.
Publicado: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
por: Jiang, Botian, et al.
Publicado: (2024)
por: Jiang, Botian, et al.
Publicado: (2024)
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
por: Wu, Xinwei, et al.
Publicado: (2025)
por: Wu, Xinwei, et al.
Publicado: (2025)
Encyclo-K: Evaluating LLMs with Dynamically Composed Knowledge Statements
por: Liang, Yiming, et al.
Publicado: (2025)
por: Liang, Yiming, et al.
Publicado: (2025)
Unveiling LLMs: The Evolution of Latent Representations in a Dynamic Knowledge Graph
por: Bronzini, Marco, et al.
Publicado: (2024)
por: Bronzini, Marco, et al.
Publicado: (2024)
Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon
por: Ma, Xuchen, et al.
Publicado: (2025)
por: Ma, Xuchen, et al.
Publicado: (2025)
CTourLLM: Enhancing LLMs with Chinese Tourism Knowledge
por: Wei, Qikai, et al.
Publicado: (2024)
por: Wei, Qikai, et al.
Publicado: (2024)
Flames: Benchmarking Value Alignment of LLMs in Chinese
por: Huang, Kexin, et al.
Publicado: (2023)
por: Huang, Kexin, et al.
Publicado: (2023)
Competition-Level Problems are Effective LLM Evaluators
por: Huang, Yiming, et al.
Publicado: (2023)
por: Huang, Yiming, et al.
Publicado: (2023)
AutoCode: LLMs as Problem Setters for Competitive Programming
por: Zhou, Shang, et al.
Publicado: (2025)
por: Zhou, Shang, et al.
Publicado: (2025)
Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark
por: Choi, Minje, et al.
Publicado: (2023)
por: Choi, Minje, et al.
Publicado: (2023)
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
por: Lin, Tianwei, et al.
Publicado: (2024)
por: Lin, Tianwei, et al.
Publicado: (2024)
InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles
por: Li, Zizhen, et al.
Publicado: (2025)
por: Li, Zizhen, et al.
Publicado: (2025)
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
por: Wu, Qinzhuo, et al.
Publicado: (2026)
por: Wu, Qinzhuo, et al.
Publicado: (2026)
Unveiling the Merits and Defects of LLMs in Automatic Review Generation for Scientific Papers
por: Li, Ruochi, et al.
Publicado: (2025)
por: Li, Ruochi, et al.
Publicado: (2025)
Discourse-Driven Evaluation: Unveiling Factual Inconsistency in Long Document Summarization
por: Zhong, Yang, et al.
Publicado: (2025)
por: Zhong, Yang, et al.
Publicado: (2025)
Unveiling Divergent Inductive Biases of LLMs on Temporal Data
por: Kishore, Sindhu, et al.
Publicado: (2024)
por: Kishore, Sindhu, et al.
Publicado: (2024)
Safety Evaluation of DeepSeek Models in Chinese Contexts
por: Zhang, Wenjing, et al.
Publicado: (2025)
por: Zhang, Wenjing, et al.
Publicado: (2025)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
por: Yang, Gao, et al.
Publicado: (2025)
por: Yang, Gao, et al.
Publicado: (2025)
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
por: Guo, Ruiling, et al.
Publicado: (2025)
por: Guo, Ruiling, et al.
Publicado: (2025)
PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations
por: Li, Jiatong, et al.
Publicado: (2024)
por: Li, Jiatong, et al.
Publicado: (2024)
Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese
por: Xu, Yunqi, et al.
Publicado: (2024)
por: Xu, Yunqi, et al.
Publicado: (2024)
Why Does New Knowledge Create Messy Ripple Effects in LLMs?
por: Qin, Jiaxin, et al.
Publicado: (2024)
por: Qin, Jiaxin, et al.
Publicado: (2024)
Evaluation Ethics of LLMs in Legal Domain
por: Zhang, Ruizhe, et al.
Publicado: (2024)
por: Zhang, Ruizhe, et al.
Publicado: (2024)
Unveiling the Lexical Sensitivity of LLMs: Combinatorial Optimization for Prompt Enhancement
por: Zhan, Pengwei, et al.
Publicado: (2024)
por: Zhan, Pengwei, et al.
Publicado: (2024)
ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation
por: Gao, Jing, et al.
Publicado: (2025)
por: Gao, Jing, et al.
Publicado: (2025)
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
por: Wang, Shanshan, et al.
Publicado: (2025)
por: Wang, Shanshan, et al.
Publicado: (2025)
Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
por: Zhang, Yichi, et al.
Publicado: (2025)
por: Zhang, Yichi, et al.
Publicado: (2025)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
por: Lu, Junyu, et al.
Publicado: (2025)
por: Lu, Junyu, et al.
Publicado: (2025)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
por: Li, Wei, et al.
Publicado: (2024)
por: Li, Wei, et al.
Publicado: (2024)
Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation
por: Luo, Kangcheng, et al.
Publicado: (2025)
por: Luo, Kangcheng, et al.
Publicado: (2025)
Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based Perspective
por: Xu, Chengyin, et al.
Publicado: (2025)
por: Xu, Chengyin, et al.
Publicado: (2025)
Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
por: Wang, Jun, et al.
Publicado: (2025)
por: Wang, Jun, et al.
Publicado: (2025)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
por: Yu, Linhao, et al.
Publicado: (2024)
por: Yu, Linhao, et al.
Publicado: (2024)
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
por: Zhao, Sihang, et al.
Publicado: (2024)
por: Zhao, Sihang, et al.
Publicado: (2024)
Ejemplares similares
-
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
por: Balunović, Mislav, et al.
Publicado: (2025) -
CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
por: Wu, Chengwei, et al.
Publicado: (2025) -
HARBOR: Exploring Persona Dynamics in Multi-Agent Competition
por: Jiang, Kenan, et al.
Publicado: (2025) -
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
por: Zhang, Jiaxin, et al.
Publicado: (2024) -
Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs
por: Zeng, Liang, et al.
Publicado: (2025)