ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Haibin, Lv, Kangtao, Hu, Chengwei, Li, Yanshi, Yuan, Yujin, He, Yancheng, Zhang, Xingyao, Liu, Langming, Liu, Shilei, Su, Wenbo, Zheng, Bo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916633692340224
author Chen, Haibin
Lv, Kangtao
Hu, Chengwei
Li, Yanshi
Yuan, Yujin
He, Yancheng
Zhang, Xingyao
Liu, Langming
Liu, Shilei
Su, Wenbo
Zheng, Bo
author_facet Chen, Haibin
Lv, Kangtao
Hu, Chengwei
Li, Yanshi
Yuan, Yujin
He, Yancheng
Zhang, Xingyao
Liu, Langming
Liu, Shilei
Su, Wenbo
Zheng, Bo
contents With the increasing use of Large Language Models (LLMs) in fields such as e-commerce, domain-specific concept evaluation benchmarks are crucial for assessing their domain capabilities. Existing LLMs may generate factually incorrect information within the complex e-commerce applications. Therefore, it is necessary to build an e-commerce concept benchmark. Existing benchmarks encounter two primary challenges: (1) handle the heterogeneous and diverse nature of tasks, (2) distinguish between generality and specificity within the e-commerce field. To address these problems, we propose \textbf{ChineseEcomQA}, a scalable question-answering benchmark focused on fundamental e-commerce concepts. ChineseEcomQA is built on three core characteristics: \textbf{Focus on Fundamental Concept}, \textbf{E-commerce Generality} and \textbf{E-commerce Expertise}. Fundamental concepts are designed to be applicable across a diverse array of e-commerce tasks, thus addressing the challenge of heterogeneity and diversity. Additionally, by carefully balancing generality and specificity, ChineseEcomQA effectively differentiates between broad e-commerce concepts, allowing for precise validation of domain capabilities. We achieve this through a scalable benchmark construction process that combines LLM validation, Retrieval-Augmented Generation (RAG) validation, and rigorous manual annotation. Based on ChineseEcomQA, we conduct extensive evaluations on mainstream LLMs and provide some valuable insights. We hope that ChineseEcomQA could guide future domain-specific evaluations, and facilitate broader LLM adoption in e-commerce applications.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20196
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models
Chen, Haibin
Lv, Kangtao
Hu, Chengwei
Li, Yanshi
Yuan, Yujin
He, Yancheng
Zhang, Xingyao
Liu, Langming
Liu, Shilei
Su, Wenbo
Zheng, Bo
Computation and Language
With the increasing use of Large Language Models (LLMs) in fields such as e-commerce, domain-specific concept evaluation benchmarks are crucial for assessing their domain capabilities. Existing LLMs may generate factually incorrect information within the complex e-commerce applications. Therefore, it is necessary to build an e-commerce concept benchmark. Existing benchmarks encounter two primary challenges: (1) handle the heterogeneous and diverse nature of tasks, (2) distinguish between generality and specificity within the e-commerce field. To address these problems, we propose \textbf{ChineseEcomQA}, a scalable question-answering benchmark focused on fundamental e-commerce concepts. ChineseEcomQA is built on three core characteristics: \textbf{Focus on Fundamental Concept}, \textbf{E-commerce Generality} and \textbf{E-commerce Expertise}. Fundamental concepts are designed to be applicable across a diverse array of e-commerce tasks, thus addressing the challenge of heterogeneity and diversity. Additionally, by carefully balancing generality and specificity, ChineseEcomQA effectively differentiates between broad e-commerce concepts, allowing for precise validation of domain capabilities. We achieve this through a scalable benchmark construction process that combines LLM validation, Retrieval-Augmented Generation (RAG) validation, and rigorous manual annotation. Based on ChineseEcomQA, we conduct extensive evaluations on mainstream LLMs and provide some valuable insights. We hope that ChineseEcomQA could guide future domain-specific evaluations, and facilitate broader LLM adoption in e-commerce applications.
title ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2502.20196