Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chang, Chen-Chi, Chen, Ching-Yuan, Lee, Hung-Shin, Lee, Chih-Cheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909325049462784
author Chang, Chen-Chi
Chen, Ching-Yuan
Lee, Hung-Shin
Lee, Chih-Cheng
author_facet Chang, Chen-Chi
Chen, Ching-Yuan
Lee, Hung-Shin
Lee, Chih-Cheng
contents This study introduces a comprehensive benchmark designed to evaluate the performance of large language models (LLMs) in understanding and processing cultural knowledge, with a specific focus on Hakka culture as a case study. Leveraging Bloom's Taxonomy, the study develops a multi-dimensional framework that systematically assesses LLMs across six cognitive domains: Remembering, Understanding, Applying, Analyzing, Evaluating, and Creating. This benchmark extends beyond traditional single-dimensional evaluations by providing a deeper analysis of LLMs' abilities to handle culturally specific content, ranging from basic recall of facts to higher-order cognitive tasks such as creative synthesis. Additionally, the study integrates Retrieval-Augmented Generation (RAG) technology to address the challenges of minority cultural knowledge representation in LLMs, demonstrating how RAG enhances the models' performance by dynamically incorporating relevant external information. The results highlight the effectiveness of RAG in improving accuracy across all cognitive domains, particularly in tasks requiring precise retrieval and application of cultural knowledge. However, the findings also reveal the limitations of RAG in creative tasks, underscoring the need for further optimization. This benchmark provides a robust tool for evaluating and comparing LLMs in culturally diverse contexts, offering valuable insights for future research and development in AI-driven cultural knowledge preservation and dissemination.
format Preprint
id arxiv_https___arxiv_org_abs_2409_01556
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture
Chang, Chen-Chi
Chen, Ching-Yuan
Lee, Hung-Shin
Lee, Chih-Cheng
Computation and Language
Artificial Intelligence
This study introduces a comprehensive benchmark designed to evaluate the performance of large language models (LLMs) in understanding and processing cultural knowledge, with a specific focus on Hakka culture as a case study. Leveraging Bloom's Taxonomy, the study develops a multi-dimensional framework that systematically assesses LLMs across six cognitive domains: Remembering, Understanding, Applying, Analyzing, Evaluating, and Creating. This benchmark extends beyond traditional single-dimensional evaluations by providing a deeper analysis of LLMs' abilities to handle culturally specific content, ranging from basic recall of facts to higher-order cognitive tasks such as creative synthesis. Additionally, the study integrates Retrieval-Augmented Generation (RAG) technology to address the challenges of minority cultural knowledge representation in LLMs, demonstrating how RAG enhances the models' performance by dynamically incorporating relevant external information. The results highlight the effectiveness of RAG in improving accuracy across all cognitive domains, particularly in tasks requiring precise retrieval and application of cultural knowledge. However, the findings also reveal the limitations of RAG in creative tasks, underscoring the need for further optimization. This benchmark provides a robust tool for evaluating and comparing LLMs in culturally diverse contexts, offering valuable insights for future research and development in AI-driven cultural knowledge preservation and dissemination.
title Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2409.01556