KoDialogBench: Evaluating Conversational Understanding of Language Models with Korean Dialogue Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jang, Seongbo, Lee, Seonghyeon, Yu, Hwanjo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Effectiveness of Integration Methods for Multimodal Dialogue Response Retrieval
von: Jang, Seongbo, et al.
Veröffentlicht: (2025)
von: Jang, Seongbo, et al.
Veröffentlicht: (2025)
Exploring Language Model's Code Generation Ability with Auxiliary Functions
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2024)
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2024)
From What to Respond to When to Respond: Timely Response Generation for Open-domain Dialogue Agents
von: Jang, Seongbo, et al.
Veröffentlicht: (2025)
von: Jang, Seongbo, et al.
Veröffentlicht: (2025)
KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness
von: Kim, Jinyoung, et al.
Veröffentlicht: (2026)
von: Kim, Jinyoung, et al.
Veröffentlicht: (2026)
How Diversely Can Language Models Solve Problems? Exploring the Algorithmic Diversity of Model-Generated Code
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2025)
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
von: Shin, Hyopil, et al.
Veröffentlicht: (2025)
von: Shin, Hyopil, et al.
Veröffentlicht: (2025)
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
von: Jung, Sungmok, et al.
Veröffentlicht: (2026)
von: Jung, Sungmok, et al.
Veröffentlicht: (2026)
Eliciting Instruction-tuned Code Language Models' Capabilities to Utilize Auxiliary Function for Code Generation
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2024)
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2024)
FunctionChat-Bench: Comprehensive Evaluation of Language Models' Generative Capabilities in Korean Tool-use Dialogs
von: Lee, Shinbok, et al.
Veröffentlicht: (2024)
von: Lee, Shinbok, et al.
Veröffentlicht: (2024)
Verbosity-Aware Rationale Reduction: Effective Reduction of Redundant Rationale via Principled Criteria
von: Jang, Joonwon, et al.
Veröffentlicht: (2024)
von: Jang, Joonwon, et al.
Veröffentlicht: (2024)
Everyday Physics in Korean Contexts: A Culturally Grounded Physical Reasoning Benchmark
von: Jeong, Jihae, et al.
Veröffentlicht: (2025)
von: Jeong, Jihae, et al.
Veröffentlicht: (2025)
KoBBQ: Korean Bias Benchmark for Question Answering
von: Jin, Jiho, et al.
Veröffentlicht: (2023)
von: Jin, Jiho, et al.
Veröffentlicht: (2023)
TriBench-Ko: Evaluating LLM Risks in Judicial Workflows
von: Lee, Haesung, et al.
Veröffentlicht: (2026)
von: Lee, Haesung, et al.
Veröffentlicht: (2026)
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
DialogBench: Evaluating LLMs as Human-like Dialogue Systems
von: Ou, Jiao, et al.
Veröffentlicht: (2023)
von: Ou, Jiao, et al.
Veröffentlicht: (2023)
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
von: Ko, Donghyeon, et al.
Veröffentlicht: (2025)
von: Ko, Donghyeon, et al.
Veröffentlicht: (2025)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
von: Son, Guijin, et al.
Veröffentlicht: (2023)
von: Son, Guijin, et al.
Veröffentlicht: (2023)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
CARBD-Ko: A Contextually Annotated Review Benchmark Dataset for Aspect-Level Sentiment Classification in Korean
von: Jang, Dongjun, et al.
Veröffentlicht: (2024)
von: Jang, Dongjun, et al.
Veröffentlicht: (2024)
STEPER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language Models
von: Lee, Kyumin, et al.
Veröffentlicht: (2025)
von: Lee, Kyumin, et al.
Veröffentlicht: (2025)
SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents
von: Lee, Jaehoon, et al.
Veröffentlicht: (2025)
von: Lee, Jaehoon, et al.
Veröffentlicht: (2025)
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
von: Kim, Kyuhee, et al.
Veröffentlicht: (2025)
von: Kim, Kyuhee, et al.
Veröffentlicht: (2025)
KoCoNovel: Annotated Dataset of Character Coreference in Korean Novels
von: Kim, Kyuhee, et al.
Veröffentlicht: (2024)
von: Kim, Kyuhee, et al.
Veröffentlicht: (2024)
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
von: Kim, Haechan, et al.
Veröffentlicht: (2026)
von: Kim, Haechan, et al.
Veröffentlicht: (2026)
KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context
von: Wang, Xiaonan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaonan, et al.
Veröffentlicht: (2024)
REFIND at SemEval-2025 Task 3: Retrieval-Augmented Factuality Hallucination Detection in Large Language Models
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
Making Sense of Korean Sentences: A Comprehensive Evaluation of LLMs through KoSEnd Dataset
von: Yu, Seunguk, et al.
Veröffentlicht: (2025)
von: Yu, Seunguk, et al.
Veröffentlicht: (2025)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2024)
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2024)
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs
von: Fan, Zhiting, et al.
Veröffentlicht: (2024)
von: Fan, Zhiting, et al.
Veröffentlicht: (2024)
When LLM Therapists Become Salespeople: Evaluating Large Language Models for Ethical Motivational Interviewing
von: Kong, Haein, et al.
Veröffentlicht: (2025)
von: Kong, Haein, et al.
Veröffentlicht: (2025)
Pragmatic Competence Evaluation of Large Language Models for the Korean Language
von: Park, Dojun, et al.
Veröffentlicht: (2024)
von: Park, Dojun, et al.
Veröffentlicht: (2024)
KoCoSa: Korean Context-aware Sarcasm Detection Dataset
von: Kim, Yumin, et al.
Veröffentlicht: (2024)
von: Kim, Yumin, et al.
Veröffentlicht: (2024)
KoLA: Carefully Benchmarking World Knowledge of Large Language Models
von: Yu, Jifan, et al.
Veröffentlicht: (2023)
von: Yu, Jifan, et al.
Veröffentlicht: (2023)
INSTRUCTIR: A Benchmark for Instruction Following of Information Retrieval Models
von: Oh, Hanseok, et al.
Veröffentlicht: (2024)
von: Oh, Hanseok, et al.
Veröffentlicht: (2024)
VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
von: Yu, Haorui, et al.
Veröffentlicht: (2026)
von: Yu, Haorui, et al.
Veröffentlicht: (2026)
WXImpactBench: A Disruptive Weather Impact Understanding Benchmark for Evaluating Large Language Models
von: Yu, Yongan, et al.
Veröffentlicht: (2025)
von: Yu, Yongan, et al.
Veröffentlicht: (2025)
RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues
von: Kuo, Tzu-Lin, et al.
Veröffentlicht: (2024)
von: Kuo, Tzu-Lin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On the Effectiveness of Integration Methods for Multimodal Dialogue Response Retrieval
von: Jang, Seongbo, et al.
Veröffentlicht: (2025) -
Exploring Language Model's Code Generation Ability with Auxiliary Functions
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2024) -
From What to Respond to When to Respond: Timely Response Generation for Open-domain Dialogue Agents
von: Jang, Seongbo, et al.
Veröffentlicht: (2025) -
KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness
von: Kim, Jinyoung, et al.
Veröffentlicht: (2026) -
How Diversely Can Language Models Solve Problems? Exploring the Algorithmic Diversity of Model-Generated Code
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2025)