Saved in:
| Main Authors: | Seo, Jean, Lim, Jongwon, Jang, Dongjun, Shin, Hyopil |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.09255 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CARBD-Ko: A Contextually Annotated Review Benchmark Dataset for Aspect-Level Sentiment Classification in Korean
by: Jang, Dongjun, et al.
Published: (2024)
by: Jang, Dongjun, et al.
Published: (2024)
A Study on How Attention Scores in the BERT Model are Aware of Lexical Categories in Syntactic and Semantic Tasks on the GLUE Benchmark
by: Jang, Dongjun, et al.
Published: (2024)
by: Jang, Dongjun, et al.
Published: (2024)
P-CoT: A Pedagogically-motivated Participatory Chain-of-Thought Prompting for Phonological Reasoning in LLMs
by: Jang, Dongjun, et al.
Published: (2025)
by: Jang, Dongjun, et al.
Published: (2025)
RCScore: Quantifying Response Consistency in Large Language Models
by: Jang, Dongjun, et al.
Published: (2025)
by: Jang, Dongjun, et al.
Published: (2025)
KIT-19: A Comprehensive Korean Instruction Toolkit on 19 Tasks for Fine-Tuning Korean Large Language Models
by: Jang, Dongjun, et al.
Published: (2024)
by: Jang, Dongjun, et al.
Published: (2024)
Korean Bio-Medical Corpus (KBMC) for Medical Named Entity Recognition
by: Byun, Sungjoo, et al.
Published: (2024)
by: Byun, Sungjoo, et al.
Published: (2024)
MoFE: Mixture of Frozen Experts Architecture
by: Seo, Jean, et al.
Published: (2025)
by: Seo, Jean, et al.
Published: (2025)
How does a Language-Specific Tokenizer affect LLMs?
by: Seo, Jean, et al.
Published: (2025)
by: Seo, Jean, et al.
Published: (2025)
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
by: Shin, Hyopil, et al.
Published: (2025)
by: Shin, Hyopil, et al.
Published: (2025)
Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines
by: Seo, Jean, et al.
Published: (2026)
by: Seo, Jean, et al.
Published: (2026)
The Impact of Negated Text on Hallucination with Large Language Models
by: Seo, Jaehyung, et al.
Published: (2025)
by: Seo, Jaehyung, et al.
Published: (2025)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
by: Kim, Dongjun, et al.
Published: (2025)
by: Kim, Dongjun, et al.
Published: (2025)
CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
by: Zhong, Hanmeng, et al.
Published: (2025)
by: Zhong, Hanmeng, et al.
Published: (2025)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
by: Lim, Hyeonseok, et al.
Published: (2024)
by: Lim, Hyeonseok, et al.
Published: (2024)
Persona Dynamics: Unveiling the Impact of Personality Traits on Agents in Text-Based Games
by: Lim, Seungwon, et al.
Published: (2025)
by: Lim, Seungwon, et al.
Published: (2025)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
by: Chen, Junjie, et al.
Published: (2026)
by: Chen, Junjie, et al.
Published: (2026)
Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation
by: Qin, Chengwei, et al.
Published: (2025)
by: Qin, Chengwei, et al.
Published: (2025)
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
by: Kim, Dongjun, et al.
Published: (2025)
by: Kim, Dongjun, et al.
Published: (2025)
Sanity Checks for Long-Form Hallucination Detection
by: Zollicoffer, Geigh, et al.
Published: (2026)
by: Zollicoffer, Geigh, et al.
Published: (2026)
How language models extrapolate outside the training data: A case study in Textualized Gridworld
by: Kim, Doyoung, et al.
Published: (2024)
by: Kim, Doyoung, et al.
Published: (2024)
Careless Whisper: Speech-to-Text Hallucination Harms
by: Koenecke, Allison, et al.
Published: (2024)
by: Koenecke, Allison, et al.
Published: (2024)
From Raw Corpora to Domain Benchmarks: Automated Evaluation of LLM Domain Expertise
by: Sharma, Nitin, et al.
Published: (2025)
by: Sharma, Nitin, et al.
Published: (2025)
C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
by: Wang, Shuting, et al.
Published: (2024)
by: Wang, Shuting, et al.
Published: (2024)
MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents
by: Shin, Joongmin, et al.
Published: (2026)
by: Shin, Joongmin, et al.
Published: (2026)
INSTRUCTIR: A Benchmark for Instruction Following of Information Retrieval Models
by: Oh, Hanseok, et al.
Published: (2024)
by: Oh, Hanseok, et al.
Published: (2024)
Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction
by: Park, Sejun, et al.
Published: (2026)
by: Park, Sejun, et al.
Published: (2026)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Automated Collection of Evaluation Dataset for Semantic Search in Low-Resource Domain Language
by: Zhukova, Anastasia, et al.
Published: (2024)
by: Zhukova, Anastasia, et al.
Published: (2024)
Benchmarking Automated Clinical Language Simplification: Dataset, Algorithm, and Evaluation
by: Luo, Junyu, et al.
Published: (2020)
by: Luo, Junyu, et al.
Published: (2020)
Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models
by: Moon, Hyeonseok, et al.
Published: (2025)
by: Moon, Hyeonseok, et al.
Published: (2025)
AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
by: Liu, Xuannan, et al.
Published: (2026)
by: Liu, Xuannan, et al.
Published: (2026)
Real-Time Detection of Hallucinated Entities in Long-Form Generation
by: Obeso, Oscar, et al.
Published: (2025)
by: Obeso, Oscar, et al.
Published: (2025)
Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
by: Hwang, Jena D., et al.
Published: (2026)
by: Hwang, Jena D., et al.
Published: (2026)
KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness
by: Kim, Jinyoung, et al.
Published: (2026)
by: Kim, Jinyoung, et al.
Published: (2026)
DOLOMITES: Domain-Specific Long-Form Methodical Tasks
by: Malaviya, Chaitanya, et al.
Published: (2024)
by: Malaviya, Chaitanya, et al.
Published: (2024)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
Evaluating Text Creativity across Diverse Domains: A Dataset and Large Language Model Evaluator
by: Cao, Qian, et al.
Published: (2025)
by: Cao, Qian, et al.
Published: (2025)
AutoHall: Automated Factuality Hallucination Dataset Generation for Large Language Models
by: Cao, Zouying, et al.
Published: (2023)
by: Cao, Zouying, et al.
Published: (2023)
Similar Items
-
CARBD-Ko: A Contextually Annotated Review Benchmark Dataset for Aspect-Level Sentiment Classification in Korean
by: Jang, Dongjun, et al.
Published: (2024) -
A Study on How Attention Scores in the BERT Model are Aware of Lexical Categories in Syntactic and Semantic Tasks on the GLUE Benchmark
by: Jang, Dongjun, et al.
Published: (2024) -
P-CoT: A Pedagogically-motivated Participatory Chain-of-Thought Prompting for Phonological Reasoning in LLMs
by: Jang, Dongjun, et al.
Published: (2025) -
RCScore: Quantifying Response Consistency in Large Language Models
by: Jang, Dongjun, et al.
Published: (2025) -
KIT-19: A Comprehensive Korean Instruction Toolkit on 19 Tasks for Fine-Tuning Korean Large Language Models
by: Jang, Dongjun, et al.
Published: (2024)