Enregistré dans:
| Auteurs principaux: | Seo, Jean, Lim, Jongwon, Jang, Dongjun, Shin, Hyopil |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2411.09255 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
CARBD-Ko: A Contextually Annotated Review Benchmark Dataset for Aspect-Level Sentiment Classification in Korean
par: Jang, Dongjun, et autres
Publié: (2024)
par: Jang, Dongjun, et autres
Publié: (2024)
A Study on How Attention Scores in the BERT Model are Aware of Lexical Categories in Syntactic and Semantic Tasks on the GLUE Benchmark
par: Jang, Dongjun, et autres
Publié: (2024)
par: Jang, Dongjun, et autres
Publié: (2024)
P-CoT: A Pedagogically-motivated Participatory Chain-of-Thought Prompting for Phonological Reasoning in LLMs
par: Jang, Dongjun, et autres
Publié: (2025)
par: Jang, Dongjun, et autres
Publié: (2025)
RCScore: Quantifying Response Consistency in Large Language Models
par: Jang, Dongjun, et autres
Publié: (2025)
par: Jang, Dongjun, et autres
Publié: (2025)
KIT-19: A Comprehensive Korean Instruction Toolkit on 19 Tasks for Fine-Tuning Korean Large Language Models
par: Jang, Dongjun, et autres
Publié: (2024)
par: Jang, Dongjun, et autres
Publié: (2024)
Korean Bio-Medical Corpus (KBMC) for Medical Named Entity Recognition
par: Byun, Sungjoo, et autres
Publié: (2024)
par: Byun, Sungjoo, et autres
Publié: (2024)
MoFE: Mixture of Frozen Experts Architecture
par: Seo, Jean, et autres
Publié: (2025)
par: Seo, Jean, et autres
Publié: (2025)
How does a Language-Specific Tokenizer affect LLMs?
par: Seo, Jean, et autres
Publié: (2025)
par: Seo, Jean, et autres
Publié: (2025)
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
par: Shin, Hyopil, et autres
Publié: (2025)
par: Shin, Hyopil, et autres
Publié: (2025)
Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines
par: Seo, Jean, et autres
Publié: (2026)
par: Seo, Jean, et autres
Publié: (2026)
The Impact of Negated Text on Hallucination with Large Language Models
par: Seo, Jaehyung, et autres
Publié: (2025)
par: Seo, Jaehyung, et autres
Publié: (2025)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
par: Kim, Dongjun, et autres
Publié: (2025)
par: Kim, Dongjun, et autres
Publié: (2025)
CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
par: Zhong, Hanmeng, et autres
Publié: (2025)
par: Zhong, Hanmeng, et autres
Publié: (2025)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
par: Costa-jussà, Marta R., et autres
Publié: (2024)
par: Costa-jussà, Marta R., et autres
Publié: (2024)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
par: Lim, Hyeonseok, et autres
Publié: (2024)
par: Lim, Hyeonseok, et autres
Publié: (2024)
Persona Dynamics: Unveiling the Impact of Personality Traits on Agents in Text-Based Games
par: Lim, Seungwon, et autres
Publié: (2025)
par: Lim, Seungwon, et autres
Publié: (2025)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
par: Chen, Junjie, et autres
Publié: (2026)
par: Chen, Junjie, et autres
Publié: (2026)
Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation
par: Qin, Chengwei, et autres
Publié: (2025)
par: Qin, Chengwei, et autres
Publié: (2025)
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
par: Kim, Dongjun, et autres
Publié: (2025)
par: Kim, Dongjun, et autres
Publié: (2025)
Sanity Checks for Long-Form Hallucination Detection
par: Zollicoffer, Geigh, et autres
Publié: (2026)
par: Zollicoffer, Geigh, et autres
Publié: (2026)
How language models extrapolate outside the training data: A case study in Textualized Gridworld
par: Kim, Doyoung, et autres
Publié: (2024)
par: Kim, Doyoung, et autres
Publié: (2024)
Careless Whisper: Speech-to-Text Hallucination Harms
par: Koenecke, Allison, et autres
Publié: (2024)
par: Koenecke, Allison, et autres
Publié: (2024)
From Raw Corpora to Domain Benchmarks: Automated Evaluation of LLM Domain Expertise
par: Sharma, Nitin, et autres
Publié: (2025)
par: Sharma, Nitin, et autres
Publié: (2025)
C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation
par: Zhang, Xu, et autres
Publié: (2025)
par: Zhang, Xu, et autres
Publié: (2025)
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
par: Wang, Shuting, et autres
Publié: (2024)
par: Wang, Shuting, et autres
Publié: (2024)
MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents
par: Shin, Joongmin, et autres
Publié: (2026)
par: Shin, Joongmin, et autres
Publié: (2026)
INSTRUCTIR: A Benchmark for Instruction Following of Information Retrieval Models
par: Oh, Hanseok, et autres
Publié: (2024)
par: Oh, Hanseok, et autres
Publié: (2024)
Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction
par: Park, Sejun, et autres
Publié: (2026)
par: Park, Sejun, et autres
Publié: (2026)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
par: Wang, Lei, et autres
Publié: (2024)
par: Wang, Lei, et autres
Publié: (2024)
Automated Collection of Evaluation Dataset for Semantic Search in Low-Resource Domain Language
par: Zhukova, Anastasia, et autres
Publié: (2024)
par: Zhukova, Anastasia, et autres
Publié: (2024)
Benchmarking Automated Clinical Language Simplification: Dataset, Algorithm, and Evaluation
par: Luo, Junyu, et autres
Publié: (2020)
par: Luo, Junyu, et autres
Publié: (2020)
Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models
par: Moon, Hyeonseok, et autres
Publié: (2025)
par: Moon, Hyeonseok, et autres
Publié: (2025)
AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
par: Liu, Xuannan, et autres
Publié: (2026)
par: Liu, Xuannan, et autres
Publié: (2026)
Real-Time Detection of Hallucinated Entities in Long-Form Generation
par: Obeso, Oscar, et autres
Publié: (2025)
par: Obeso, Oscar, et autres
Publié: (2025)
Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
par: Hwang, Jena D., et autres
Publié: (2026)
par: Hwang, Jena D., et autres
Publié: (2026)
KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness
par: Kim, Jinyoung, et autres
Publié: (2026)
par: Kim, Jinyoung, et autres
Publié: (2026)
DOLOMITES: Domain-Specific Long-Form Methodical Tasks
par: Malaviya, Chaitanya, et autres
Publié: (2024)
par: Malaviya, Chaitanya, et autres
Publié: (2024)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
par: Wu, Yuhao, et autres
Publié: (2024)
par: Wu, Yuhao, et autres
Publié: (2024)
Evaluating Text Creativity across Diverse Domains: A Dataset and Large Language Model Evaluator
par: Cao, Qian, et autres
Publié: (2025)
par: Cao, Qian, et autres
Publié: (2025)
AutoHall: Automated Factuality Hallucination Dataset Generation for Large Language Models
par: Cao, Zouying, et autres
Publié: (2023)
par: Cao, Zouying, et autres
Publié: (2023)
Documents similaires
-
CARBD-Ko: A Contextually Annotated Review Benchmark Dataset for Aspect-Level Sentiment Classification in Korean
par: Jang, Dongjun, et autres
Publié: (2024) -
A Study on How Attention Scores in the BERT Model are Aware of Lexical Categories in Syntactic and Semantic Tasks on the GLUE Benchmark
par: Jang, Dongjun, et autres
Publié: (2024) -
P-CoT: A Pedagogically-motivated Participatory Chain-of-Thought Prompting for Phonological Reasoning in LLMs
par: Jang, Dongjun, et autres
Publié: (2025) -
RCScore: Quantifying Response Consistency in Large Language Models
par: Jang, Dongjun, et autres
Publié: (2025) -
KIT-19: A Comprehensive Korean Instruction Toolkit on 19 Tasks for Fine-Tuning Korean Large Language Models
par: Jang, Dongjun, et autres
Publié: (2024)