Saved in:
| Main Authors: | Lee, Taehyun, Hong, Seokhee, Ahn, Jaewoo, Hong, Ilgee, Lee, Hwaran, Yun, Sangdoo, Shin, Jamin, Kim, Gunhee |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2305.15060 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TimeChara: Evaluating Point-in-Time Character Hallucination of Role-Playing Large Language Models
by: Ahn, Jaewoo, et al.
Published: (2024)
by: Ahn, Jaewoo, et al.
Published: (2024)
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs
by: Yoo, Haneul, et al.
Published: (2024)
by: Yoo, Haneul, et al.
Published: (2024)
Is a Peeled Apple Still Red? Evaluating LLMs' Ability for Conceptual Combination with Property Type
by: Song, Seokwon, et al.
Published: (2025)
by: Song, Seokwon, et al.
Published: (2025)
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
by: Kim, Seungone, et al.
Published: (2023)
by: Kim, Seungone, et al.
Published: (2023)
Calibrating Large Language Models Using Their Generations Only
by: Ulmer, Dennis, et al.
Published: (2024)
by: Ulmer, Dennis, et al.
Published: (2024)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
by: Lim, Junyoung, et al.
Published: (2025)
by: Lim, Junyoung, et al.
Published: (2025)
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
by: Ahn, Jaewoo, et al.
Published: (2025)
by: Ahn, Jaewoo, et al.
Published: (2025)
Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
by: Yoo, Haneul, et al.
Published: (2024)
by: Yoo, Haneul, et al.
Published: (2024)
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
by: Gubri, Martin, et al.
Published: (2024)
by: Gubri, Martin, et al.
Published: (2024)
Who Wrote This? Identifying Machine vs Human-Generated Text in Hausa
by: Sani, Babangida, et al.
Published: (2025)
by: Sani, Babangida, et al.
Published: (2025)
Who Wrote the Book? Detecting and Attributing LLM Ghostwriters
by: Shetty, Anudeex, et al.
Published: (2026)
by: Shetty, Anudeex, et al.
Published: (2026)
Teaching Language Models to Think in Code
by: Hwang, Hyeon, et al.
Published: (2026)
by: Hwang, Hyeon, et al.
Published: (2026)
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
by: Ahn, Jaewoo, et al.
Published: (2025)
by: Ahn, Jaewoo, et al.
Published: (2025)
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
by: Hong, Seokhee, et al.
Published: (2025)
by: Hong, Seokhee, et al.
Published: (2025)
Who Wrote This Line? Evaluating the Detection of LLM-Generated Classical Chinese Poetry
by: Li, Jiang, et al.
Published: (2026)
by: Li, Jiang, et al.
Published: (2026)
Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore
by: Wu, Junchao, et al.
Published: (2024)
by: Wu, Junchao, et al.
Published: (2024)
Alignment Data Map for Efficient Preference Data Selection and Diagnosis
by: Lee, Seohyeong, et al.
Published: (2025)
by: Lee, Seohyeong, et al.
Published: (2025)
Toward Interactive Regional Understanding in Vision-Large Language Models
by: Lee, Jungbeom, et al.
Published: (2024)
by: Lee, Jungbeom, et al.
Published: (2024)
Generative Visual Code Mobile World Models
by: Koh, Woosung, et al.
Published: (2026)
by: Koh, Woosung, et al.
Published: (2026)
AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective Intelligence
by: Kim, Minbeom, et al.
Published: (2024)
by: Kim, Minbeom, et al.
Published: (2024)
Guaranteed Generation from Large Language Models
by: Kim, Minbeom, et al.
Published: (2024)
by: Kim, Minbeom, et al.
Published: (2024)
MAQA: Evaluating Uncertainty Quantification in LLMs Regarding Data Uncertainty
by: Yang, Yongjin, et al.
Published: (2024)
by: Yang, Yongjin, et al.
Published: (2024)
Can Language Models Laugh at YouTube Short-form Videos?
by: Ko, Dayoon, et al.
Published: (2023)
by: Ko, Dayoon, et al.
Published: (2023)
KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge
by: Lee, Jiyoung, et al.
Published: (2024)
by: Lee, Jiyoung, et al.
Published: (2024)
LifeTox: Unveiling Implicit Toxicity in Life Advice
by: Kim, Minbeom, et al.
Published: (2023)
by: Kim, Minbeom, et al.
Published: (2023)
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
by: Kim, Tae Soo, et al.
Published: (2023)
by: Kim, Tae Soo, et al.
Published: (2023)
Drift: Decoding-time Personalized Alignments with Implicit User Preferences
by: Kim, Minbeom, et al.
Published: (2025)
by: Kim, Minbeom, et al.
Published: (2025)
KoBBQ: Korean Bias Benchmark for Question Answering
by: Jin, Jiho, et al.
Published: (2023)
by: Jin, Jiho, et al.
Published: (2023)
HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
by: Paik, Gio, et al.
Published: (2025)
by: Paik, Gio, et al.
Published: (2025)
OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
by: Liu, Tianci, et al.
Published: (2025)
by: Liu, Tianci, et al.
Published: (2025)
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
by: Kim, Jang-Hyun, et al.
Published: (2026)
by: Kim, Jang-Hyun, et al.
Published: (2026)
UniCoM: A Universal Code-Switching Speech Generator
by: Lee, Sangmin, et al.
Published: (2025)
by: Lee, Sangmin, et al.
Published: (2025)
Assessing the Answerability of Queries in Retrieval-Augmented Code Generation
by: Kim, Geonmin, et al.
Published: (2024)
by: Kim, Geonmin, et al.
Published: (2024)
Eliciting Instruction-tuned Code Language Models' Capabilities to Utilize Auxiliary Function for Code Generation
by: Lee, Seonghyeon, et al.
Published: (2024)
by: Lee, Seonghyeon, et al.
Published: (2024)
LingoQ: Bridging the Gap between EFL Learning and Work through AI-Generated Work-Related Quizzes
by: Yang, Yeonsun, et al.
Published: (2025)
by: Yang, Yeonsun, et al.
Published: (2025)
Revealing User Familiarity Bias in Task-Oriented Dialogue via Interactive Evaluation
by: Kim, Takyoung, et al.
Published: (2023)
by: Kim, Takyoung, et al.
Published: (2023)
Self-Correcting Code Generation Using Small Language Models
by: Cho, Jeonghun, et al.
Published: (2025)
by: Cho, Jeonghun, et al.
Published: (2025)
ArchCode: Incorporating Software Requirements in Code Generation with Large Language Models
by: Han, Hojae, et al.
Published: (2024)
by: Han, Hojae, et al.
Published: (2024)
Learning from Negative Samples in Biomedical Generative Entity Linking
by: Kim, Chanhwi, et al.
Published: (2024)
by: Kim, Chanhwi, et al.
Published: (2024)
Think, Verbalize, then Speak: Bridging Complex Thoughts and Comprehensible Speech
by: Woo, Sang Hoon, et al.
Published: (2025)
by: Woo, Sang Hoon, et al.
Published: (2025)
Similar Items
-
TimeChara: Evaluating Point-in-Time Character Hallucination of Role-Playing Large Language Models
by: Ahn, Jaewoo, et al.
Published: (2024) -
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs
by: Yoo, Haneul, et al.
Published: (2024) -
Is a Peeled Apple Still Red? Evaluating LLMs' Ability for Conceptual Combination with Property Type
by: Song, Seokwon, et al.
Published: (2025) -
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
by: Kim, Seungone, et al.
Published: (2023) -
Calibrating Large Language Models Using Their Generations Only
by: Ulmer, Dennis, et al.
Published: (2024)