Saved in:
| Main Authors: | Lee, Yukyung, Lim, Yebin, Jung, Woojun, Choi, Wonjun, Yoon, Susik |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.19250 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-level Diagnosis and Evaluation for Robust Tabular Feature Engineering with Large Language Models
by: Lim, Yebin, et al.
Published: (2025)
by: Lim, Yebin, et al.
Published: (2025)
Segment-driven Structural Induction and Semantic Alignment for Heterogeneous Tabular Representation
by: Jung, Woojun, et al.
Published: (2026)
by: Jung, Woojun, et al.
Published: (2026)
CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models
by: Lee, Sangin, et al.
Published: (2026)
by: Lee, Sangin, et al.
Published: (2026)
Navigating the Path of Writing: Outline-guided Text Generation with Large Language Models
by: Lee, Yukyung, et al.
Published: (2024)
by: Lee, Yukyung, et al.
Published: (2024)
CIRF: Tokenizing Chain-of-Thoughts into Reusable Functional Units for Efficient Latent Reasoning in Large Language Models
by: Lee, Yukyung, et al.
Published: (2026)
by: Lee, Yukyung, et al.
Published: (2026)
RExBench: Can coding agents autonomously implement AI research extensions?
by: Edwards, Nicholas, et al.
Published: (2025)
by: Edwards, Nicholas, et al.
Published: (2025)
ELITE: Enhanced Language-Image Toxicity Evaluation for Safety
by: Lee, Wonjun, et al.
Published: (2025)
by: Lee, Wonjun, et al.
Published: (2025)
Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams
by: Kim, Jiyeon, et al.
Published: (2026)
by: Kim, Jiyeon, et al.
Published: (2026)
SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models
by: Jeong, Wonjun, et al.
Published: (2025)
by: Jeong, Wonjun, et al.
Published: (2025)
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
by: Lee, Yebin, et al.
Published: (2024)
by: Lee, Yebin, et al.
Published: (2024)
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
by: Yoon, Eunseop, et al.
Published: (2025)
by: Yoon, Eunseop, et al.
Published: (2025)
References Indeed Matter? Reference-Free Preference Optimization for Conversational Query Reformulation
by: Kim, Doyoung, et al.
Published: (2025)
by: Kim, Doyoung, et al.
Published: (2025)
ProtoMed-LLM: An Automatic Evaluation Framework for Large Language Models in Medical Protocol Formulation
by: Yi, Seungjun, et al.
Published: (2024)
by: Yi, Seungjun, et al.
Published: (2024)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
FLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language Models
by: Jung, Dahyun, et al.
Published: (2025)
by: Jung, Dahyun, et al.
Published: (2025)
Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues
by: Cho, Ye-eun, et al.
Published: (2025)
by: Cho, Ye-eun, et al.
Published: (2025)
Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action Localization
by: Lim, Geuntaek, et al.
Published: (2024)
by: Lim, Geuntaek, et al.
Published: (2024)
Can MLLMs Perform Text-to-Image In-Context Learning?
by: Zeng, Yuchen, et al.
Published: (2024)
by: Zeng, Yuchen, et al.
Published: (2024)
Draft-based Approximate Inference for LLMs
by: Galim, Kevin, et al.
Published: (2025)
by: Galim, Kevin, et al.
Published: (2025)
HerO at AVeriTeC: The Herd of Open Large Language Models for Verifying Real-World Claims
by: Yoon, Yejun, et al.
Published: (2024)
by: Yoon, Yejun, et al.
Published: (2024)
BridG MT: Enhancing LLMs' Machine Translation Capabilities with Sentence Bridging and Gradual MT
by: Choi, Seung-Woo, et al.
Published: (2024)
by: Choi, Seung-Woo, et al.
Published: (2024)
Saving the legacy of Hero Ibash: Evaluating Four Language Models for Aminoacian
by: Xiao, Yunze, et al.
Published: (2024)
by: Xiao, Yunze, et al.
Published: (2024)
StreamingThinker: Large Language Models Can Think While Reading
by: Tong, Junlong, et al.
Published: (2025)
by: Tong, Junlong, et al.
Published: (2025)
Bias in LLMs as Annotators: The Effect of Party Cues on Labelling Decision by Large Language Models
by: Vera, Sebastian Vallejo, et al.
Published: (2024)
by: Vera, Sebastian Vallejo, et al.
Published: (2024)
Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluation
by: Cheng, Jiahao, et al.
Published: (2025)
by: Cheng, Jiahao, et al.
Published: (2025)
Theme-Explanation Structure for Table Summarization using Large Language Models: A Case Study on Korean Tabular Data
by: Kwack, TaeYoon, et al.
Published: (2025)
by: Kwack, TaeYoon, et al.
Published: (2025)
KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness
by: Kim, Jinyoung, et al.
Published: (2026)
by: Kim, Jinyoung, et al.
Published: (2026)
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
by: Jung, Woojun, et al.
Published: (2025)
by: Jung, Woojun, et al.
Published: (2025)
Knowledge Integration Decay in Search-Augmented Reasoning of Large Language Models
by: Yu, Sangwon, et al.
Published: (2026)
by: Yu, Sangwon, et al.
Published: (2026)
QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering
by: Jung, Woojun, et al.
Published: (2026)
by: Jung, Woojun, et al.
Published: (2026)
Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
by: Su, Guinan, et al.
Published: (2026)
by: Su, Guinan, et al.
Published: (2026)
Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion
by: Yoon, Yejun, et al.
Published: (2025)
by: Yoon, Yejun, et al.
Published: (2025)
A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls
by: Shafayat, Sheikh, et al.
Published: (2024)
by: Shafayat, Sheikh, et al.
Published: (2024)
CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
by: Lee, Yukyung, et al.
Published: (2024)
by: Lee, Yukyung, et al.
Published: (2024)
CLARIFID: Improving Radiology Report Generation by Reinforcing Clinically Accurate Impressions and Enforcing Detailed Findings
by: Lee, Kyeongkyu, et al.
Published: (2025)
by: Lee, Kyeongkyu, et al.
Published: (2025)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
by: Son, Guijin, et al.
Published: (2023)
by: Son, Guijin, et al.
Published: (2023)
How Language Models Prioritize Contextual Grammatical Cues?
by: Amirzadeh, Hamidreza, et al.
Published: (2024)
by: Amirzadeh, Hamidreza, et al.
Published: (2024)
Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs
by: Wang, Siyuan, et al.
Published: (2024)
by: Wang, Siyuan, et al.
Published: (2024)
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
by: Lee, Wonjun, et al.
Published: (2026)
by: Lee, Wonjun, et al.
Published: (2026)
Similar Items
-
Multi-level Diagnosis and Evaluation for Robust Tabular Feature Engineering with Large Language Models
by: Lim, Yebin, et al.
Published: (2025) -
Segment-driven Structural Induction and Semantic Alignment for Heterogeneous Tabular Representation
by: Jung, Woojun, et al.
Published: (2026) -
CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models
by: Lee, Sangin, et al.
Published: (2026) -
Navigating the Path of Writing: Outline-guided Text Generation with Large Language Models
by: Lee, Yukyung, et al.
Published: (2024) -
CIRF: Tokenizing Chain-of-Thoughts into Reusable Functional Units for Efficient Latent Reasoning in Large Language Models
by: Lee, Yukyung, et al.
Published: (2026)