LLMs can be easily Confused by Instructional Distractions
Fuente:
arXiv
Saved in:
| Main Authors: | Hwang, Yerin, Kim, Yongil, Koo, Jahyun, Kang, Taegwan, Bae, Hyunkyung, Jung, Kyomin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models
by: Koo, Jahyun, et al.
Published: (2024)
by: Koo, Jahyun, et al.
Published: (2024)
MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs
by: Hwang, Yerin, et al.
Published: (2024)
by: Hwang, Yerin, et al.
Published: (2024)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
When Wording Steers the Evaluation: Framing Bias in LLM judges
by: Hwang, Yerin, et al.
Published: (2026)
by: Hwang, Yerin, et al.
Published: (2026)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
by: Moon, Jiwon, et al.
Published: (2025)
by: Moon, Jiwon, et al.
Published: (2025)
Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
by: Joo, Seongho, et al.
Published: (2025)
by: Joo, Seongho, et al.
Published: (2025)
Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
by: Yang, Nakyeong, et al.
Published: (2023)
by: Yang, Nakyeong, et al.
Published: (2023)
Program Synthesis via Test-Time Transduction
by: Lee, Kang-il, et al.
Published: (2025)
by: Lee, Kang-il, et al.
Published: (2025)
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
by: Lee, Dongryeol, et al.
Published: (2024)
by: Lee, Dongryeol, et al.
Published: (2024)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
by: Lee, Dongryeol, et al.
Published: (2026)
by: Lee, Dongryeol, et al.
Published: (2026)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Uncovering the Potential Risks in Unlearning: Danger of English-only Unlearning in Multilingual LLMs
by: Hwang, Kyomin, et al.
Published: (2025)
by: Hwang, Kyomin, et al.
Published: (2025)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
by: Joo, Seongho, et al.
Published: (2025)
by: Joo, Seongho, et al.
Published: (2025)
Do not think about pink elephant!
by: Hwang, Kyomin, et al.
Published: (2024)
by: Hwang, Kyomin, et al.
Published: (2024)
Avoidance Decoding for Diverse Multi-Branch Story Generation
by: Park, Kyeongman, et al.
Published: (2025)
by: Park, Kyeongman, et al.
Published: (2025)
Public Data Assisted Differentially Private In-Context Learning
by: Joo, Seongho, et al.
Published: (2025)
by: Joo, Seongho, et al.
Published: (2025)
LongStory: Coherent, Complete and Length Controlled Long story Generation
by: Park, Kyeongman, et al.
Published: (2023)
by: Park, Kyeongman, et al.
Published: (2023)
BiCon-Gate: Consistency-Gated De-colloquialisation for Dialogue Fact-Checking
by: Park, Hyunkyung, et al.
Published: (2026)
by: Park, Hyunkyung, et al.
Published: (2026)
Optimal path for Biomedical Text Summarization Using Pointer GPT
by: Han, Hyunkyung, et al.
Published: (2024)
by: Han, Hyunkyung, et al.
Published: (2024)
LifeTox: Unveiling Implicit Toxicity in Life Advice
by: Kim, Minbeom, et al.
Published: (2023)
by: Kim, Minbeom, et al.
Published: (2023)
BenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMs
by: Yoon, Sangyeon, et al.
Published: (2026)
by: Yoon, Sangyeon, et al.
Published: (2026)
FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledge
by: Yang, Nakyeong, et al.
Published: (2025)
by: Yang, Nakyeong, et al.
Published: (2025)
Conditional [MASK] Discrete Diffusion Language Model
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
Unplug and Play Language Models: Decomposing Experts in Language Models at Inference Time
by: Yang, Nakyeong, et al.
Published: (2024)
by: Yang, Nakyeong, et al.
Published: (2024)
Format Inertia: A Failure Mechanism of LLMs in Medical Pre-Consultation
by: Lim, Seungseop, et al.
Published: (2025)
by: Lim, Seungseop, et al.
Published: (2025)
Reducing Peak Memory Usage for Modern Multimodal Large Language Model Pipelines
by: Kim, Junwan, et al.
Published: (2026)
by: Kim, Junwan, et al.
Published: (2026)
EMCEE: Improving Multilingual Capability of LLMs via Bridging Knowledge and Reasoning with Extracted Synthetic Multilingual Context
by: Koo, Hamin, et al.
Published: (2025)
by: Koo, Hamin, et al.
Published: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
by: Min, Kyungmin, et al.
Published: (2024)
by: Min, Kyungmin, et al.
Published: (2024)
Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines
by: Seo, Jean, et al.
Published: (2026)
by: Seo, Jean, et al.
Published: (2026)
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
by: Kim, Minsung, et al.
Published: (2025)
by: Kim, Minsung, et al.
Published: (2025)
SelectLLM: Can LLMs Select Important Instructions to Annotate?
by: Parkar, Ritik Sachin, et al.
Published: (2024)
by: Parkar, Ritik Sachin, et al.
Published: (2024)
Medical large language models are easily distracted
by: Vishwanath, Krithik, et al.
Published: (2025)
by: Vishwanath, Krithik, et al.
Published: (2025)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
by: Lee, Kang-il, et al.
Published: (2024)
by: Lee, Kang-il, et al.
Published: (2024)
MVMR: A New Framework for Evaluating Faithfulness of Video Moment Retrieval against Multiple Distractors
by: Yang, Nakyeong, et al.
Published: (2023)
by: Yang, Nakyeong, et al.
Published: (2023)
Draft-based Approximate Inference for LLMs
by: Galim, Kevin, et al.
Published: (2025)
by: Galim, Kevin, et al.
Published: (2025)
Reasoning Models Better Express Their Confidence
by: Yoon, Dongkeun, et al.
Published: (2025)
by: Yoon, Dongkeun, et al.
Published: (2025)
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
by: Kim, Jeonghye, et al.
Published: (2025)
by: Kim, Jeonghye, et al.
Published: (2025)
Contrastive and Consistency Learning for Neural Noisy-Channel Model in Spoken Language Understanding
by: Kim, Suyoung, et al.
Published: (2024)
by: Kim, Suyoung, et al.
Published: (2024)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
Similar Items
-
SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models
by: Koo, Jahyun, et al.
Published: (2024) -
MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs
by: Hwang, Yerin, et al.
Published: (2024) -
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
by: Hwang, Yerin, et al.
Published: (2025) -
When Wording Steers the Evaluation: Framing Bias in LLM judges
by: Hwang, Yerin, et al.
Published: (2026) -
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
by: Moon, Jiwon, et al.
Published: (2025)