Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Ziling, Cao, Meng, Rondeau, Marc-Antoine, Cheung, Jackie Chi Kit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation
by: Cheng, Ziling, et al.
Published: (2025)
by: Cheng, Ziling, et al.
Published: (2025)
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
CItruS: Chunked Instruction-aware State Eviction for Long Sequence Modeling
by: Bai, Yu, et al.
Published: (2024)
by: Bai, Yu, et al.
Published: (2024)
Identifying and Analyzing Performance-Critical Tokens in Large Language Models
by: Bai, Yu, et al.
Published: (2024)
by: Bai, Yu, et al.
Published: (2024)
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
by: Porada, Ian, et al.
Published: (2024)
by: Porada, Ian, et al.
Published: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
by: Chehbouni, Khaoula, et al.
Published: (2025)
by: Chehbouni, Khaoula, et al.
Published: (2025)
A Controlled Reevaluation of Coreference Resolution Models
by: Porada, Ian, et al.
Published: (2024)
by: Porada, Ian, et al.
Published: (2024)
PreSumm: Predicting Summarization Performance Without Summarizing
by: Koniaev, Steven, et al.
Published: (2025)
by: Koniaev, Steven, et al.
Published: (2025)
Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2025)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2025)
Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective
by: Porada, Ian, et al.
Published: (2023)
by: Porada, Ian, et al.
Published: (2023)
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
by: Darrin, Maxime, et al.
Published: (2024)
by: Darrin, Maxime, et al.
Published: (2024)
Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
$(RSA)^2$: A Rhetorical-Strategy-Aware Rational Speech Act Framework for Figurative Language Understanding
by: Piano, Cesare Spinoso-Di, et al.
Published: (2025)
by: Piano, Cesare Spinoso-Di, et al.
Published: (2025)
Unveiling LLMs' Metaphorical Understanding: Exploring Conceptual Irrelevance, Context Leveraging and Syntactic Influence
by: Ye, Fengying, et al.
Published: (2025)
by: Ye, Fengying, et al.
Published: (2025)
Do LLMs Know Tool Irrelevance? Demystifying Structural Alignment Bias in Tool Invocations
by: Liu, Yilong, et al.
Published: (2026)
by: Liu, Yilong, et al.
Published: (2026)
Can Vision Language Models Be Adaptive in Mathematics Education? A Learner Model-based Rubric Study
by: Gao, Jie, et al.
Published: (2026)
by: Gao, Jie, et al.
Published: (2026)
ECBD: Evidence-Centered Benchmark Design for NLP
by: Liu, Yu Lu, et al.
Published: (2024)
by: Liu, Yu Lu, et al.
Published: (2024)
Testing the Assumptions of Active Learning for Translation Tasks with Few Samples
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
Instructing Large Language Models to Identify and Ignore Irrelevant Conditions
by: Wu, Zhenyu, et al.
Published: (2024)
by: Wu, Zhenyu, et al.
Published: (2024)
Error Diversity Matters: An Error-Resistant Ensemble Method for Unsupervised Dependency Parsing
by: Shayegh, Behzad, et al.
Published: (2024)
by: Shayegh, Behzad, et al.
Published: (2024)
Making Retrieval-Augmented Language Models Robust to Irrelevant Context
by: Yoran, Ori, et al.
Published: (2023)
by: Yoran, Ori, et al.
Published: (2023)
Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning
by: Liao, Yiming, et al.
Published: (2026)
by: Liao, Yiming, et al.
Published: (2026)
Rowen: Adaptive Retrieval-Augmented Generation for Hallucination Mitigation in LLMs
by: Ding, Hanxing, et al.
Published: (2024)
by: Ding, Hanxing, et al.
Published: (2024)
Hallucination Detection with Small Language Models
by: Cheung, Ming
Published: (2025)
by: Cheung, Ming
Published: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
Fine-Grained Detection of Context-Grounded Hallucinations Using LLMs
by: Peisakhovsky, Yehonatan, et al.
Published: (2025)
by: Peisakhovsky, Yehonatan, et al.
Published: (2025)
Evaluating LLMs' Assessment of Mixed-Context Hallucination Through the Lens of Summarization
by: Qi, Siya, et al.
Published: (2025)
by: Qi, Siya, et al.
Published: (2025)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
by: Lucchetti, Francesca, et al.
Published: (2024)
by: Lucchetti, Francesca, et al.
Published: (2024)
Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge Conflicts
by: Xie, Jian, et al.
Published: (2023)
by: Xie, Jian, et al.
Published: (2023)
Are Language Models Sensitive to Morally Irrelevant Distractors?
by: Shaw, Andrew, et al.
Published: (2026)
by: Shaw, Andrew, et al.
Published: (2026)
Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
by: Cohen-Inger, Nurit, et al.
Published: (2025)
by: Cohen-Inger, Nurit, et al.
Published: (2025)
Few-shot Personalization of LLMs with Mis-aligned Responses
by: Kim, Jaehyung, et al.
Published: (2024)
by: Kim, Jaehyung, et al.
Published: (2024)
Shakespearean Sparks: The Dance of Hallucination and Creativity in LLMs' Decoding Layers
by: He, Zicong, et al.
Published: (2025)
by: He, Zicong, et al.
Published: (2025)
Look Within, Why LLMs Hallucinate: A Causal Perspective
by: Li, He, et al.
Published: (2024)
by: Li, He, et al.
Published: (2024)
Does Less Hallucination Mean Less Creativity? An Empirical Investigation in LLMs
by: Banerjee, Mohor, et al.
Published: (2025)
by: Banerjee, Mohor, et al.
Published: (2025)
Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and Disinformation
by: Yuan, Xin, et al.
Published: (2023)
by: Yuan, Xin, et al.
Published: (2023)
How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
Sensivity of LLMs' Explanations to the Training Randomness:Context, Class & Task Dependencies
by: Loncour, Romain, et al.
Published: (2026)
by: Loncour, Romain, et al.
Published: (2026)
GLIMPSE: Pragmatically Informative Multi-Document Summarization for Scholarly Reviews
by: Darrin, Maxime, et al.
Published: (2024)
by: Darrin, Maxime, et al.
Published: (2024)
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
by: Wu, Jiarong, et al.
Published: (2025)
by: Wu, Jiarong, et al.
Published: (2025)
Similar Items
-
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation
by: Cheng, Ziling, et al.
Published: (2025) -
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
by: Yu, Lei, et al.
Published: (2024) -
CItruS: Chunked Instruction-aware State Eviction for Long Sequence Modeling
by: Bai, Yu, et al.
Published: (2024) -
Identifying and Analyzing Performance-Critical Tokens in Large Language Models
by: Bai, Yu, et al.
Published: (2024) -
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
by: Porada, Ian, et al.
Published: (2024)