Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Sewon, Kim, Jiwon, Shin, Seungwoo, Chung, Hyejin, Moon, Daeun, Kwon, Yejin, Yoon, Hyunsoo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LogicQA: Logical Anomaly Detection with Vision Language Model Generated Questions
di: Kwon, Yejin, et al.
Pubblicazione: (2025)
di: Kwon, Yejin, et al.
Pubblicazione: (2025)
M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models
di: Kwon, Yejin, et al.
Pubblicazione: (2025)
di: Kwon, Yejin, et al.
Pubblicazione: (2025)
Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detection
di: Lee, Yujin, et al.
Pubblicazione: (2024)
di: Lee, Yujin, et al.
Pubblicazione: (2024)
Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate
di: Wynn, Andrea, et al.
Pubblicazione: (2025)
di: Wynn, Andrea, et al.
Pubblicazione: (2025)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
di: Barkett, Emilio, et al.
Pubblicazione: (2025)
di: Barkett, Emilio, et al.
Pubblicazione: (2025)
Inverse Scaling: When Bigger Isn't Better
di: McKenzie, Ian R., et al.
Pubblicazione: (2023)
di: McKenzie, Ian R., et al.
Pubblicazione: (2023)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
di: Xu, Xiaoyu, et al.
Pubblicazione: (2025)
di: Xu, Xiaoyu, et al.
Pubblicazione: (2025)
Word Boundary Information Isn't Useful for Encoder Language Models
di: Gow-Smith, Edward, et al.
Pubblicazione: (2024)
di: Gow-Smith, Edward, et al.
Pubblicazione: (2024)
Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive Diagnosis
di: Long, Zhuohan, et al.
Pubblicazione: (2026)
di: Long, Zhuohan, et al.
Pubblicazione: (2026)
When Meaning Isn't Literal: Exploring Idiomatic Meaning Across Languages and Modalities
di: Das, Sarmistha, et al.
Pubblicazione: (2026)
di: Das, Sarmistha, et al.
Pubblicazione: (2026)
Mathematics Isn't Culture-Free: Probing Cultural Gaps via Entity and Scenario Perturbations
di: Tomar, Aditya, et al.
Pubblicazione: (2025)
di: Tomar, Aditya, et al.
Pubblicazione: (2025)
SAIE Framework: Support Alone Isn't Enough -- Advancing LLM Training with Adversarial Remarks
di: Loem, Mengsay, et al.
Pubblicazione: (2023)
di: Loem, Mengsay, et al.
Pubblicazione: (2023)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
di: Khan, Mohammed Safi Ur Rahman, et al.
Pubblicazione: (2026)
di: Khan, Mohammed Safi Ur Rahman, et al.
Pubblicazione: (2026)
When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning
di: Barale, Claire, et al.
Pubblicazione: (2025)
di: Barale, Claire, et al.
Pubblicazione: (2025)
Recall Isn't Enough: Bounding Commitments in Personalized Language Systems
di: Tang, Rui, et al.
Pubblicazione: (2026)
di: Tang, Rui, et al.
Pubblicazione: (2026)
Latent Preference Modeling for Cross-Session Personalized Tool Calling
di: Yoon, Yejin, et al.
Pubblicazione: (2026)
di: Yoon, Yejin, et al.
Pubblicazione: (2026)
Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding
di: Kim, Sungkyun, et al.
Pubblicazione: (2025)
di: Kim, Sungkyun, et al.
Pubblicazione: (2025)
When Bigger Isn't Better: A Comprehensive Fairness Evaluation of Political Bias in Multi-News Summarisation
di: Huang, Nannan, et al.
Pubblicazione: (2026)
di: Huang, Nannan, et al.
Pubblicazione: (2026)
Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form Discourse
di: Tsipidi, Eleftheria, et al.
Pubblicazione: (2024)
di: Tsipidi, Eleftheria, et al.
Pubblicazione: (2024)
"Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak
di: Mei, Lingrui, et al.
Pubblicazione: (2024)
di: Mei, Lingrui, et al.
Pubblicazione: (2024)
Seeing Isn't Believing: Mitigating Belief Inertia via Active Intervention in Embodied Agents
di: Wang, Hanlin, et al.
Pubblicazione: (2026)
di: Wang, Hanlin, et al.
Pubblicazione: (2026)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
di: Zhang, Yue, et al.
Pubblicazione: (2026)
di: Zhang, Yue, et al.
Pubblicazione: (2026)
Why Synthetic Isn't Real Yet: A Diagnostic Framework for Contact Center Dialogue Generation
di: Devanathan, Rishikesh, et al.
Pubblicazione: (2025)
di: Devanathan, Rishikesh, et al.
Pubblicazione: (2025)
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
di: Galeone, Cosimo, et al.
Pubblicazione: (2026)
di: Galeone, Cosimo, et al.
Pubblicazione: (2026)
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
di: Kim, Jongsuk, et al.
Pubblicazione: (2024)
di: Kim, Jongsuk, et al.
Pubblicazione: (2024)
Are Today's LLMs Ready to Explain Well-Being Concepts?
di: Jiang, Bohan, et al.
Pubblicazione: (2025)
di: Jiang, Bohan, et al.
Pubblicazione: (2025)
Predicting Psychological Well-Being from Spontaneous Speech using LLMs
di: Loweimi, Erfan, et al.
Pubblicazione: (2026)
di: Loweimi, Erfan, et al.
Pubblicazione: (2026)
More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
di: Tsuchiya, Yuta, et al.
Pubblicazione: (2026)
di: Tsuchiya, Yuta, et al.
Pubblicazione: (2026)
Truth-Aware Context Selection: Mitigating Hallucinations of Large Language Models Being Misled by Untruthful Contexts
di: Yu, Tian, et al.
Pubblicazione: (2024)
di: Yu, Tian, et al.
Pubblicazione: (2024)
BlendX: Complex Multi-Intent Detection with Blended Patterns
di: Yoon, Yejin, et al.
Pubblicazione: (2024)
di: Yoon, Yejin, et al.
Pubblicazione: (2024)
Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
di: Sun, Yiyou, et al.
Pubblicazione: (2025)
di: Sun, Yiyou, et al.
Pubblicazione: (2025)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
di: Raval, Shivam, et al.
Pubblicazione: (2026)
di: Raval, Shivam, et al.
Pubblicazione: (2026)
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
di: Song, Jiwon, et al.
Pubblicazione: (2024)
di: Song, Jiwon, et al.
Pubblicazione: (2024)
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
di: Phute, Mansi, et al.
Pubblicazione: (2023)
di: Phute, Mansi, et al.
Pubblicazione: (2023)
The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead?
di: Choi, Alexander S., et al.
Pubblicazione: (2024)
di: Choi, Alexander S., et al.
Pubblicazione: (2024)
What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection
di: Nguyen, Binh, et al.
Pubblicazione: (2025)
di: Nguyen, Binh, et al.
Pubblicazione: (2025)
Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer
di: Yeom, Jewon, et al.
Pubblicazione: (2026)
di: Yeom, Jewon, et al.
Pubblicazione: (2026)
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking
di: LeVine, Will, et al.
Pubblicazione: (2025)
di: LeVine, Will, et al.
Pubblicazione: (2025)
ML Interpretability: Simple Isn't Easy
di: Räz, Tim
Pubblicazione: (2022)
di: Räz, Tim
Pubblicazione: (2022)
I'm Fine, But My Voice Isn't: Cross-Modal Affective Dissonance Detection for Reflective Journaling
di: Lee, Sumin
Pubblicazione: (2026)
di: Lee, Sumin
Pubblicazione: (2026)
Documenti analoghi
-
LogicQA: Logical Anomaly Detection with Vision Language Model Generated Questions
di: Kwon, Yejin, et al.
Pubblicazione: (2025) -
M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models
di: Kwon, Yejin, et al.
Pubblicazione: (2025) -
Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detection
di: Lee, Yujin, et al.
Pubblicazione: (2024) -
Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate
di: Wynn, Andrea, et al.
Pubblicazione: (2025) -
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
di: Barkett, Emilio, et al.
Pubblicazione: (2025)