Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiao, Pengkun, Zhu, Bin, Chen, Jingjing, Ngo, Chong-Wah, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models
von: Tang, Ziyao, et al.
Veröffentlicht: (2026)
von: Tang, Ziyao, et al.
Veröffentlicht: (2026)
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
Benchmarking Gaslighting Negation Attacks Against Reasoning Models
von: Zhu, Bin, et al.
Veröffentlicht: (2025)
von: Zhu, Bin, et al.
Veröffentlicht: (2025)
Retrieval Augmented Recipe Generation
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
FoodLMM: A Versatile Food Assistant using Large Multi-modal Model
von: Yin, Yuehao, et al.
Veröffentlicht: (2023)
von: Yin, Yuehao, et al.
Veröffentlicht: (2023)
Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs
von: Tu, Chongjun, et al.
Veröffentlicht: (2025)
von: Tu, Chongjun, et al.
Veröffentlicht: (2025)
Domain Expansion and Boundary Growth for Open-Set Single-Source Domain Generalization
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
Unlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
Advancing Food Nutrition Estimation via Visual-Ingredient Feature Fusion
von: Qi, Huiyan, et al.
Veröffentlicht: (2025)
von: Qi, Huiyan, et al.
Veröffentlicht: (2025)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Interpretable Embedding for Ad-hoc Video Search
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
Improving Interpretable Embeddings for Ad-hoc Video Search with Generative Captions and Multi-word Concept Bank
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
von: Zhao, Kairan, et al.
Veröffentlicht: (2026)
von: Zhao, Kairan, et al.
Veröffentlicht: (2026)
From Canteen Food to Daily Meals: Generalizing Food Recognition to More Practical Scenarios
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
von: Han, Feng, et al.
Veröffentlicht: (2025)
von: Han, Feng, et al.
Veröffentlicht: (2025)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
von: Li, Yian, et al.
Veröffentlicht: (2026)
von: Li, Yian, et al.
Veröffentlicht: (2026)
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
von: Meng, Lingchen, et al.
Veröffentlicht: (2024)
von: Meng, Lingchen, et al.
Veröffentlicht: (2024)
LLMs-based Augmentation for Domain Adaptation in Long-tailed Food Datasets
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
CookingDiffusion: Cooking Procedural Image Generation with Stable Diffusion
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
EAGLE: Towards Efficient Arbitrary Referring Visual Prompts Comprehension for Multimodal Large Language Models
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
PosMLP-Video: Spatial and Temporal Relative Position Encoding for Efficient Video Recognition
von: Hao, Yanbin, et al.
Veröffentlicht: (2024)
von: Hao, Yanbin, et al.
Veröffentlicht: (2024)
Vision Transformers Don't Need Trained Registers
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
von: Wu, Xiongwei, et al.
Veröffentlicht: (2024)
von: Wu, Xiongwei, et al.
Veröffentlicht: (2024)
Don't Look into the Dark: Latent Codes for Pluralistic Image Inpainting
von: Chen, Haiwei, et al.
Veröffentlicht: (2024)
von: Chen, Haiwei, et al.
Veröffentlicht: (2024)
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2026)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2026)
Now You See Me, Now You Don't: A Unified Framework for Expression Consistent Anonymization in Talking Head Videos
von: Egin, Anil, et al.
Veröffentlicht: (2026)
von: Egin, Anil, et al.
Veröffentlicht: (2026)
NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
von: Qian, Tianwen, et al.
Veröffentlicht: (2023)
von: Qian, Tianwen, et al.
Veröffentlicht: (2023)
ReToMe-VA: Recursive Token Merging for Video Diffusion-based Unrestricted Adversarial Attack
von: Gao, Ziyi, et al.
Veröffentlicht: (2024)
von: Gao, Ziyi, et al.
Veröffentlicht: (2024)
NeIn: Telling What You Don't Want
von: Bui, Nhat-Tan, et al.
Veröffentlicht: (2024)
von: Bui, Nhat-Tan, et al.
Veröffentlicht: (2024)
Don't Get Me Wrong: How to Apply Deep Visual Interpretations to Time Series
von: Loeffler, Christoffer, et al.
Veröffentlicht: (2022)
von: Loeffler, Christoffer, et al.
Veröffentlicht: (2022)
Don't let the information slip away
von: Li, Taozhe, et al.
Veröffentlicht: (2026)
von: Li, Taozhe, et al.
Veröffentlicht: (2026)
Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flow
von: Zhao, Jianfei, et al.
Veröffentlicht: (2025)
von: Zhao, Jianfei, et al.
Veröffentlicht: (2025)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models
von: Jiao, Yang, et al.
Veröffentlicht: (2024)
von: Jiao, Yang, et al.
Veröffentlicht: (2024)
MIBench: Evaluating LMMs on Multimodal Interaction
von: Miao, Yu, et al.
Veröffentlicht: (2026)
von: Miao, Yu, et al.
Veröffentlicht: (2026)
Adversarial Magnification to Deceive Deepfake Detection through Super Resolution
von: Coccomini, Davide Alessandro, et al.
Veröffentlicht: (2024)
von: Coccomini, Davide Alessandro, et al.
Veröffentlicht: (2024)
RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation
von: Zhou, Shijie, et al.
Veröffentlicht: (2026)
von: Zhou, Shijie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models
von: Tang, Ziyao, et al.
Veröffentlicht: (2026) -
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024) -
RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024) -
Benchmarking Gaslighting Negation Attacks Against Reasoning Models
von: Zhu, Bin, et al.
Veröffentlicht: (2025) -
Retrieval Augmented Recipe Generation
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)