Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Chengsheng, Sun, Chenghao, Xie, Zhining, Tian, Xinmei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models
von: Zhang, Chengsheng, et al.
Veröffentlicht: (2026)
von: Zhang, Chengsheng, et al.
Veröffentlicht: (2026)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
von: Che, Liwei, et al.
Veröffentlicht: (2026)
von: Che, Liwei, et al.
Veröffentlicht: (2026)
GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models
von: Shen, Guanxi
Veröffentlicht: (2025)
von: Shen, Guanxi
Veröffentlicht: (2025)
Cross-Layer Vision Smoothing: Enhancing Visual Understanding via Sustained Focus on Key Objects in Large Vision-Language Models
von: Zhao, Jianfei, et al.
Veröffentlicht: (2025)
von: Zhao, Jianfei, et al.
Veröffentlicht: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation
von: Guo, Shasha, et al.
Veröffentlicht: (2025)
von: Guo, Shasha, et al.
Veröffentlicht: (2025)
AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
von: Xu, Shixiong, et al.
Veröffentlicht: (2025)
von: Xu, Shixiong, et al.
Veröffentlicht: (2025)
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models
von: Yan, Hanqi, et al.
Veröffentlicht: (2025)
von: Yan, Hanqi, et al.
Veröffentlicht: (2025)
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
von: Zhong, Yi, et al.
Veröffentlicht: (2026)
von: Zhong, Yi, et al.
Veröffentlicht: (2026)
Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
Large Vision-Language Models as Emotion Recognizers in Context Awareness
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
MBQ: Modality-Balanced Quantization for Large Vision-Language Models
von: Li, Shiyao, et al.
Veröffentlicht: (2024)
von: Li, Shiyao, et al.
Veröffentlicht: (2024)
Enhancing Interpretability for Vision Models via Shapley Value Optimization
von: Fan, Kanglong, et al.
Veröffentlicht: (2025)
von: Fan, Kanglong, et al.
Veröffentlicht: (2025)
SPARC: Concept-Aligned Sparse Autoencoders for Cross-Model and Cross-Modal Interpretability
von: Nasiri-Sarvi, Ali, et al.
Veröffentlicht: (2025)
von: Nasiri-Sarvi, Ali, et al.
Veröffentlicht: (2025)
Enhancing Medical Large Vision-Language Models via Alignment Distillation
von: Chang, Aofei, et al.
Veröffentlicht: (2025)
von: Chang, Aofei, et al.
Veröffentlicht: (2025)
GPT-4 Enhanced Multimodal Grounding for Autonomous Driving: Leveraging Cross-Modal Attention with Large Language Models
von: Liao, Haicheng, et al.
Veröffentlicht: (2023)
von: Liao, Haicheng, et al.
Veröffentlicht: (2023)
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model
von: Feng, Qianhan, et al.
Veröffentlicht: (2024)
von: Feng, Qianhan, et al.
Veröffentlicht: (2024)
Cross-modal Information Flow in Multimodal Large Language Models
von: Zhang, Zhi, et al.
Veröffentlicht: (2024)
von: Zhang, Zhi, et al.
Veröffentlicht: (2024)
PIP: Detecting Adversarial Examples in Large Vision-Language Models via Attention Patterns of Irrelevant Probe Questions
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
Cross-Modal Attention Analysis and Optimization in Vision-Language Models: A Study on Visual Reliability
von: Zhou, Lijie
Veröffentlicht: (2026)
von: Zhou, Lijie
Veröffentlicht: (2026)
CapRecover: A Cross-Modality Feature Inversion Attack Framework on Vision Language Models
von: Xiu, Kedong, et al.
Veröffentlicht: (2025)
von: Xiu, Kedong, et al.
Veröffentlicht: (2025)
QAVA: Query-Agnostic Visual Attack to Large Vision-Language Models
von: Zhang, Yudong, et al.
Veröffentlicht: (2025)
von: Zhang, Yudong, et al.
Veröffentlicht: (2025)
Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout
von: QI, Anbin, et al.
Veröffentlicht: (2024)
von: QI, Anbin, et al.
Veröffentlicht: (2024)
FlowLearn: Evaluating Large Vision-Language Models on Flowchart Understanding
von: Pan, Huitong, et al.
Veröffentlicht: (2024)
von: Pan, Huitong, et al.
Veröffentlicht: (2024)
Seeing is Believing: Robust Vision-Guided Cross-Modal Prompt Learning under Label Noise
von: Geng, Zibin, et al.
Veröffentlicht: (2026)
von: Geng, Zibin, et al.
Veröffentlicht: (2026)
Cross-Modal Adapter for Vision-Language Retrieval
von: Jiang, Haojun, et al.
Veröffentlicht: (2022)
von: Jiang, Haojun, et al.
Veröffentlicht: (2022)
ViSA-Enhanced Aerial VLN: A Visual-Spatial Reasoning Enhanced Framework for Aerial Vision-Language Navigation
von: Tong, Haoyu, et al.
Veröffentlicht: (2026)
von: Tong, Haoyu, et al.
Veröffentlicht: (2026)
AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrieval
von: Wang, Yihan, et al.
Veröffentlicht: (2026)
von: Wang, Yihan, et al.
Veröffentlicht: (2026)
FOCA: Frequency-Oriented Cross-Domain Forgery Detection, Localization and Explanation via Multi-Modal Large Language Model
von: Liu, Zhou, et al.
Veröffentlicht: (2026)
von: Liu, Zhou, et al.
Veröffentlicht: (2026)
FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language Models
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registration
von: Liu, Xuecong, et al.
Veröffentlicht: (2026)
von: Liu, Xuecong, et al.
Veröffentlicht: (2026)
Interpretable Debiasing of Vision-Language Models for Social Fairness
von: An, Na Min, et al.
Veröffentlicht: (2026)
von: An, Na Min, et al.
Veröffentlicht: (2026)
Large Vision-Language Models Get Lost in Attention
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
How Do Vision-Language Models Process Conflicting Information Across Modalities?
von: Hua, Tianze, et al.
Veröffentlicht: (2025)
von: Hua, Tianze, et al.
Veröffentlicht: (2025)
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
von: Yan, Bei, et al.
Veröffentlicht: (2024)
von: Yan, Bei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models
von: Zhang, Chengsheng, et al.
Veröffentlicht: (2026) -
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
von: Che, Liwei, et al.
Veröffentlicht: (2026) -
GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models
von: Shen, Guanxi
Veröffentlicht: (2025) -
Cross-Layer Vision Smoothing: Enhancing Visual Understanding via Sustained Focus on Key Objects in Large Vision-Language Models
von: Zhao, Jianfei, et al.
Veröffentlicht: (2025) -
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)