STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Fan, Linfeng, Tian, Yuan, Li, Ziwei, Lu, Zhiwu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
di: Fazli, Mehrdad, et al.
Pubblicazione: (2025)
di: Fazli, Mehrdad, et al.
Pubblicazione: (2025)
Mitigating Image Captioning Hallucinations in Vision-Language Models
di: Zhao, Fei, et al.
Pubblicazione: (2025)
di: Zhao, Fei, et al.
Pubblicazione: (2025)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
di: Song, Jiale, et al.
Pubblicazione: (2026)
di: Song, Jiale, et al.
Pubblicazione: (2026)
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
di: Fu, Yuhan, et al.
Pubblicazione: (2024)
di: Fu, Yuhan, et al.
Pubblicazione: (2024)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
di: Poppi, Tobia, et al.
Pubblicazione: (2026)
di: Poppi, Tobia, et al.
Pubblicazione: (2026)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
di: Wang, Xintong, et al.
Pubblicazione: (2024)
di: Wang, Xintong, et al.
Pubblicazione: (2024)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
di: Lan, Xiaohan, et al.
Pubblicazione: (2024)
di: Lan, Xiaohan, et al.
Pubblicazione: (2024)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
di: Ding, Peng, et al.
Pubblicazione: (2024)
di: Ding, Peng, et al.
Pubblicazione: (2024)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
di: Qu, Mengxue, et al.
Pubblicazione: (2024)
di: Qu, Mengxue, et al.
Pubblicazione: (2024)
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination
di: Chen, Yangneng, et al.
Pubblicazione: (2026)
di: Chen, Yangneng, et al.
Pubblicazione: (2026)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
di: Zhang, Pingping, et al.
Pubblicazione: (2024)
di: Zhang, Pingping, et al.
Pubblicazione: (2024)
Failures to Surface Harmful Contents in Video Large Language Models
di: Cao, Yuxin, et al.
Pubblicazione: (2025)
di: Cao, Yuxin, et al.
Pubblicazione: (2025)
Context-Enhanced Video Moment Retrieval with Large Language Models
di: Liu, Weijia, et al.
Pubblicazione: (2024)
di: Liu, Weijia, et al.
Pubblicazione: (2024)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
di: Xu, Yifang, et al.
Pubblicazione: (2025)
di: Xu, Yifang, et al.
Pubblicazione: (2025)
SMC++: Masked Learning of Unsupervised Video Semantic Compression
di: Tian, Yuan, et al.
Pubblicazione: (2024)
di: Tian, Yuan, et al.
Pubblicazione: (2024)
On the Audio Hallucinations in Large Audio-Video Language Models
di: Nishimura, Taichi, et al.
Pubblicazione: (2024)
di: Nishimura, Taichi, et al.
Pubblicazione: (2024)
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
di: He, Xin, et al.
Pubblicazione: (2024)
di: He, Xin, et al.
Pubblicazione: (2024)
OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment
di: Chen, Rongjun, et al.
Pubblicazione: (2025)
di: Chen, Rongjun, et al.
Pubblicazione: (2025)
Context and Pixel Aware Large Language Model for Video Quality Assessment
di: Wen, Wen, et al.
Pubblicazione: (2025)
di: Wen, Wen, et al.
Pubblicazione: (2025)
Modality-Aware Shot Relating and Comparing for Video Scene Detection
di: Tan, Jiawei, et al.
Pubblicazione: (2024)
di: Tan, Jiawei, et al.
Pubblicazione: (2024)
EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model
di: Li, Deng, et al.
Pubblicazione: (2024)
di: Li, Deng, et al.
Pubblicazione: (2024)
Scene Graph Generation with Role-Playing Large Language Models
di: Chen, Guikun, et al.
Pubblicazione: (2024)
di: Chen, Guikun, et al.
Pubblicazione: (2024)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
di: Meng, Jiahao, et al.
Pubblicazione: (2026)
di: Meng, Jiahao, et al.
Pubblicazione: (2026)
HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video Retrieval
di: Chen, Zhiwei, et al.
Pubblicazione: (2025)
di: Chen, Zhiwei, et al.
Pubblicazione: (2025)
OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
di: Hao, Jing, et al.
Pubblicazione: (2025)
di: Hao, Jing, et al.
Pubblicazione: (2025)
LAVA: Language Driven Scalable and Versatile Traffic Video Analytics
di: Yu, Yanrui, et al.
Pubblicazione: (2025)
di: Yu, Yanrui, et al.
Pubblicazione: (2025)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
di: Deng, Ailin, et al.
Pubblicazione: (2024)
di: Deng, Ailin, et al.
Pubblicazione: (2024)
MindCine: Multimodal EEG-to-Video Reconstruction with Large-Scale Pretrained Models
di: Zhou, Tian-Yi, et al.
Pubblicazione: (2026)
di: Zhou, Tian-Yi, et al.
Pubblicazione: (2026)
ChronusOmni: Improving Time Awareness of Omni Large Language Models
di: Chen, Yijing, et al.
Pubblicazione: (2025)
di: Chen, Yijing, et al.
Pubblicazione: (2025)
CPSL: Representing Volumetric Video via Content-Promoted Scene Layers
di: Hu, Kaiyuan, et al.
Pubblicazione: (2025)
di: Hu, Kaiyuan, et al.
Pubblicazione: (2025)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
di: Zhou, Sheng, et al.
Pubblicazione: (2025)
di: Zhou, Sheng, et al.
Pubblicazione: (2025)
Spatial-Aware Efficient Projector for MLLMs via Multi-Layer Feature Aggregation
di: Qian, Shun, et al.
Pubblicazione: (2024)
di: Qian, Shun, et al.
Pubblicazione: (2024)
POINTS1.5: Building a Vision-Language Model towards Real World Applications
di: Liu, Yuan, et al.
Pubblicazione: (2024)
di: Liu, Yuan, et al.
Pubblicazione: (2024)
Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product Retrieval
di: Hu, Xiaowan, et al.
Pubblicazione: (2024)
di: Hu, Xiaowan, et al.
Pubblicazione: (2024)
VideoSTF: Stress-Testing Output Repetition in Video Large Language Models
di: Cao, Yuxin, et al.
Pubblicazione: (2026)
di: Cao, Yuxin, et al.
Pubblicazione: (2026)
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
di: Lyu, Yibo, et al.
Pubblicazione: (2025)
di: Lyu, Yibo, et al.
Pubblicazione: (2025)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
di: Chow, Wei, et al.
Pubblicazione: (2024)
di: Chow, Wei, et al.
Pubblicazione: (2024)
P-GSVC: Layered Progressive 2D Gaussian Splatting for Scalable Image and Video
di: Wang, Longan, et al.
Pubblicazione: (2026)
di: Wang, Longan, et al.
Pubblicazione: (2026)
Unveiling Encoder-Free Vision-Language Models
di: Diao, Haiwen, et al.
Pubblicazione: (2024)
di: Diao, Haiwen, et al.
Pubblicazione: (2024)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
di: Yuan, Shenghai, et al.
Pubblicazione: (2024)
di: Yuan, Shenghai, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
di: Fazli, Mehrdad, et al.
Pubblicazione: (2025) -
Mitigating Image Captioning Hallucinations in Vision-Language Models
di: Zhao, Fei, et al.
Pubblicazione: (2025) -
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
di: Song, Jiale, et al.
Pubblicazione: (2026) -
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
di: Fu, Yuhan, et al.
Pubblicazione: (2024) -
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
di: Poppi, Tobia, et al.
Pubblicazione: (2026)