MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bae, Kyungho, Kim, Jinhyung, Lee, Sihaeng, Lee, Soonyoung, Lee, Gunhee, Choi, Jinwoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DEVIAS: Learning Disentangled Video Representations of Action and Scene
von: Bae, Kyungho, et al.
Veröffentlicht: (2023)
von: Bae, Kyungho, et al.
Veröffentlicht: (2023)
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
von: Kim, Jisoo, et al.
Veröffentlicht: (2026)
von: Kim, Jisoo, et al.
Veröffentlicht: (2026)
ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams
von: Kim, Chris Dongjoo, et al.
Veröffentlicht: (2025)
von: Kim, Chris Dongjoo, et al.
Veröffentlicht: (2025)
ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
CAST: Cross-Attention in Space and Time for Video Action Recognition
von: Lee, Dongho, et al.
Veröffentlicht: (2023)
von: Lee, Dongho, et al.
Veröffentlicht: (2023)
Bi-directional Contextual Attention for 3D Dense Captioning
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
EXAONEPath 1.0 Patch-level Foundation Model for Pathology
von: Yun, Juseung, et al.
Veröffentlicht: (2024)
von: Yun, Juseung, et al.
Veröffentlicht: (2024)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs
von: Lee, Jongseo, et al.
Veröffentlicht: (2026)
von: Lee, Jongseo, et al.
Veröffentlicht: (2026)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
von: Lee, Kyuho, et al.
Veröffentlicht: (2025)
von: Lee, Kyuho, et al.
Veröffentlicht: (2025)
Fourier-Guided Attention Upsampling for Image Super-Resolution
von: Choi, Daejune, et al.
Veröffentlicht: (2025)
von: Choi, Daejune, et al.
Veröffentlicht: (2025)
Discovering and Mitigating Visual Biases through Keyword Explanation
von: Kim, Younghyun, et al.
Veröffentlicht: (2023)
von: Kim, Younghyun, et al.
Veröffentlicht: (2023)
See It All: Contextualized Late Aggregation for 3D Dense Captioning
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
von: Kim, Minjung, et al.
Veröffentlicht: (2024)
PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2026)
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2026)
PCBEAR: Pose Concept Bottleneck for Explainable Action Recognition
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition
von: Oh, Youngtaek, et al.
Veröffentlicht: (2024)
von: Oh, Youngtaek, et al.
Veröffentlicht: (2024)
Can Language Models Laugh at YouTube Short-form Videos?
von: Ko, Dayoon, et al.
Veröffentlicht: (2023)
von: Ko, Dayoon, et al.
Veröffentlicht: (2023)
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
IB-GAN: Disentangled Representation Learning with Information Bottleneck Generative Adversarial Networks
von: Jeon, Insu, et al.
Veröffentlicht: (2025)
von: Jeon, Insu, et al.
Veröffentlicht: (2025)
GaussianVideo: Efficient Video Representation and Compression by Gaussian Splatting
von: Lee, Inseo, et al.
Veröffentlicht: (2025)
von: Lee, Inseo, et al.
Veröffentlicht: (2025)
ChatEXAONEPath: An Expert-level Multimodal Large Language Model for Histopathology Using Whole Slide Images
von: Kim, Sangwook, et al.
Veröffentlicht: (2025)
von: Kim, Sangwook, et al.
Veröffentlicht: (2025)
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
von: Baik, Sangwon, et al.
Veröffentlicht: (2026)
von: Baik, Sangwon, et al.
Veröffentlicht: (2026)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
von: Cai, Jianfeng, et al.
Veröffentlicht: (2025)
von: Cai, Jianfeng, et al.
Veröffentlicht: (2025)
EO-VLM: VLM-Guided Energy Overload Attacks on Vision Models
von: Seo, Minjae, et al.
Veröffentlicht: (2025)
von: Seo, Minjae, et al.
Veröffentlicht: (2025)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
von: Fateh, Fawad Javed, et al.
Veröffentlicht: (2024)
von: Fateh, Fawad Javed, et al.
Veröffentlicht: (2024)
Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance
von: Choi, June Suk, et al.
Veröffentlicht: (2025)
von: Choi, June Suk, et al.
Veröffentlicht: (2025)
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
von: Han, Jiwook, et al.
Veröffentlicht: (2026)
von: Han, Jiwook, et al.
Veröffentlicht: (2026)
HalLoc: Token-level Localization of Hallucinations for Vision Language Models
von: Park, Eunkyu, et al.
Veröffentlicht: (2025)
von: Park, Eunkyu, et al.
Veröffentlicht: (2025)
Towards Temporal Fusion Beyond the Field of View for Camera-based Semantic Scene Completion
von: Bae, Jongseong, et al.
Veröffentlicht: (2025)
von: Bae, Jongseong, et al.
Veröffentlicht: (2025)
JARViS: Detecting Actions in Video Using Unified Actor-Scene Context Relation Modeling
von: Lee, Seok Hwan, et al.
Veröffentlicht: (2024)
von: Lee, Seok Hwan, et al.
Veröffentlicht: (2024)
Isometric Representation Learning for Disentangled Latent Space of Diffusion Models
von: Hahm, Jaehoon, et al.
Veröffentlicht: (2024)
von: Hahm, Jaehoon, et al.
Veröffentlicht: (2024)
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
von: Jeon, Byungwoo, et al.
Veröffentlicht: (2026)
von: Jeon, Byungwoo, et al.
Veröffentlicht: (2026)
ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
Infusing Environmental Captions for Long-Form Video Language Grounding
von: Lee, Hyogun, et al.
Veröffentlicht: (2024)
von: Lee, Hyogun, et al.
Veröffentlicht: (2024)
SpatialMosaic: A Multiview VLM Dataset for Partial Visibility
von: Lee, Kanghee, et al.
Veröffentlicht: (2025)
von: Lee, Kanghee, et al.
Veröffentlicht: (2025)
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DEVIAS: Learning Disentangled Video Representations of Action and Scene
von: Bae, Kyungho, et al.
Veröffentlicht: (2023) -
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
von: Kim, Jisoo, et al.
Veröffentlicht: (2026) -
ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams
von: Kim, Chris Dongjoo, et al.
Veröffentlicht: (2025) -
ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
von: Lee, Jongseo, et al.
Veröffentlicht: (2025) -
Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)