Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jongseo, Lee, Hyuntak, Kim, Sunghun, Kim, Sooa, Chung, Jihoon, Choi, Jinwoo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAST: Cross-Attention in Space and Time for Video Action Recognition
by: Lee, Dongho, et al.
Published: (2023)
by: Lee, Dongho, et al.
Published: (2023)
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
by: Lee, Jongseo, et al.
Published: (2024)
by: Lee, Jongseo, et al.
Published: (2024)
PCBEAR: Pose Concept Bottleneck for Explainable Action Recognition
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
Restoration-Aligned Generative Flow Models for Blind Motion Deblurring
by: Kim, Insoo, et al.
Published: (2026)
by: Kim, Insoo, et al.
Published: (2026)
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
by: Bae, Kyungho, et al.
Published: (2025)
by: Bae, Kyungho, et al.
Published: (2025)
TransFlow: Motion Knowledge Transfer from Video Diffusion Models to Video Salient Object Detection
by: Cho, Suhwan, et al.
Published: (2025)
by: Cho, Suhwan, et al.
Published: (2025)
Real-World Efficient Blind Motion Deblurring via Blur Pixel Discretization
by: Kim, Insoo, et al.
Published: (2024)
by: Kim, Insoo, et al.
Published: (2024)
Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance
by: Choi, June Suk, et al.
Published: (2025)
by: Choi, June Suk, et al.
Published: (2025)
Controllable Long-term Motion Generation with Extended Joint Targets
by: Lee, Eunjong, et al.
Published: (2025)
by: Lee, Eunjong, et al.
Published: (2025)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
Video Diffusion Models are Strong Video Inpainter
by: Lee, Minhyeok, et al.
Published: (2024)
by: Lee, Minhyeok, et al.
Published: (2024)
STATIC : Surface Temporal Affine for TIme Consistency in Video Monocular Depth Estimation
by: Yang, Sunghun, et al.
Published: (2024)
by: Yang, Sunghun, et al.
Published: (2024)
Effective SAM Combination for Open-Vocabulary Semantic Segmentation
by: Lee, Minhyeok, et al.
Published: (2024)
by: Lee, Minhyeok, et al.
Published: (2024)
DEVIAS: Learning Disentangled Video Representations of Action and Scene
by: Bae, Kyungho, et al.
Published: (2023)
by: Bae, Kyungho, et al.
Published: (2023)
PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition
by: Lee, Sanghyeon, et al.
Published: (2026)
by: Lee, Sanghyeon, et al.
Published: (2026)
How to Move Your Dragon: Text-to-Motion Synthesis for Large-Vocabulary Objects
by: Lee, Wonkwang, et al.
Published: (2025)
by: Lee, Wonkwang, et al.
Published: (2025)
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
by: Han, Jiwook, et al.
Published: (2026)
by: Han, Jiwook, et al.
Published: (2026)
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
Leveraging Image Augmentation for Object Manipulation: Towards Interpretable Controllability in Object-Centric Learning
by: Kim, Jinwoo, et al.
Published: (2023)
by: Kim, Jinwoo, et al.
Published: (2023)
Text to Blind Motion
by: Kim, Hee Jae, et al.
Published: (2024)
by: Kim, Hee Jae, et al.
Published: (2024)
MATRIX: Mask Track Alignment for Interaction-aware Video Generation
by: Jin, Siyoon, et al.
Published: (2025)
by: Jin, Siyoon, et al.
Published: (2025)
Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models
by: Kim, Sanghyun, et al.
Published: (2024)
by: Kim, Sanghyun, et al.
Published: (2024)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
by: Lee, Kyuho, et al.
Published: (2025)
by: Lee, Kyuho, et al.
Published: (2025)
A Review of Image Retrieval Techniques: Data Augmentation and Adversarial Learning Approaches
by: Jinwoo, Kim
Published: (2024)
by: Jinwoo, Kim
Published: (2024)
MonoCLUE : Object-Aware Clustering Enhances Monocular 3D Object Detection
by: Yang, Sunghun, et al.
Published: (2025)
by: Yang, Sunghun, et al.
Published: (2025)
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
by: Chung, Hyungjin, et al.
Published: (2025)
by: Chung, Hyungjin, et al.
Published: (2025)
DAFT-GAN: Dual Affine Transformation Generative Adversarial Network for Text-Guided Image Inpainting
by: Lee, Jihoon, et al.
Published: (2024)
by: Lee, Jihoon, et al.
Published: (2024)
CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Images
by: Lee, Jungho, et al.
Published: (2025)
by: Lee, Jungho, et al.
Published: (2025)
Infusing Environmental Captions for Long-Form Video Language Grounding
by: Lee, Hyogun, et al.
Published: (2024)
by: Lee, Hyogun, et al.
Published: (2024)
Flashback: Memory-Driven Zero-shot, Real-time Video Anomaly Detection
by: Lee, Hyogun, et al.
Published: (2025)
by: Lee, Hyogun, et al.
Published: (2025)
Context-Aware Input Orchestration for Video Inpainting
by: Kim, Hoyoung, et al.
Published: (2024)
by: Kim, Hoyoung, et al.
Published: (2024)
Treating Motion as Option with Output Selection for Unsupervised Video Object Segmentation
by: Cho, Suhwan, et al.
Published: (2023)
by: Cho, Suhwan, et al.
Published: (2023)
NVS-Adapter: Plug-and-Play Novel View Synthesis from a Single Image
by: Jeong, Yoonwoo, et al.
Published: (2023)
by: Jeong, Yoonwoo, et al.
Published: (2023)
HypeVPR: Exploring Hyperbolic Space for Perspective to Equirectangular Visual Place Recognition
by: Woo, Suhan, et al.
Published: (2025)
by: Woo, Suhan, et al.
Published: (2025)
Generating Human Motion Videos using a Cascaded Text-to-Video Framework
by: Nam, Hyelin, et al.
Published: (2025)
by: Nam, Hyelin, et al.
Published: (2025)
SMURF: Continuous Dynamics for Motion-Deblurring Radiance Fields
by: Lee, Jungho, et al.
Published: (2024)
by: Lee, Jungho, et al.
Published: (2024)
Similar Items
-
CAST: Cross-Attention in Space and Time for Video Action Recognition
by: Lee, Dongho, et al.
Published: (2023) -
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
by: Lee, Jongseo, et al.
Published: (2025) -
Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
by: Lee, Jongseo, et al.
Published: (2025) -
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
by: Lee, Jongseo, et al.
Published: (2024) -
PCBEAR: Pose Concept Bottleneck for Explainable Action Recognition
by: Lee, Jongseo, et al.
Published: (2025)