PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
Fuente:
arXiv
Saved in:
| Main Author: | Siam, Mennatullah |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?
by: Siam, Mennatullah
Published: (2025)
by: Siam, Mennatullah
Published: (2025)
Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving
by: Cheshmi, Leila, et al.
Published: (2025)
by: Cheshmi, Leila, et al.
Published: (2025)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
by: Goyal, Raghav, et al.
Published: (2023)
by: Goyal, Raghav, et al.
Published: (2023)
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024)
MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer
by: Karim, Rezaul, et al.
Published: (2023)
by: Karim, Rezaul, et al.
Published: (2023)
Dynamics Based Neural Encoding with Inter-Intra Region Connectivity
by: Gamal, Mai, et al.
Published: (2024)
by: Gamal, Mai, et al.
Published: (2024)
The Power of One: A Single Example is All it Takes for Segmentation in VLMs
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)
Generalized Few-Shot Semantic Segmentation in Remote Sensing: Challenge and Benchmark
by: Broni-Bediako, Clifford, et al.
Published: (2024)
by: Broni-Bediako, Clifford, et al.
Published: (2024)
Enhanced Pix2Pix GAN for Visual Defect Removal in UAV-Captured Images
by: Rizun, Volodymyr
Published: (2024)
by: Rizun, Volodymyr
Published: (2024)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
by: Ou, Ruizhe, et al.
Published: (2025)
by: Ou, Ruizhe, et al.
Published: (2025)
Pix2Gif: Motion-Guided Diffusion for GIF Generation
by: Kandala, Hitesh, et al.
Published: (2024)
by: Kandala, Hitesh, et al.
Published: (2024)
ReasonPix2Pix: Instruction Reasoning Dataset for Advanced Image Editing
by: Jin, Ying, et al.
Published: (2024)
by: Jin, Ying, et al.
Published: (2024)
Mapping New Realities: Ground Truth Image Creation with Pix2Pix Image-to-Image Translation
by: Li, Zhenglin, et al.
Published: (2024)
by: Li, Zhenglin, et al.
Published: (2024)
Quantifying and Learning Static vs. Dynamic Information in Deep Spatiotemporal Networks
by: Kowal, Matthew, et al.
Published: (2022)
by: Kowal, Matthew, et al.
Published: (2022)
Multimodal Crowd Counting with Pix2Pix GANs
by: Khan, Muhammad Asif, et al.
Published: (2024)
by: Khan, Muhammad Asif, et al.
Published: (2024)
InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following
by: Li, Shufan, et al.
Published: (2023)
by: Li, Shufan, et al.
Published: (2023)
MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
by: Ding, Henghui, et al.
Published: (2025)
by: Ding, Henghui, et al.
Published: (2025)
Multi-Modal Hallucination Control by Visual Information Grounding
by: Favero, Alessandro, et al.
Published: (2024)
by: Favero, Alessandro, et al.
Published: (2024)
BEDLAM2.0: Synthetic Humans and Cameras in Motion
by: Tesch, Joachim, et al.
Published: (2025)
by: Tesch, Joachim, et al.
Published: (2025)
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
Large Motion Model for Unified Multi-Modal Motion Generation
by: Zhang, Mingyuan, et al.
Published: (2024)
by: Zhang, Mingyuan, et al.
Published: (2024)
OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog
by: Abdessaied, Adnen, et al.
Published: (2024)
by: Abdessaied, Adnen, et al.
Published: (2024)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
by: Padhi, Trilok, et al.
Published: (2025)
by: Padhi, Trilok, et al.
Published: (2025)
Pix2Next: Leveraging Vision Foundation Models for RGB to NIR Image Translation
by: Jin, Youngwan, et al.
Published: (2024)
by: Jin, Youngwan, et al.
Published: (2024)
MotiMotion: Motion-Controlled Video Generation with Visual Reasoning
by: Hsin-Ying, Lee, et al.
Published: (2026)
by: Hsin-Ying, Lee, et al.
Published: (2026)
LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding
by: Xiao, Feng, et al.
Published: (2025)
by: Xiao, Feng, et al.
Published: (2025)
MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
by: Feng, Yue, et al.
Published: (2025)
by: Feng, Yue, et al.
Published: (2025)
Exploiting Modality-Specific Features For Multi-Modal Manipulation Detection And Grounding
by: Wang, Jiazhen, et al.
Published: (2023)
by: Wang, Jiazhen, et al.
Published: (2023)
MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer
by: Liu, Penghui, et al.
Published: (2025)
by: Liu, Penghui, et al.
Published: (2025)
Pix2Key: Controllable Open-Vocabulary Retrieval with Semantic Decomposition and Self-Supervised Visual Dictionary Learning
by: Wei, Guoyizhe, et al.
Published: (2026)
by: Wei, Guoyizhe, et al.
Published: (2026)
TempCompass: Do Video LLMs Really Understand Videos?
by: Liu, Yuanxin, et al.
Published: (2024)
by: Liu, Yuanxin, et al.
Published: (2024)
KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
by: Li, Zhiyang, et al.
Published: (2026)
by: Li, Zhiyang, et al.
Published: (2026)
Prithvi-EO-2.0: A Versatile Multi-Temporal Foundation Model for Earth Observation Applications
by: Szwarcman, Daniela, et al.
Published: (2024)
by: Szwarcman, Daniela, et al.
Published: (2024)
SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
by: Tan, Shuai, et al.
Published: (2025)
by: Tan, Shuai, et al.
Published: (2025)
EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
by: Yang, Yuxiao, et al.
Published: (2025)
by: Yang, Yuxiao, et al.
Published: (2025)
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
by: Ghazanfari, Sara, et al.
Published: (2024)
by: Ghazanfari, Sara, et al.
Published: (2024)
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
by: Plizzari, Chiara, et al.
Published: (2025)
by: Plizzari, Chiara, et al.
Published: (2025)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
by: Peng, Haosong, et al.
Published: (2025)
by: Peng, Haosong, et al.
Published: (2025)
Similar Items
-
PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?
by: Siam, Mennatullah
Published: (2025) -
Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving
by: Cheshmi, Leila, et al.
Published: (2025) -
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
by: Goyal, Raghav, et al.
Published: (2023) -
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024) -
MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer
by: Karim, Rezaul, et al.
Published: (2023)