Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Yikang, Zhang, Tao, Xu, Shilin, Chen, Shihao, Zhou, Qianyu, Tong, Yunhai, Ji, Shunping, Zhang, Jiangning, Qi, Lu, Li, Xiangtai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dense360: Dense Understanding from Omnidirectional Panoramas
by: Zhou, Yikang, et al.
Published: (2025)
by: Zhou, Yikang, et al.
Published: (2025)
Point Cloud Mamba: Point Cloud Learning via State Space Model
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries
by: Zhou, Yikang, et al.
Published: (2024)
by: Zhou, Yikang, et al.
Published: (2024)
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation
by: Niu, Quanzhu, et al.
Published: (2025)
by: Niu, Quanzhu, et al.
Published: (2025)
The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
by: Niu, Quanzhu, et al.
Published: (2025)
by: Niu, Quanzhu, et al.
Published: (2025)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
LLAVADI: What Matters For Multimodal Large Language Models Distillation
by: Xu, Shilin, et al.
Published: (2024)
by: Xu, Shilin, et al.
Published: (2024)
MotionBooth: Motion-Aware Customized Text-to-Video Generation
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
Explore In-Context Segmentation via Latent Diffusion Models
by: Wang, Chaoyang, et al.
Published: (2024)
by: Wang, Chaoyang, et al.
Published: (2024)
Conditional Panoramic Image Generation via Masked Autoregressive Modeling
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer
by: Shi, Qingyu, et al.
Published: (2025)
by: Shi, Qingyu, et al.
Published: (2025)
VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track
by: Gong, Dengxian, et al.
Published: (2026)
by: Gong, Dengxian, et al.
Published: (2026)
P2PFormer: A Primitive-to-polygon Method for Regular Building Contour Extraction from Remote Sensing Images
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
by: Xu, Shilin, et al.
Published: (2023)
by: Xu, Shilin, et al.
Published: (2023)
Generative Classifier for Domain Generalization
by: Long, Shaocong, et al.
Published: (2025)
by: Long, Shaocong, et al.
Published: (2025)
Towards Language-Driven Video Inpainting via Multimodal Large Language Models
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
by: Xu, Shilin, et al.
Published: (2025)
by: Xu, Shilin, et al.
Published: (2025)
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
Towards Open Vocabulary Learning: A Survey
by: Wu, Jianzong, et al.
Published: (2023)
by: Wu, Jianzong, et al.
Published: (2023)
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
PanopticPartFormer++: A Unified and Decoupled View for Panoptic Part Segmentation
by: Li, Xiangtai, et al.
Published: (2023)
by: Li, Xiangtai, et al.
Published: (2023)
SemFlow: Binding Semantic Segmentation and Image Synthesis via Rectified Flow
by: Wang, Chaoyang, et al.
Published: (2024)
by: Wang, Chaoyang, et al.
Published: (2024)
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
by: Li, Xiangtai, et al.
Published: (2025)
by: Li, Xiangtai, et al.
Published: (2025)
CyberV: Cybernetics for Test-time Scaling in Video Understanding
by: Meng, Jiahao, et al.
Published: (2025)
by: Meng, Jiahao, et al.
Published: (2025)
DreamRelation: Bridging Customization and Relation Generation
by: Shi, Qingyu, et al.
Published: (2024)
by: Shi, Qingyu, et al.
Published: (2024)
SAMTok: Representing Any Mask with Two Words
by: Zhou, Yikang, et al.
Published: (2026)
by: Zhou, Yikang, et al.
Published: (2026)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
Exploring Plain ViT Reconstruction for Multi-class Unsupervised Anomaly Detection
by: Zhang, Jiangning, et al.
Published: (2023)
by: Zhang, Jiangning, et al.
Published: (2023)
4th PVUW MeViS 3rd Place Report: Sa2VA
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything
by: Xu, Shilin, et al.
Published: (2024)
by: Xu, Shilin, et al.
Published: (2024)
DeH4R: A Decoupled and Hybrid Method for Road Network Graph Extraction
by: Gong, Dengxian, et al.
Published: (2025)
by: Gong, Dengxian, et al.
Published: (2025)
BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything Model
by: Song, Yiran, et al.
Published: (2024)
by: Song, Yiran, et al.
Published: (2024)
DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers
by: Wang, Zitong, et al.
Published: (2025)
by: Wang, Zitong, et al.
Published: (2025)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
by: Xue, Zhucun, et al.
Published: (2025)
by: Xue, Zhucun, et al.
Published: (2025)
PointDGRWKV: Generalizing RWKV-like Architecture to Unseen Domains for Point Cloud Classification
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time Scaling
by: Yu, Xinlei, et al.
Published: (2025)
by: Yu, Xinlei, et al.
Published: (2025)
Bridge Feature Matching and Cross-Modal Alignment with Mutual-filtering for Zero-shot Anomaly Detection
by: Bai, Yuhu, et al.
Published: (2025)
by: Bai, Yuhu, et al.
Published: (2025)
Similar Items
-
Dense360: Dense Understanding from Omnidirectional Panoramas
by: Zhou, Yikang, et al.
Published: (2025) -
Point Cloud Mamba: Point Cloud Learning via State Space Model
by: Zhang, Tao, et al.
Published: (2024) -
DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries
by: Zhou, Yikang, et al.
Published: (2024) -
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation
by: Niu, Quanzhu, et al.
Published: (2025) -
The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
by: Niu, Quanzhu, et al.
Published: (2025)