CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Wei, Yuan, Yuqian, Lin, Tianwei, Zhang, Wenqiao, Tang, Siliang, Xiao, Jun, Zhuang, Yueting |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CrossView-GS: Cross-view Gaussian Splatting For Large-scale Scene Reconstruction
por: Zhang, Chenhao, et al.
Publicado: (2025)
por: Zhang, Chenhao, et al.
Publicado: (2025)
InstructSAM: Segment Any Instance with Any Instructions
por: Yuan, Yuqian, et al.
Publicado: (2026)
por: Yuan, Yuqian, et al.
Publicado: (2026)
Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales
por: Gao, Minghe, et al.
Publicado: (2024)
por: Gao, Minghe, et al.
Publicado: (2024)
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
por: Yuan, Yuqian, et al.
Publicado: (2025)
por: Yuan, Yuqian, et al.
Publicado: (2025)
EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs
por: Dai, Yang, et al.
Publicado: (2026)
por: Dai, Yang, et al.
Publicado: (2026)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
por: Miao, Bingchen, et al.
Publicado: (2024)
por: Miao, Bingchen, et al.
Publicado: (2024)
EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
por: Li, Sijing, et al.
Publicado: (2025)
por: Li, Sijing, et al.
Publicado: (2025)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
por: Yuan, Yuqian, et al.
Publicado: (2024)
por: Yuan, Yuqian, et al.
Publicado: (2024)
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
por: Lin, Tianwei, et al.
Publicado: (2025)
por: Lin, Tianwei, et al.
Publicado: (2025)
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
por: Zheng, Haoyu, et al.
Publicado: (2025)
por: Zheng, Haoyu, et al.
Publicado: (2025)
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness
por: Qiu, Haiyi, et al.
Publicado: (2026)
por: Qiu, Haiyi, et al.
Publicado: (2026)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
por: Dai, Guangyu, et al.
Publicado: (2025)
por: Dai, Guangyu, et al.
Publicado: (2025)
AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization
por: Yu, Binhe, et al.
Publicado: (2025)
por: Yu, Binhe, et al.
Publicado: (2025)
CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents
por: Wang, Keyu, et al.
Publicado: (2026)
por: Wang, Keyu, et al.
Publicado: (2026)
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
por: Gao, Mingjian, et al.
Publicado: (2026)
por: Gao, Mingjian, et al.
Publicado: (2026)
Unified Personalized Understanding, Generating and Editing
por: Zhong, Yu, et al.
Publicado: (2026)
por: Zhong, Yu, et al.
Publicado: (2026)
De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
por: Gao, Minghe, et al.
Publicado: (2023)
por: Gao, Minghe, et al.
Publicado: (2023)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
por: Li, Juncheng, et al.
Publicado: (2023)
por: Li, Juncheng, et al.
Publicado: (2023)
ICG-MVSNet: Learning Intra-view and Cross-view Relationships for Guidance in Multi-View Stereo
por: Hu, Yuxi, et al.
Publicado: (2025)
por: Hu, Yuxi, et al.
Publicado: (2025)
T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from Text
por: Yin, Aoxiong, et al.
Publicado: (2024)
por: Yin, Aoxiong, et al.
Publicado: (2024)
Graft: Integrating the Domain Knowledge via Efficient Parameter Synergy for MLLMs
por: Dai, Yang, et al.
Publicado: (2025)
por: Dai, Yang, et al.
Publicado: (2025)
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
por: Ge, Zhiqi, et al.
Publicado: (2024)
por: Ge, Zhiqi, et al.
Publicado: (2024)
Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
por: Yu, Qifan, et al.
Publicado: (2024)
por: Yu, Qifan, et al.
Publicado: (2024)
LASER: Tuning-Free LLM-Driven Attention Control for Efficient Text-conditioned Image-to-Animation
por: Zheng, Haoyu, et al.
Publicado: (2024)
por: Zheng, Haoyu, et al.
Publicado: (2024)
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
por: Yuan, Yuqian, et al.
Publicado: (2026)
por: Yuan, Yuqian, et al.
Publicado: (2026)
XLD: A Cross-Lane Dataset for Benchmarking Novel Driving View Synthesis
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation
por: Zheng, Haoyu, et al.
Publicado: (2024)
por: Zheng, Haoyu, et al.
Publicado: (2024)
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
por: Zhang, Wenqiao, et al.
Publicado: (2024)
por: Zhang, Wenqiao, et al.
Publicado: (2024)
OpenView: Empowering MLLMs with Out-of-view VQA
por: Chen, Qixiang, et al.
Publicado: (2025)
por: Chen, Qixiang, et al.
Publicado: (2025)
Geometry-guided Cross-view Diffusion for One-to-many Cross-view Image Synthesis
por: Lin, Tao Jun, et al.
Publicado: (2024)
por: Lin, Tao Jun, et al.
Publicado: (2024)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
por: Fan, Zhaoyu, et al.
Publicado: (2025)
por: Fan, Zhaoyu, et al.
Publicado: (2025)
Physically Plausible Human-Object Rendering from Sparse Views via 3D Gaussian Splatting
por: Wang, Weiquan, et al.
Publicado: (2025)
por: Wang, Weiquan, et al.
Publicado: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
por: Qin, Bosheng, et al.
Publicado: (2023)
por: Qin, Bosheng, et al.
Publicado: (2023)
MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
por: Kristoffersen, Oskar, et al.
Publicado: (2025)
por: Kristoffersen, Oskar, et al.
Publicado: (2025)
CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis
por: Li, Weijia, et al.
Publicado: (2024)
por: Li, Weijia, et al.
Publicado: (2024)
OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
por: Bu, Wendong, et al.
Publicado: (2025)
por: Bu, Wendong, et al.
Publicado: (2025)
Enhancing Post-Training Quantization via Future Activation Awareness
por: Lv, Zheqi, et al.
Publicado: (2026)
por: Lv, Zheqi, et al.
Publicado: (2026)
Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
por: Gao, Minghe, et al.
Publicado: (2025)
por: Gao, Minghe, et al.
Publicado: (2025)
Cross-view Localization and Synthesis -- Datasets, Challenges and Opportunities
por: Xu, Ningli, et al.
Publicado: (2025)
por: Xu, Ningli, et al.
Publicado: (2025)
On the Generalization Capacities of MLLMs for Spatial Intelligence
por: Zhang, Gongjie, et al.
Publicado: (2026)
por: Zhang, Gongjie, et al.
Publicado: (2026)
Ejemplares similares
-
CrossView-GS: Cross-view Gaussian Splatting For Large-scale Scene Reconstruction
por: Zhang, Chenhao, et al.
Publicado: (2025) -
InstructSAM: Segment Any Instance with Any Instructions
por: Yuan, Yuqian, et al.
Publicado: (2026) -
Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales
por: Gao, Minghe, et al.
Publicado: (2024) -
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
por: Yuan, Yuqian, et al.
Publicado: (2025) -
EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs
por: Dai, Yang, et al.
Publicado: (2026)