OpenView: Empowering MLLMs with Out-of-view VQA
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Qixiang, Zhang, Cheng, Fu, Chi-Wing, Ye, Jingwen, Cai, Jianfei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenViewer: Openness-Aware Multi-View Learning
by: Du, Shide, et al.
Published: (2024)
by: Du, Shide, et al.
Published: (2024)
Delta-SVD: Efficient Compression for Personalized Text-to-Image Models
by: Zhang, Tangyuan, et al.
Published: (2025)
by: Zhang, Tangyuan, et al.
Published: (2025)
FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs
by: Wang, Xiaoqin, et al.
Published: (2025)
by: Wang, Xiaoqin, et al.
Published: (2025)
Ray Denoising: Depth-aware Hard Negative Sampling for Multi-view 3D Object Detection
by: Liu, Feng, et al.
Published: (2024)
by: Liu, Feng, et al.
Published: (2024)
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
by: Li, Shuai, et al.
Published: (2025)
by: Li, Shuai, et al.
Published: (2025)
ChatCam: Empowering Camera Control through Conversational AI
by: Liu, Xinhang, et al.
Published: (2024)
by: Liu, Xinhang, et al.
Published: (2024)
SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View Adaptation
by: Wang, Yinqiao, et al.
Published: (2024)
by: Wang, Yinqiao, et al.
Published: (2024)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Heavy Labels Out! Dataset Distillation with Label Space Lightening
by: Yu, Ruonan, et al.
Published: (2024)
by: Yu, Ruonan, et al.
Published: (2024)
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
by: Ou, Linyu, et al.
Published: (2025)
by: Ou, Linyu, et al.
Published: (2025)
iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
by: Zhao, Zhaoran, et al.
Published: (2025)
by: Zhao, Zhaoran, et al.
Published: (2025)
SpatialTree: How Spatial Abilities Branch Out in MLLMs
by: Xiao, Yuxi, et al.
Published: (2025)
by: Xiao, Yuxi, et al.
Published: (2025)
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
ICG-MVSNet: Learning Intra-view and Cross-view Relationships for Guidance in Multi-View Stereo
by: Hu, Yuxi, et al.
Published: (2025)
by: Hu, Yuxi, et al.
Published: (2025)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
by: Zhang, Shan, et al.
Published: (2025)
by: Zhang, Shan, et al.
Published: (2025)
Differentiable Convex Polyhedra Optimization from Multi-view Images
by: Ren, Daxuan, et al.
Published: (2024)
by: Ren, Daxuan, et al.
Published: (2024)
CloseUpShot: Close-up Novel View Synthesis from Sparse-views via Point-conditioned Diffusion Model
by: Zhang, Yuqi, et al.
Published: (2025)
by: Zhang, Yuqi, et al.
Published: (2025)
Generative Region-Language Pretraining for Open-Ended Object Detection
by: Lin, Chuang, et al.
Published: (2024)
by: Lin, Chuang, et al.
Published: (2024)
VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving
by: Wang, Jie, et al.
Published: (2026)
by: Wang, Jie, et al.
Published: (2026)
Delving Deep into Semantic Relation Distillation
by: Yan, Zhaoyi, et al.
Published: (2025)
by: Yan, Zhaoyi, et al.
Published: (2025)
Expandable Residual Approximation for Knowledge Distillation
by: Yan, Zhaoyi, et al.
Published: (2025)
by: Yan, Zhaoyi, et al.
Published: (2025)
Uncertainty-guided Optimal Transport in Depth Supervised Sparse-View 3D Gaussian
by: Sun, Wei, et al.
Published: (2024)
by: Sun, Wei, et al.
Published: (2024)
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations
by: Hsu, Peng-Hao, et al.
Published: (2025)
by: Hsu, Peng-Hao, et al.
Published: (2025)
Enhancing Multi-view Open-set Learning via Ambiguity Uncertainty Calibration and View-wise Debiasing
by: Fang, Zihan, et al.
Published: (2025)
by: Fang, Zihan, et al.
Published: (2025)
MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse Views
by: Chen, Yuedong, et al.
Published: (2024)
by: Chen, Yuedong, et al.
Published: (2024)
Mixture of Physical Priors Adapter for Parameter-Efficient Fine-Tuning
by: Wang, Zhaozhi, et al.
Published: (2024)
by: Wang, Zhaozhi, et al.
Published: (2024)
AceTone: Bridging Words and Colors for Conditional Image Grading
by: Ma, Tianren, et al.
Published: (2026)
by: Ma, Tianren, et al.
Published: (2026)
Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-Observations for High-Quality Sparse-View Reconstruction
by: Liu, Xinhang, et al.
Published: (2023)
by: Liu, Xinhang, et al.
Published: (2023)
Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles
by: Lin, Chuang, et al.
Published: (2024)
by: Lin, Chuang, et al.
Published: (2024)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
by: He, Haibin, et al.
Published: (2026)
by: He, Haibin, et al.
Published: (2026)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
by: Zhang, Chengyi, et al.
Published: (2026)
by: Zhang, Chengyi, et al.
Published: (2026)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
ProFuse: Efficient Cross-View Context Fusion for Open-Vocabulary 3D Gaussian Splatting
by: Chiou, Yen-Jen, et al.
Published: (2026)
by: Chiou, Yen-Jen, et al.
Published: (2026)
DropoutGS: Dropping Out Gaussians for Better Sparse-view Rendering
by: Xu, Yexing, et al.
Published: (2025)
by: Xu, Yexing, et al.
Published: (2025)
Knowledge Condensation and Reasoning for Knowledge-based VQA
by: Hao, Dongze, et al.
Published: (2024)
by: Hao, Dongze, et al.
Published: (2024)
Empowering Vector Graphics with Consistently Arbitrary Viewing and View-dependent Visibility
by: Li, Yidi, et al.
Published: (2025)
by: Li, Yidi, et al.
Published: (2025)
Similar Items
-
OpenViewer: Openness-Aware Multi-View Learning
by: Du, Shide, et al.
Published: (2024) -
Delta-SVD: Efficient Compression for Personalized Text-to-Image Models
by: Zhang, Tangyuan, et al.
Published: (2025) -
FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs
by: Wang, Xiaoqin, et al.
Published: (2025) -
Ray Denoising: Depth-aware Hard Negative Sampling for Multi-view 3D Object Detection
by: Liu, Feng, et al.
Published: (2024) -
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
by: Wang, Yi, et al.
Published: (2025)