DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Jingzhou, Liu, Yang, Chen, Weixing, Li, Zhen, Wang, Yaowei, Li, Guanbin, Lin, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering
by: Jiang, Kaixuan, et al.
Published: (2025)
by: Jiang, Kaixuan, et al.
Published: (2025)
3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians
by: Wei, Zeming, et al.
Published: (2025)
by: Wei, Zeming, et al.
Published: (2025)
MEIA: Multimodal Embodied Perception and Interaction in Unknown Environments
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation
by: Chen, Weixing, et al.
Published: (2026)
by: Chen, Weixing, et al.
Published: (2026)
3D Question Answering for City Scene Understanding
by: Sun, Penglei, et al.
Published: (2024)
by: Sun, Penglei, et al.
Published: (2024)
Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
by: Song, Xinshuai, et al.
Published: (2024)
by: Song, Xinshuai, et al.
Published: (2024)
DDP-WM: Disentangled Dynamics Prediction for Efficient World Models
by: Yin, Shicheng, et al.
Published: (2026)
by: Yin, Shicheng, et al.
Published: (2026)
Cross-Modal Causal Intervention for Medical Report Generation
by: Chen, Weixing, et al.
Published: (2023)
by: Chen, Weixing, et al.
Published: (2023)
Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
by: Li, Zechuan, et al.
Published: (2025)
by: Li, Zechuan, et al.
Published: (2025)
Credible Teacher for Semi-Supervised Object Detection in Open Scene
by: Zhuang, Jingyu, et al.
Published: (2024)
by: Zhuang, Jingyu, et al.
Published: (2024)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
SceneExpander: Expanding 3D Scenes with Free-Form Inserted Views
by: He, Zijian, et al.
Published: (2026)
by: He, Zijian, et al.
Published: (2026)
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
by: Shen, Zhixuan, et al.
Published: (2024)
by: Shen, Zhixuan, et al.
Published: (2024)
Improving Data Augmentation for Robust Visual Question Answering with Effective Curriculum Learning
by: Zheng, Yuhang, et al.
Published: (2024)
by: Zheng, Yuhang, et al.
Published: (2024)
AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering
by: Zhang, Jiayu, et al.
Published: (2025)
by: Zhang, Jiayu, et al.
Published: (2025)
EgoSplat: Open-Vocabulary Egocentric Scene Understanding with Language Embedded 3D Gaussian Splatting
by: Li, Di, et al.
Published: (2025)
by: Li, Di, et al.
Published: (2025)
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
by: Zhang, Yanjie, et al.
Published: (2026)
by: Zhang, Yanjie, et al.
Published: (2026)
SeqAffordSplat: Scene-level Sequential Affordance Reasoning on 3D Gaussian Splatting
by: Li, Di, et al.
Published: (2025)
by: Li, Di, et al.
Published: (2025)
Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
by: Bai, Yongjie, et al.
Published: (2025)
by: Bai, Yongjie, et al.
Published: (2025)
A3R: Agentic Affordance Reasoning via Cross-Dimensional Evidence in 3D Gaussian Scenes
by: Li, Di, et al.
Published: (2026)
by: Li, Di, et al.
Published: (2026)
3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
by: Xu, Rongtao, et al.
Published: (2025)
by: Xu, Rongtao, et al.
Published: (2025)
SDesc3D: Towards Layout-Aware 3D Indoor Scene Generation from Short Descriptions
by: Feng, Jie, et al.
Published: (2026)
by: Feng, Jie, et al.
Published: (2026)
Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
by: Yang, Xuyi, et al.
Published: (2025)
by: Yang, Xuyi, et al.
Published: (2025)
VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
by: Li, Yuyi, et al.
Published: (2025)
by: Li, Yuyi, et al.
Published: (2025)
Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering
by: Maryam, Hiba, et al.
Published: (2024)
by: Maryam, Hiba, et al.
Published: (2024)
DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding
by: Tao, Mingzhe, et al.
Published: (2026)
by: Tao, Mingzhe, et al.
Published: (2026)
UAV-MM3D: A Large-Scale Synthetic Benchmark for 3D Perception of Unmanned Aerial Vehicles with Multi-Modal Data
by: Zou, Longkun, et al.
Published: (2025)
by: Zou, Longkun, et al.
Published: (2025)
AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning
by: Chen, Weixing, et al.
Published: (2025)
by: Chen, Weixing, et al.
Published: (2025)
Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
by: Chen, Tieyuan, et al.
Published: (2025)
by: Chen, Tieyuan, et al.
Published: (2025)
IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
by: Zhou, Sheng, et al.
Published: (2025)
by: Zhou, Sheng, et al.
Published: (2025)
NeRF-HuGS: Improved Neural Radiance Fields in Non-static Scenes Using Heuristics-Guided Segmentation
by: Chen, Jiahao, et al.
Published: (2024)
by: Chen, Jiahao, et al.
Published: (2024)
Resonance4D: Frequency-Domain Motion Supervision for Preset-Free Physical Parameter Learning in 4D Dynamic Physical Scene Simulation
by: Zhang, Changshe, et al.
Published: (2026)
by: Zhang, Changshe, et al.
Published: (2026)
Robust and Efficient 3D Gaussian Splatting for Urban Scene Reconstruction
by: Yuan, Zhensheng, et al.
Published: (2025)
by: Yuan, Zhensheng, et al.
Published: (2025)
Scene-Text Grounding for Text-Based Video Question Answering
by: Zhou, Sheng, et al.
Published: (2024)
by: Zhou, Sheng, et al.
Published: (2024)
SeGPruner: Semantic-Geometric Visual Token Pruner for 3D Question Answering
by: Li, Wenli, et al.
Published: (2026)
by: Li, Wenli, et al.
Published: (2026)
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving
by: Lin, Yuxin, et al.
Published: (2025)
by: Lin, Yuxin, et al.
Published: (2025)
DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
by: Al-Mohannadi, Aisha, et al.
Published: (2026)
by: Al-Mohannadi, Aisha, et al.
Published: (2026)
Cross-modal Causal Relation Alignment for Video Question Grounding
by: Chen, Weixing, et al.
Published: (2025)
by: Chen, Weixing, et al.
Published: (2025)
Similar Items
-
Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering
by: Jiang, Kaixuan, et al.
Published: (2025) -
3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians
by: Wei, Zeming, et al.
Published: (2025) -
MEIA: Multimodal Embodied Perception and Interaction in Unknown Environments
by: Liu, Yang, et al.
Published: (2024) -
PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation
by: Chen, Weixing, et al.
Published: (2026) -
3D Question Answering for City Scene Understanding
by: Sun, Penglei, et al.
Published: (2024)