3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xiaoye, Tang, Chen, Yue, Xiangyu, Li, Wei-Hong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis
by: Chen, Weiliang, et al.
Published: (2025)
by: Chen, Weiliang, et al.
Published: (2025)
MTPano: Multi-Task Panoramic Scene Understanding via Label-Free Integration of Dense Prediction Priors
by: Zhang, Jingdong, et al.
Published: (2026)
by: Zhang, Jingdong, et al.
Published: (2026)
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
by: Yu, Hanxun, et al.
Published: (2025)
by: Yu, Hanxun, et al.
Published: (2025)
ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency
by: Feng, Haitang, et al.
Published: (2025)
by: Feng, Haitang, et al.
Published: (2025)
MTMamba: Enhancing Multi-Task Dense Scene Understanding by Mamba-Based Decoders
by: Lin, Baijiong, et al.
Published: (2024)
by: Lin, Baijiong, et al.
Published: (2024)
Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
by: Li, Ruihuang, et al.
Published: (2024)
by: Li, Ruihuang, et al.
Published: (2024)
DenseScan: Advancing 3D Scene Understanding with 2D Dense Annotation
by: Wang, Zirui, et al.
Published: (2025)
by: Wang, Zirui, et al.
Published: (2025)
VSFormer: Mining Correlations in Flexible View Set for Multi-view 3D Shape Understanding
by: Sun, Hongyu, et al.
Published: (2024)
by: Sun, Hongyu, et al.
Published: (2024)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
by: Li, Jinlong, et al.
Published: (2025)
by: Li, Jinlong, et al.
Published: (2025)
MTMamba++: Enhancing Multi-Task Dense Scene Understanding via Mamba-Based Decoders
by: Lin, Baijiong, et al.
Published: (2024)
by: Lin, Baijiong, et al.
Published: (2024)
Cross-Task Affinity Learning for Multitask Dense Scene Predictions
by: Sinodinos, Dimitrios, et al.
Published: (2024)
by: Sinodinos, Dimitrios, et al.
Published: (2024)
From Sparse to Dense: Spatio-Temporal Fusion for Multi-View 3D Human Pose Estimation with DenseWarper
by: Li, Ling, et al.
Published: (2026)
by: Li, Ling, et al.
Published: (2026)
WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration
by: Ni, Chaojun, et al.
Published: (2025)
by: Ni, Chaojun, et al.
Published: (2025)
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
by: Chen, Jiahua, et al.
Published: (2026)
by: Chen, Jiahua, et al.
Published: (2026)
CAGS: Open-Vocabulary 3D Scene Understanding with Context-Aware Gaussian Splatting
by: Sun, Wei, et al.
Published: (2025)
by: Sun, Wei, et al.
Published: (2025)
DisCo3D: Distilling Multi-View Consistency for 3D Scene Editing
by: Chi, Yufeng, et al.
Published: (2025)
by: Chi, Yufeng, et al.
Published: (2025)
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding
by: Yin, Xingyilang, et al.
Published: (2025)
by: Yin, Xingyilang, et al.
Published: (2025)
DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding
by: Ge, Luzhou, et al.
Published: (2026)
by: Ge, Luzhou, et al.
Published: (2026)
Boosting Instance Awareness via Cross-View Correlation with 4D Radar and Camera for 3D Object Detection
by: Bai, Xiaokai, et al.
Published: (2026)
by: Bai, Xiaokai, et al.
Published: (2026)
3D Scene Change Modeling With Consistent Multi-View Aggregation
by: Zhou, Zirui, et al.
Published: (2025)
by: Zhou, Zirui, et al.
Published: (2025)
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
by: Zheng, Duo, et al.
Published: (2024)
by: Zheng, Duo, et al.
Published: (2024)
SurgCUT3R: Surgical Scene-Aware Continuous Understanding of Temporal 3D Representation
by: Xu, Kaiyuan, et al.
Published: (2026)
by: Xu, Kaiyuan, et al.
Published: (2026)
ConDense: Consistent 2D/3D Pre-training for Dense and Sparse Features from Multi-View Images
by: Zhang, Xiaoshuai, et al.
Published: (2024)
by: Zhang, Xiaoshuai, et al.
Published: (2024)
Dense Dynamic Scene Reconstruction and Camera Pose Estimation from Multi-View Videos
by: Sun, Shuo, et al.
Published: (2026)
by: Sun, Shuo, et al.
Published: (2026)
Cross-Temporal 3D Gaussian Splatting for Sparse-View Guided Scene Update
by: An, Zeyuan, et al.
Published: (2025)
by: An, Zeyuan, et al.
Published: (2025)
SceneExpander: Expanding 3D Scenes with Free-Form Inserted Views
by: He, Zijian, et al.
Published: (2026)
by: He, Zijian, et al.
Published: (2026)
FusionBERT: Multi-View Image-3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder
by: Li, Wei, et al.
Published: (2026)
by: Li, Wei, et al.
Published: (2026)
Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images
by: Sun, Xiangyu, et al.
Published: (2025)
by: Sun, Xiangyu, et al.
Published: (2025)
TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
by: Jin, Bu, et al.
Published: (2024)
by: Jin, Bu, et al.
Published: (2024)
Curvature-Aware Captioning:Leveraging Geodesic Attention for 3D Scene Understanding
by: He, Ziyao, et al.
Published: (2026)
by: He, Ziyao, et al.
Published: (2026)
FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation
by: Jiang, Chenhan, et al.
Published: (2026)
by: Jiang, Chenhan, et al.
Published: (2026)
DC-Scene: Data-Centric Learning for 3D Scene Understanding
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
Task-oriented Sequential Grounding and Navigation in 3D Scenes
by: Zhang, Zhuofan, et al.
Published: (2024)
by: Zhang, Zhuofan, et al.
Published: (2024)
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
by: Huang, Wencan, et al.
Published: (2025)
by: Huang, Wencan, et al.
Published: (2025)
HetScene: Heterogeneity-Aware Diffusion for Dense Indoor Scene Generation
by: Chen, Zini, et al.
Published: (2026)
by: Chen, Zini, et al.
Published: (2026)
GaussianGraph: 3D Gaussian-based Scene Graph Generation for Open-world Scene Understanding
by: Wang, Xihan, et al.
Published: (2025)
by: Wang, Xihan, et al.
Published: (2025)
AdaptiveFusion: Adaptive Multi-Modal Multi-View Fusion for 3D Human Body Reconstruction
by: Chen, Anjun, et al.
Published: (2024)
by: Chen, Anjun, et al.
Published: (2024)
GP-NeRF: Generalized Perception NeRF for Context-Aware 3D Scene Understanding
by: Li, Hao, et al.
Published: (2023)
by: Li, Hao, et al.
Published: (2023)
OpenGS-SLAM: Open-Set Dense Semantic SLAM with 3D Gaussian Splatting for Object-Level Scene Understanding
by: Yang, Dianyi, et al.
Published: (2025)
by: Yang, Dianyi, et al.
Published: (2025)
ToLL: Topological Layout Learning with Asymmetric Cross-View Structural Distillation for 3D Scene Graph Generation Pretraining
by: Huang, Yucheng, et al.
Published: (2026)
by: Huang, Yucheng, et al.
Published: (2026)
Similar Items
-
SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis
by: Chen, Weiliang, et al.
Published: (2025) -
MTPano: Multi-Task Panoramic Scene Understanding via Label-Free Integration of Dense Prediction Priors
by: Zhang, Jingdong, et al.
Published: (2026) -
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
by: Yu, Hanxun, et al.
Published: (2025) -
ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency
by: Feng, Haitang, et al.
Published: (2025) -
MTMamba: Enhancing Multi-Task Dense Scene Understanding by Mamba-Based Decoders
by: Lin, Baijiong, et al.
Published: (2024)