Efficient RGB-D Scene Understanding via Multi-task Adaptive Learning and Cross-dimensional Feature Guidance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Guodong, Liu, Junjie, Zhang, Gaoyang, Wu, Bo, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation
von: Li, Bohan, et al.
Veröffentlicht: (2024)
von: Li, Bohan, et al.
Veröffentlicht: (2024)
Dynamic Graph Neural Network with Adaptive Features Selection for RGB-D Based Indoor Scene Recognition
von: Liu, Qiong, et al.
Veröffentlicht: (2026)
von: Liu, Qiong, et al.
Veröffentlicht: (2026)
PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding
von: Wen, Junjie, et al.
Veröffentlicht: (2026)
von: Wen, Junjie, et al.
Veröffentlicht: (2026)
Efficient Multi-Task Scene Analysis with RGB-D Transformers
von: Fischedick, Söhnke Benedikt, et al.
Veröffentlicht: (2023)
von: Fischedick, Söhnke Benedikt, et al.
Veröffentlicht: (2023)
Real-Time Glottis Detection Framework via Spatial-decoupled Feature Learning for Nasal Transnasal Intubation
von: Liu, Jinyu, et al.
Veröffentlicht: (2026)
von: Liu, Jinyu, et al.
Veröffentlicht: (2026)
DC-Scene: Data-Centric Learning for 3D Scene Understanding
von: Huang, Ting, et al.
Veröffentlicht: (2025)
von: Huang, Ting, et al.
Veröffentlicht: (2025)
HAPNet: Toward Superior RGB-Thermal Scene Parsing via Hybrid, Asymmetric, and Progressive Heterogeneous Feature Fusion
von: Li, Jiahang, et al.
Veröffentlicht: (2024)
von: Li, Jiahang, et al.
Veröffentlicht: (2024)
MLVU: Benchmarking Multi-task Long Video Understanding
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
Delving into Multi-modal Multi-task Foundation Models for Road Scene Understanding: From Learning Paradigm Perspectives
von: Luo, Sheng, et al.
Veröffentlicht: (2024)
von: Luo, Sheng, et al.
Veröffentlicht: (2024)
BridgeNet: Comprehensive and Effective Feature Interactions via Bridge Feature for Multi-task Dense Predictions
von: Zhang, Jingdong, et al.
Veröffentlicht: (2023)
von: Zhang, Jingdong, et al.
Veröffentlicht: (2023)
NIS-SLAM: Neural Implicit Semantic RGB-D SLAM for 3D Consistent Scene Understanding
von: Zhai, Hongjia, et al.
Veröffentlicht: (2024)
von: Zhai, Hongjia, et al.
Veröffentlicht: (2024)
RGB-D Video Object Segmentation via Enhanced Multi-store Feature Memory
von: Xu, Boyue, et al.
Veröffentlicht: (2025)
von: Xu, Boyue, et al.
Veröffentlicht: (2025)
Elite360M: Efficient 360 Multi-task Learning via Bi-projection Fusion and Cross-task Collaboration
von: Ai, Hao, et al.
Veröffentlicht: (2024)
von: Ai, Hao, et al.
Veröffentlicht: (2024)
Learning Adaptive Lighting via Channel-Aware Guidance
von: Yang, Qirui, et al.
Veröffentlicht: (2024)
von: Yang, Qirui, et al.
Veröffentlicht: (2024)
Learning Temporal 3D Semantic Scene Completion via Optical Flow Guidance
von: Wang, Meng, et al.
Veröffentlicht: (2025)
von: Wang, Meng, et al.
Veröffentlicht: (2025)
FAMNet: Integrating 2D and 3D Features for Micro-expression Recognition via Multi-task Learning and Hierarchical Attention
von: Fu, Liangyu, et al.
Veröffentlicht: (2025)
von: Fu, Liangyu, et al.
Veröffentlicht: (2025)
Adaptive Visual Scene Understanding: Incremental Scene Graph Generation
von: Khandelwal, Naitik, et al.
Veröffentlicht: (2023)
von: Khandelwal, Naitik, et al.
Veröffentlicht: (2023)
MAESTRO: Task-Relevant Optimization via Adaptive Feature Enhancement and Suppression for Multi-task 3D Perception
von: Kang, Changwon, et al.
Veröffentlicht: (2025)
von: Kang, Changwon, et al.
Veröffentlicht: (2025)
3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
von: Wang, Xiaoye, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoye, et al.
Veröffentlicht: (2025)
CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image
von: Yao, Kaixin, et al.
Veröffentlicht: (2025)
von: Yao, Kaixin, et al.
Veröffentlicht: (2025)
RGB-D Indiscernible Object Counting in Underwater Scenes
von: Sun, Guolei, et al.
Veröffentlicht: (2023)
von: Sun, Guolei, et al.
Veröffentlicht: (2023)
Expressive yet Efficient Feature Expansion with Adaptive Cross-Hadamard Products
von: Zhang, Xuyang, et al.
Veröffentlicht: (2025)
von: Zhang, Xuyang, et al.
Veröffentlicht: (2025)
Language-Assisted 3D Scene Understanding
von: Wu, Yanmin, et al.
Veröffentlicht: (2023)
von: Wu, Yanmin, et al.
Veröffentlicht: (2023)
RGB-Phase Speckle: Cross-Scene Stereo 3D Reconstruction via Wrapped Pre-Normalization
von: Yang, Kai, et al.
Veröffentlicht: (2025)
von: Yang, Kai, et al.
Veröffentlicht: (2025)
RGB-Sonar Tracking Benchmark and Spatial Cross-Attention Transformer Tracker
von: Li, Yunfeng, et al.
Veröffentlicht: (2024)
von: Li, Yunfeng, et al.
Veröffentlicht: (2024)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
AVS-Net: Point Sampling with Adaptive Voxel Size for 3D Scene Understanding
von: Yang, Hongcheng, et al.
Veröffentlicht: (2024)
von: Yang, Hongcheng, et al.
Veröffentlicht: (2024)
SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB Sequences
von: Kim, Seok-Young, et al.
Veröffentlicht: (2026)
von: Kim, Seok-Young, et al.
Veröffentlicht: (2026)
DoGFlow: Self-Supervised LiDAR Scene Flow via Cross-Modal Doppler Guidance
von: Khoche, Ajinkya, et al.
Veröffentlicht: (2025)
von: Khoche, Ajinkya, et al.
Veröffentlicht: (2025)
Multiple Prior Representation Learning for Self-Supervised Monocular Depth Estimation via Hybrid Transformer
von: Sun, Guodong, et al.
Veröffentlicht: (2024)
von: Sun, Guodong, et al.
Veröffentlicht: (2024)
Dissecting RGB-D Learning for Improved Multi-modal Fusion
von: Chen, Hao, et al.
Veröffentlicht: (2023)
von: Chen, Hao, et al.
Veröffentlicht: (2023)
Improving Viewpoint Consistency in 3D Generation via Structure Feature and CLIP Guidance
von: Zhang, Qing, et al.
Veröffentlicht: (2024)
von: Zhang, Qing, et al.
Veröffentlicht: (2024)
Universal 3D Shape Matching via Coarse-to-Fine Language Guidance
von: Xiao, Qinfeng, et al.
Veröffentlicht: (2026)
von: Xiao, Qinfeng, et al.
Veröffentlicht: (2026)
GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
SceneReVis: A Self-Reflective Vision-Grounded Framework for 3D Indoor Scene Synthesis via Multi-turn RL
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training?
von: Sun, Lingchen, et al.
Veröffentlicht: (2026)
von: Sun, Lingchen, et al.
Veröffentlicht: (2026)
LocalDyGS: Multi-view Global Dynamic Scene Modeling via Adaptive Local Implicit Feature Decoupling
von: Wu, Jiahao, et al.
Veröffentlicht: (2025)
von: Wu, Jiahao, et al.
Veröffentlicht: (2025)
EAGLE: Efficient Adaptive Geometry-based Learning in Cross-view Understanding
von: Truong, Thanh-Dat, et al.
Veröffentlicht: (2024)
von: Truong, Thanh-Dat, et al.
Veröffentlicht: (2024)
3D Question Answering for City Scene Understanding
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
Towards Balanced RGB-TSDF Fusion for Consistent Semantic Scene Completion by 3D RGB Feature Completion and a Classwise Entropy Loss Function
von: Ding, Laiyan, et al.
Veröffentlicht: (2024)
von: Ding, Laiyan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation
von: Li, Bohan, et al.
Veröffentlicht: (2024) -
Dynamic Graph Neural Network with Adaptive Features Selection for RGB-D Based Indoor Scene Recognition
von: Liu, Qiong, et al.
Veröffentlicht: (2026) -
PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding
von: Wen, Junjie, et al.
Veröffentlicht: (2026) -
Efficient Multi-Task Scene Analysis with RGB-D Transformers
von: Fischedick, Söhnke Benedikt, et al.
Veröffentlicht: (2023) -
Real-Time Glottis Detection Framework via Spatial-decoupled Feature Learning for Nasal Transnasal Intubation
von: Liu, Jinyu, et al.
Veröffentlicht: (2026)