Two-Stream Interactive Joint Learning of Scene Parsing and Geometric Vision Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Guanfeng, Zhao, Hongbo, Long, Ziwei, Li, Jiayao, Xiao, Bohong, Ye, Wei, Wang, Hanli, Fan, Rui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing
by: Fan, Jiahe, et al.
Published: (2026)
by: Fan, Jiahe, et al.
Published: (2026)
A Birotation Solution for Relative Pose Problems
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
Discriminately Treating Motion Components Evolves Joint Depth and Ego-Motion Learning
by: Zhang, Mengtan, et al.
Published: (2025)
by: Zhang, Mengtan, et al.
Published: (2025)
TiCoSS: Tightening the Coupling between Semantic Segmentation and Stereo Matching within A Joint Learning Framework
by: Tang, Guanfeng, et al.
Published: (2024)
by: Tang, Guanfeng, et al.
Published: (2024)
Fully Exploiting Vision Foundation Model's Profound Prior Knowledge for Generalizable RGB-Depth Driving Scene Parsing
by: Guo, Sicen, et al.
Published: (2025)
by: Guo, Sicen, et al.
Published: (2025)
DyStream: Streaming Dyadic Talking Heads Generation via Flow Matching-based Autoregressive Model
by: Chen, Bohong, et al.
Published: (2025)
by: Chen, Bohong, et al.
Published: (2025)
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
by: Li, Haoyuan, et al.
Published: (2025)
by: Li, Haoyuan, et al.
Published: (2025)
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding
by: Xu, Pengxin, et al.
Published: (2026)
by: Xu, Pengxin, et al.
Published: (2026)
RoadFormer: Duplex Transformer for RGB-Normal Semantic Road Scene Parsing
by: Li, Jiahang, et al.
Published: (2023)
by: Li, Jiahang, et al.
Published: (2023)
HAPNet: Toward Superior RGB-Thermal Scene Parsing via Hybrid, Asymmetric, and Progressive Heterogeneous Feature Fusion
by: Li, Jiahang, et al.
Published: (2024)
by: Li, Jiahang, et al.
Published: (2024)
DepthMatch: Semi-Supervised RGB-D Scene Parsing through Depth-Guided Regularization
by: Huang, Jianxin, et al.
Published: (2025)
by: Huang, Jianxin, et al.
Published: (2025)
The Midas Touch for Metric Depth
by: Ma, Yu, et al.
Published: (2026)
by: Ma, Yu, et al.
Published: (2026)
PIG: Prompt Images Guidance for Night-Time Scene Parsing
by: Xie, Zhifeng, et al.
Published: (2024)
by: Xie, Zhifeng, et al.
Published: (2024)
Leveraging Geometric Priors for Unaligned Scene Change Detection
by: Liu, Ziling, et al.
Published: (2025)
by: Liu, Ziling, et al.
Published: (2025)
Integrating Disparity Confidence Estimation into Relative Depth Prior-Guided Unsupervised Stereo Matching
by: Liu, Chuang-Wei, et al.
Published: (2025)
by: Liu, Chuang-Wei, et al.
Published: (2025)
Robotic Surgery Remote Mentoring via AR with 3D Scene Streaming and Hand Interaction
by: Long, Yonghao, et al.
Published: (2022)
by: Long, Yonghao, et al.
Published: (2022)
Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning
by: Liang, Dayong, et al.
Published: (2025)
by: Liang, Dayong, et al.
Published: (2025)
Instant Gaussian Stream: Fast and Generalizable Streaming of Dynamic Scene Reconstruction via Gaussian Splatting
by: Yan, Jinbo, et al.
Published: (2025)
by: Yan, Jinbo, et al.
Published: (2025)
StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
LIX: Implicitly Infusing Spatial Geometric Prior Knowledge into Visual Semantic Segmentation for Autonomous Driving
by: Guo, Sicen, et al.
Published: (2024)
by: Guo, Sicen, et al.
Published: (2024)
RoadFormer+: Delivering RGB-X Scene Parsing through Scale-Aware Information Decoupling and Advanced Heterogeneous Feature Fusion
by: Huang, Jianxin, et al.
Published: (2024)
by: Huang, Jianxin, et al.
Published: (2024)
Towards Two-Stream Foveation-based Active Vision Learning
by: Ibrayev, Timur, et al.
Published: (2024)
by: Ibrayev, Timur, et al.
Published: (2024)
Generative Face Parsing Map Guided 3D Face Reconstruction Under Occluded Scenes
by: Zhao, Dapeng, et al.
Published: (2024)
by: Zhao, Dapeng, et al.
Published: (2024)
Learning AND-OR Templates for Professional Photograph Parsing and Guidance
by: Jin, Xin, et al.
Published: (2024)
by: Jin, Xin, et al.
Published: (2024)
HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions
by: Cao, Yukang, et al.
Published: (2026)
by: Cao, Yukang, et al.
Published: (2026)
Controlling Vision-Language Models for Multi-Task Image Restoration
by: Luo, Ziwei, et al.
Published: (2023)
by: Luo, Ziwei, et al.
Published: (2023)
Radar Spectra-Language Model for Automotive Scene Parsing
by: Pushkareva, Mariia, et al.
Published: (2024)
by: Pushkareva, Mariia, et al.
Published: (2024)
3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
by: Wang, Xiaoye, et al.
Published: (2025)
by: Wang, Xiaoye, et al.
Published: (2025)
Traffic Scene Parsing through the TSP6K Dataset
by: Jiang, Peng-Tao, et al.
Published: (2023)
by: Jiang, Peng-Tao, et al.
Published: (2023)
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2024)
by: Zhan, Yufei, et al.
Published: (2024)
Sparse3DPR: Training-Free 3D Hierarchical Scene Parsing and Task-Adaptive Subgraph Reasoning from Sparse RGB Views
by: Feng, Haida, et al.
Published: (2025)
by: Feng, Haida, et al.
Published: (2025)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
by: Qi, Zhangyang, et al.
Published: (2025)
by: Qi, Zhangyang, et al.
Published: (2025)
HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene Perception
by: Yao, Wei, et al.
Published: (2025)
by: Yao, Wei, et al.
Published: (2025)
Environment-Driven Online LiDAR-Camera Extrinsic Calibration
by: Huang, Zhiwei, et al.
Published: (2025)
by: Huang, Zhiwei, et al.
Published: (2025)
Aligning Neuronal Coding of Dynamic Visual Scenes with Foundation Vision Models
by: Wu, Rining, et al.
Published: (2024)
by: Wu, Rining, et al.
Published: (2024)
An Instance-Centric Panoptic Occupancy Prediction Benchmark for Autonomous Driving
by: Feng, Yi, et al.
Published: (2026)
by: Feng, Yi, et al.
Published: (2026)
Dynamic Scene Reconstruction: Recent Advance in Real-time Rendering and Streaming
by: Zhu, Jiaxuan, et al.
Published: (2025)
by: Zhu, Jiaxuan, et al.
Published: (2025)
A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects
by: Xie, Guohuan, et al.
Published: (2025)
by: Xie, Guohuan, et al.
Published: (2025)
Towards Geometric and Textural Consistency 3D Scene Generation via Single Image-guided Model Generation and Layout Optimization
by: Tang, Xiang, et al.
Published: (2025)
by: Tang, Xiang, et al.
Published: (2025)
Image-Plane Geometric Decoding for View-Invariant Indoor Scene Reconstruction
by: Li, Mingyang, et al.
Published: (2025)
by: Li, Mingyang, et al.
Published: (2025)
Similar Items
-
Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing
by: Fan, Jiahe, et al.
Published: (2026) -
A Birotation Solution for Relative Pose Problems
by: Zhao, Hongbo, et al.
Published: (2025) -
Discriminately Treating Motion Components Evolves Joint Depth and Ego-Motion Learning
by: Zhang, Mengtan, et al.
Published: (2025) -
TiCoSS: Tightening the Coupling between Semantic Segmentation and Stereo Matching within A Joint Learning Framework
by: Tang, Guanfeng, et al.
Published: (2024) -
Fully Exploiting Vision Foundation Model's Profound Prior Knowledge for Generalizable RGB-Depth Driving Scene Parsing
by: Guo, Sicen, et al.
Published: (2025)