Cross-modal State Space Modeling for Real-time RGB-thermal Wild Scene Semantic Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Xiaodong, Lin, Zi'ang, Hu, Luwen, Deng, Zhihong, Liu, Tong, Zhou, Wujie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TUNI: Real-time RGB-T Semantic Segmentation with Unified Multi-Modal Feature Extraction and Cross-Modal Feature Fusion
by: Guo, Xiaodong, et al.
Published: (2025)
by: Guo, Xiaodong, et al.
Published: (2025)
Semantic Scene Segmentation for Robotics
by: Hurtado, Juana Valeria, et al.
Published: (2024)
by: Hurtado, Juana Valeria, et al.
Published: (2024)
TOSS: Real-time Tracking and Moving Object Segmentation for Static Scene Mapping
by: Jang, Seoyeon, et al.
Published: (2024)
by: Jang, Seoyeon, et al.
Published: (2024)
Semantic Segmentation and Scene Reconstruction of RGB-D Image Frames: An End-to-End Modular Pipeline for Robotic Applications
by: Zheng, Zhiwu, et al.
Published: (2024)
by: Zheng, Zhiwu, et al.
Published: (2024)
S3M: Semantic Segmentation Sparse Mapping for UAVs with RGB-D Camera
by: Canh, Thanh Nguyen, et al.
Published: (2024)
by: Canh, Thanh Nguyen, et al.
Published: (2024)
Real-time 3D Semantic Scene Perception for Egocentric Robots with Binocular Vision
by: Nguyen, K., et al.
Published: (2024)
by: Nguyen, K., et al.
Published: (2024)
WildScenes: A Benchmark for 2D and 3D Semantic Segmentation in Large-scale Natural Environments
by: Vidanapathirana, Kavisha, et al.
Published: (2023)
by: Vidanapathirana, Kavisha, et al.
Published: (2023)
FeasibleCap: Real-Time Embodiment Constraint Guidance for In-the-Wild Robot Demonstration Collection
by: Yin, Zi, et al.
Published: (2026)
by: Yin, Zi, et al.
Published: (2026)
Excavating in the Wild: The GOOSE-Ex Dataset for Semantic Segmentation
by: Hagmanns, Raphael, et al.
Published: (2024)
by: Hagmanns, Raphael, et al.
Published: (2024)
RoadFormer: Duplex Transformer for RGB-Normal Semantic Road Scene Parsing
by: Li, Jiahang, et al.
Published: (2023)
by: Li, Jiahang, et al.
Published: (2023)
Complementary Random Masking for RGB-Thermal Semantic Segmentation
by: Shin, Ukcheol, et al.
Published: (2023)
by: Shin, Ukcheol, et al.
Published: (2023)
Caltech Aerial RGB-Thermal Dataset in the Wild
by: Lee, Connor, et al.
Published: (2024)
by: Lee, Connor, et al.
Published: (2024)
VLM-Auto: VLM-based Autonomous Driving Assistant with Human-like Behavior and Understanding for Complex Road Scenes
by: Guo, Ziang, et al.
Published: (2024)
by: Guo, Ziang, et al.
Published: (2024)
METDrive: Multi-modal End-to-end Autonomous Driving with Temporal Guidance
by: Guo, Ziang, et al.
Published: (2024)
by: Guo, Ziang, et al.
Published: (2024)
Privacy-Preserving Semantic Segmentation from Ultra-Low-Resolution RGB Inputs
by: Huang, Xuying, et al.
Published: (2025)
by: Huang, Xuying, et al.
Published: (2025)
DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation
by: Gong, Yan, et al.
Published: (2025)
by: Gong, Yan, et al.
Published: (2025)
VDRive: Leveraging Reinforced VLA and Diffusion Policy for End-to-end Autonomous Driving
by: Guo, Ziang, et al.
Published: (2025)
by: Guo, Ziang, et al.
Published: (2025)
EquiBot: SIM(3)-Equivariant Diffusion Policy for Generalizable and Data Efficient Learning
by: Yang, Jingyun, et al.
Published: (2024)
by: Yang, Jingyun, et al.
Published: (2024)
Wild-Drive: Off-Road Scene Captioning and Path Planning via Robust Multi-modal Routing and Efficient Large Language Model
by: Wang, Zihang, et al.
Published: (2026)
by: Wang, Zihang, et al.
Published: (2026)
HawkDrive: A Transformer-driven Visual Perception System for Autonomous Driving in Night Scene
by: Guo, Ziang, et al.
Published: (2024)
by: Guo, Ziang, et al.
Published: (2024)
SGFormer: Satellite-Ground Fusion for 3D Semantic Scene Completion
by: Guo, Xiyue, et al.
Published: (2025)
by: Guo, Xiyue, et al.
Published: (2025)
SR-SLAM: Scene-reliability Based RGB-D SLAM in Diverse Environments
by: Zhang, Haolan, et al.
Published: (2025)
by: Zhang, Haolan, et al.
Published: (2025)
FEAST: A Flexible Mealtime-Assistance System Towards In-the-Wild Personalization
by: Jenamani, Rajat Kumar, et al.
Published: (2025)
by: Jenamani, Rajat Kumar, et al.
Published: (2025)
BikeScenes: Online LiDAR Semantic Segmentation for Bicycles
by: Goren, Denniz, et al.
Published: (2025)
by: Goren, Denniz, et al.
Published: (2025)
Unveiling the Potential of Segment Anything Model 2 for RGB-Thermal Semantic Segmentation with Language Guidance
by: Zhao, Jiayi, et al.
Published: (2025)
by: Zhao, Jiayi, et al.
Published: (2025)
A Semantic Communication System for Real-time 3D Reconstruction Tasks
by: Zhang, Jiaxing, et al.
Published: (2024)
by: Zhang, Jiaxing, et al.
Published: (2024)
Real-time Monocular 2D and 3D Perception of Endoluminal Scenes for Controlling Flexible Robotic Endoscopic Instruments
by: Wei, Ruofeng, et al.
Published: (2026)
by: Wei, Ruofeng, et al.
Published: (2026)
Multi-modal NeRF Self-Supervision for LiDAR Semantic Segmentation
by: Timoneda, Xavier, et al.
Published: (2024)
by: Timoneda, Xavier, et al.
Published: (2024)
RoboOcc: Enhancing the Geometric and Semantic Scene Understanding for Robots
by: Zhang, Zhang, et al.
Published: (2025)
by: Zhang, Zhang, et al.
Published: (2025)
TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
by: Zhang, Zongzheng, et al.
Published: (2025)
by: Zhang, Zongzheng, et al.
Published: (2025)
Real-time Recognition of Human Interactions from a Single RGB-D Camera for Socially-Aware Robot Navigation
by: Nguyen, Thanh Long, et al.
Published: (2025)
by: Nguyen, Thanh Long, et al.
Published: (2025)
GeomPrompt: Geometric Prompt Learning for RGB-D Semantic Segmentation Under Missing and Degraded Depth
by: Jaganathan, Krishna, et al.
Published: (2026)
by: Jaganathan, Krishna, et al.
Published: (2026)
CEI: A Unified Interface for Cross-Embodiment Visuomotor Policy Learning in 3D Space
by: Wu, Tong, et al.
Published: (2026)
by: Wu, Tong, et al.
Published: (2026)
EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models
by: Yue, Hu, et al.
Published: (2025)
by: Yue, Hu, et al.
Published: (2025)
Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
by: Morin, Sacha, et al.
Published: (2025)
by: Morin, Sacha, et al.
Published: (2025)
Clio: Real-time Task-Driven Open-Set 3D Scene Graphs
by: Maggio, Dominic, et al.
Published: (2024)
by: Maggio, Dominic, et al.
Published: (2024)
MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding
by: Liang, Jing, et al.
Published: (2025)
by: Liang, Jing, et al.
Published: (2025)
CSCPR: Cross-Source-Context Indoor RGB-D Place Recognition
by: Liang, Jing, et al.
Published: (2024)
by: Liang, Jing, et al.
Published: (2024)
Efficient Multi-Task Scene Analysis with RGB-D Transformers
by: Fischedick, Söhnke Benedikt, et al.
Published: (2023)
by: Fischedick, Söhnke Benedikt, et al.
Published: (2023)
ARMOR: Adaptive Meshing with Reinforcement Optimization for Real-time 3D Monitoring in Unexposed Scenes
by: Zhang, Yizhe, et al.
Published: (2025)
by: Zhang, Yizhe, et al.
Published: (2025)
Similar Items
-
TUNI: Real-time RGB-T Semantic Segmentation with Unified Multi-Modal Feature Extraction and Cross-Modal Feature Fusion
by: Guo, Xiaodong, et al.
Published: (2025) -
Semantic Scene Segmentation for Robotics
by: Hurtado, Juana Valeria, et al.
Published: (2024) -
TOSS: Real-time Tracking and Moving Object Segmentation for Static Scene Mapping
by: Jang, Seoyeon, et al.
Published: (2024) -
Semantic Segmentation and Scene Reconstruction of RGB-D Image Frames: An End-to-End Modular Pipeline for Robotic Applications
by: Zheng, Zhiwu, et al.
Published: (2024) -
S3M: Semantic Segmentation Sparse Mapping for UAVs with RGB-D Camera
by: Canh, Thanh Nguyen, et al.
Published: (2024)