DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Gong, Yan, Lu, Jianli, Gao, Yongsheng, Zhao, Jie, Zhang, Xiaojuan, Rahardja, Susanto |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey
by: Gong, Yan, et al.
Published: (2025)
by: Gong, Yan, et al.
Published: (2025)
RoadFormer: Duplex Transformer for RGB-Normal Semantic Road Scene Parsing
by: Li, Jiahang, et al.
Published: (2023)
by: Li, Jiahang, et al.
Published: (2023)
PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding
by: Wen, Junjie, et al.
Published: (2026)
by: Wen, Junjie, et al.
Published: (2026)
Efficient Multi-Task Scene Analysis with RGB-D Transformers
by: Fischedick, Söhnke Benedikt, et al.
Published: (2023)
by: Fischedick, Söhnke Benedikt, et al.
Published: (2023)
Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation
by: Bao, Muyi, et al.
Published: (2026)
by: Bao, Muyi, et al.
Published: (2026)
RGB-only Active 3D Scene Graph Generation for Indoor Mobile Robots
by: Modi, Giorgia, et al.
Published: (2026)
by: Modi, Giorgia, et al.
Published: (2026)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
by: Liang, Wenqi, et al.
Published: (2025)
by: Liang, Wenqi, et al.
Published: (2025)
CSCPR: Cross-Source-Context Indoor RGB-D Place Recognition
by: Liang, Jing, et al.
Published: (2024)
by: Liang, Jing, et al.
Published: (2024)
PixelDiT: Pixel Diffusion Transformers for Image Generation
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
Pixel Motion as Universal Representation for Robot Control
by: Ranasinghe, Kanchana, et al.
Published: (2025)
by: Ranasinghe, Kanchana, et al.
Published: (2025)
Semantic Segmentation and Scene Reconstruction of RGB-D Image Frames: An End-to-End Modular Pipeline for Robotic Applications
by: Zheng, Zhiwu, et al.
Published: (2024)
by: Zheng, Zhiwu, et al.
Published: (2024)
DiFuse-Net: RGB and Dual-Pixel Depth Estimation using Window Bi-directional Parallax Attention and Cross-modal Transfer Learning
by: Swami, Kunal, et al.
Published: (2025)
by: Swami, Kunal, et al.
Published: (2025)
DiffPrompter: Differentiable Implicit Visual Prompts for Semantic-Segmentation in Adverse Conditions
by: Kalwar, Sanket, et al.
Published: (2023)
by: Kalwar, Sanket, et al.
Published: (2023)
Hyp2Former: Hierarchy-Aware Hyperbolic Embeddings for Open-Set Panoptic Segmentation
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
VXP: Voxel-Cross-Pixel Large-scale Image-LiDAR Place Recognition
by: Li, Yun-Jin, et al.
Published: (2024)
by: Li, Yun-Jin, et al.
Published: (2024)
Pixel Motion Diffusion is What We Need for Robot Control
by: Nguyen, E-Ro, et al.
Published: (2025)
by: Nguyen, E-Ro, et al.
Published: (2025)
Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
by: Lee, Seokmin, et al.
Published: (2026)
by: Lee, Seokmin, et al.
Published: (2026)
SLAM-Former: Putting SLAM into One Transformer
by: Yuan, Yijun, et al.
Published: (2025)
by: Yuan, Yijun, et al.
Published: (2025)
GeoReFormer: Geometry-Aware Refinement for Lane Segment Detection and Topology Reasoning
by: Abraham, Danny, et al.
Published: (2026)
by: Abraham, Danny, et al.
Published: (2026)
Pixel-wise Smoothing for Certified Robustness against Camera Motion Perturbations
by: Hu, Hanjiang, et al.
Published: (2023)
by: Hu, Hanjiang, et al.
Published: (2023)
Anyview: Generalizable Indoor 3D Object Detection with Variable Frames
by: Wu, Zhenyu, et al.
Published: (2023)
by: Wu, Zhenyu, et al.
Published: (2023)
OmniIndoor3D: Comprehensive Indoor 3D Reconstruction
by: Wei, Xiaobao, et al.
Published: (2025)
by: Wei, Xiaobao, et al.
Published: (2025)
Cross-modal State Space Modeling for Real-time RGB-thermal Wild Scene Semantic Segmentation
by: Guo, Xiaodong, et al.
Published: (2025)
by: Guo, Xiaodong, et al.
Published: (2025)
Pixels-to-Graph: Real-time Integration of Building Information Models and Scene Graphs for Semantic-Geometric Human-Robot Understanding
by: Longo, Antonello, et al.
Published: (2025)
by: Longo, Antonello, et al.
Published: (2025)
Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
by: Xue, Haoru, et al.
Published: (2025)
by: Xue, Haoru, et al.
Published: (2025)
SutureFormer: Learning Surgical Trajectories via Goal-conditioned Offline RL in Pixel Space
by: Liu, Huanrong, et al.
Published: (2026)
by: Liu, Huanrong, et al.
Published: (2026)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
by: Lin, Yihan, et al.
Published: (2026)
by: Lin, Yihan, et al.
Published: (2026)
Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
by: Hu, Xinggang, et al.
Published: (2026)
by: Hu, Xinggang, et al.
Published: (2026)
W-PoseNet: Dense Correspondence Regularized Pixel Pair Pose Regression
by: Xu, Zelin, et al.
Published: (2019)
by: Xu, Zelin, et al.
Published: (2019)
Robust Surgical Tool Tracking with Pixel-based Probabilities for Projected Geometric Primitives
by: D'Ambrosia, Christopher, et al.
Published: (2024)
by: D'Ambrosia, Christopher, et al.
Published: (2024)
PAGaS: Pixel-Aligned 1DoF Gaussian Splatting for Depth Refinement
by: Recasens, David, et al.
Published: (2026)
by: Recasens, David, et al.
Published: (2026)
InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts
by: Zhong, Weipeng, et al.
Published: (2025)
by: Zhong, Weipeng, et al.
Published: (2025)
Entity-Centric Reinforcement Learning for Object Manipulation from Pixels
by: Haramati, Dan, et al.
Published: (2024)
by: Haramati, Dan, et al.
Published: (2024)
OmniPose6D: Towards Short-Term Object Pose Tracking in Dynamic Scenes from Monocular RGB
by: Lin, Yunzhi, et al.
Published: (2024)
by: Lin, Yunzhi, et al.
Published: (2024)
Configurable Embodied Data Generation for Class-Agnostic RGB-D Video Segmentation
by: Opipari, Anthony, et al.
Published: (2024)
by: Opipari, Anthony, et al.
Published: (2024)
SpaCeFormer: Fast Proposal-Free Open-Vocabulary 3D Instance Segmentation
by: Choy, Chris, et al.
Published: (2026)
by: Choy, Chris, et al.
Published: (2026)
Intelligent Spatial Perception by Building Hierarchical 3D Scene Graphs for Indoor Scenarios with the Help of LLMs
by: Cheng, Yao, et al.
Published: (2025)
by: Cheng, Yao, et al.
Published: (2025)
CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage
by: Liu, Jiale, et al.
Published: (2026)
by: Liu, Jiale, et al.
Published: (2026)
Exploring Object-Aware Attention Guided Frame Association for RGB-D SLAM
by: Caglayan, Ali, et al.
Published: (2025)
by: Caglayan, Ali, et al.
Published: (2025)
SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes
by: Pfaff, Nicholas, et al.
Published: (2026)
by: Pfaff, Nicholas, et al.
Published: (2026)
Similar Items
-
Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey
by: Gong, Yan, et al.
Published: (2025) -
RoadFormer: Duplex Transformer for RGB-Normal Semantic Road Scene Parsing
by: Li, Jiahang, et al.
Published: (2023) -
PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding
by: Wen, Junjie, et al.
Published: (2026) -
Efficient Multi-Task Scene Analysis with RGB-D Transformers
by: Fischedick, Söhnke Benedikt, et al.
Published: (2023) -
Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation
by: Bao, Muyi, et al.
Published: (2026)