PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Wen, Junjie, He, Junlin, Ma, Fei, Cui, Jinqiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Language-Assisted 3D Scene Understanding
by: Wu, Yanmin, et al.
Published: (2023)
by: Wu, Yanmin, et al.
Published: (2023)
TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification
by: Tao, Huaqi, et al.
Published: (2025)
by: Tao, Huaqi, et al.
Published: (2025)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
by: Liang, Wenqi, et al.
Published: (2025)
by: Liang, Wenqi, et al.
Published: (2025)
DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation
by: Gong, Yan, et al.
Published: (2025)
by: Gong, Yan, et al.
Published: (2025)
Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
by: Halacheva, Anna-Maria, et al.
Published: (2024)
by: Halacheva, Anna-Maria, et al.
Published: (2024)
SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
by: Huang, Jiaxin, et al.
Published: (2025)
by: Huang, Jiaxin, et al.
Published: (2025)
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
by: Li, Haoyuan, et al.
Published: (2025)
by: Li, Haoyuan, et al.
Published: (2025)
SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild
by: Rim, Patrick, et al.
Published: (2026)
by: Rim, Patrick, et al.
Published: (2026)
SGFormer: Satellite-Ground Fusion for 3D Semantic Scene Completion
by: Guo, Xiyue, et al.
Published: (2025)
by: Guo, Xiyue, et al.
Published: (2025)
Weather-Robust Scene Semantics with Vision-Aligned 4D Radar
by: Hamilton, Kali, et al.
Published: (2026)
by: Hamilton, Kali, et al.
Published: (2026)
Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation
by: Bao, Muyi, et al.
Published: (2026)
by: Bao, Muyi, et al.
Published: (2026)
DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding
by: Ge, Luzhou, et al.
Published: (2026)
by: Ge, Luzhou, et al.
Published: (2026)
Is Your LiDAR Placement Optimized for 3D Scene Understanding?
by: Li, Ye, et al.
Published: (2024)
by: Li, Ye, et al.
Published: (2024)
SAGE: Scalable Agentic 3D Scene Generation for Embodied AI
by: Xia, Hongchi, et al.
Published: (2026)
by: Xia, Hongchi, et al.
Published: (2026)
QueSTMaps: Queryable Semantic Topological Maps for 3D Scene Understanding
by: Mehan, Yash, et al.
Published: (2024)
by: Mehan, Yash, et al.
Published: (2024)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
by: Wu, Xianjin, et al.
Published: (2026)
by: Wu, Xianjin, et al.
Published: (2026)
PAGaS: Pixel-Aligned 1DoF Gaussian Splatting for Depth Refinement
by: Recasens, David, et al.
Published: (2026)
by: Recasens, David, et al.
Published: (2026)
Overlap-Aware Feature Learning for Robust Unsupervised Domain Adaptation for 3D Semantic Segmentation
by: Chen, Junjie, et al.
Published: (2025)
by: Chen, Junjie, et al.
Published: (2025)
LBurst: Learning-Based Robotic Burst Feature Extraction for 3D Reconstruction in Low Light
by: Ravendran, Ahalya, et al.
Published: (2024)
by: Ravendran, Ahalya, et al.
Published: (2024)
OpenSGA: Efficient 3D Scene Graph Alignment in the Open World
by: Chen, Gang, et al.
Published: (2026)
by: Chen, Gang, et al.
Published: (2026)
Pixel-wise Smoothing for Certified Robustness against Camera Motion Perturbations
by: Hu, Hanjiang, et al.
Published: (2023)
by: Hu, Hanjiang, et al.
Published: (2023)
Embodied Scene Understanding for Vision Language Models via MetaVQA
by: Wang, Weizhen, et al.
Published: (2025)
by: Wang, Weizhen, et al.
Published: (2025)
SGAligner++: Cross-Modal Language-Aided 3D Scene Graph Alignment
by: Singh, Binod, et al.
Published: (2025)
by: Singh, Binod, et al.
Published: (2025)
ZING-3D: Zero-shot Incremental 3D Scene Graphs via Vision-Language Models
by: Saxena, Pranav, et al.
Published: (2025)
by: Saxena, Pranav, et al.
Published: (2025)
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
by: Wu, Yanhao, et al.
Published: (2026)
by: Wu, Yanhao, et al.
Published: (2026)
Scalable 3D Registration via Truncated Entry-wise Absolute Residuals
by: Huang, Tianyu, et al.
Published: (2024)
by: Huang, Tianyu, et al.
Published: (2024)
MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
by: Hou, Minghui, et al.
Published: (2025)
by: Hou, Minghui, et al.
Published: (2025)
SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation
by: Munje, Michael J., et al.
Published: (2025)
by: Munje, Michael J., et al.
Published: (2025)
Efficient Multi-Task Scene Analysis with RGB-D Transformers
by: Fischedick, Söhnke Benedikt, et al.
Published: (2023)
by: Fischedick, Söhnke Benedikt, et al.
Published: (2023)
Object-Scene-Camera Decomposition and Recomposition for Data-Efficient Monocular 3D Object Detection
by: Kuang, Zhaonian, et al.
Published: (2026)
by: Kuang, Zhaonian, et al.
Published: (2026)
Calib3D: Calibrating Model Preferences for Reliable 3D Scene Understanding
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
by: Huang, Huang, et al.
Published: (2025)
by: Huang, Huang, et al.
Published: (2025)
PanoSLAM: Panoptic 3D Scene Reconstruction via Gaussian SLAM
by: Chen, Runnan, et al.
Published: (2024)
by: Chen, Runnan, et al.
Published: (2024)
IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes
by: Zhang, Haochen, et al.
Published: (2025)
by: Zhang, Haochen, et al.
Published: (2025)
Edge-Enabled VIO with Long-Tracked Features for High-Accuracy Low-Altitude IoT Navigation
by: Huang, Xiaohong, et al.
Published: (2025)
by: Huang, Xiaohong, et al.
Published: (2025)
Lightning NeRF: Efficient Hybrid Scene Representation for Autonomous Driving
by: Cao, Junyi, et al.
Published: (2024)
by: Cao, Junyi, et al.
Published: (2024)
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
by: Yang, Yuncong, et al.
Published: (2024)
by: Yang, Yuncong, et al.
Published: (2024)
REACT3D: Recovering Articulations for Interactive Physical 3D Scenes
by: Huang, Zhao, et al.
Published: (2025)
by: Huang, Zhao, et al.
Published: (2025)
arg-VU: Affordance Reasoning with Physics-Aware 3D Geometry for Visual Understanding in Robotic Surgery
by: Xiao, Nan, et al.
Published: (2026)
by: Xiao, Nan, et al.
Published: (2026)
Similar Items
-
Language-Assisted 3D Scene Understanding
by: Wu, Yanmin, et al.
Published: (2023) -
TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification
by: Tao, Huaqi, et al.
Published: (2025) -
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
by: Liang, Wenqi, et al.
Published: (2025) -
DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation
by: Gong, Yan, et al.
Published: (2025) -
Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
by: Halacheva, Anna-Maria, et al.
Published: (2024)