Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xun, Cruz, Rodrigo Santa, Xi, Mingze, Zhang, Hu, Perera, Madhawa, Wang, Ziwei, Ravendran, Ahalya, Matthews, Brandon J., Xu, Feng, Adcock, Matt, Wang, Dadong, Liu, Jiajun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Appraisal-Based Approach to Human-Centred Explanations
by: Somarathna, Rukshani, et al.
Published: (2025)
by: Somarathna, Rukshani, et al.
Published: (2025)
LBurst: Learning-Based Robotic Burst Feature Extraction for 3D Reconstruction in Low Light
by: Ravendran, Ahalya, et al.
Published: (2024)
by: Ravendran, Ahalya, et al.
Published: (2024)
QueSTMaps: Queryable Semantic Topological Maps for 3D Scene Understanding
by: Mehan, Yash, et al.
Published: (2024)
by: Mehan, Yash, et al.
Published: (2024)
SODIUM: From Open Web Data to Queryable Databases
by: Hu, Chuxuan, et al.
Published: (2026)
by: Hu, Chuxuan, et al.
Published: (2026)
Representation-Centric Survey of Supervised Skeletal Action Recognition and the New Benchmark
by: Liu, Yang, et al.
Published: (2022)
by: Liu, Yang, et al.
Published: (2022)
ESceme: Vision-and-Language Navigation with Episodic Scene Memory
by: Zheng, Qi, et al.
Published: (2023)
by: Zheng, Qi, et al.
Published: (2023)
CHARM: Collaborative Harmonization across Arbitrary Modalities for Modality-agnostic Semantic Segmentation
by: Wen, Lekang, et al.
Published: (2025)
by: Wen, Lekang, et al.
Published: (2025)
GeoNDC: A Queryable Neural Data Cube for Planetary-Scale Earth Observation
by: Qi, Jianbo, et al.
Published: (2026)
by: Qi, Jianbo, et al.
Published: (2026)
SGC-VQGAN: Towards Complex Scene Representation via Semantic Guided Clustering Codebook
by: Ding, Chenjing, et al.
Published: (2024)
by: Ding, Chenjing, et al.
Published: (2024)
Learning Modality-agnostic Representation for Semantic Segmentation from Any Modalities
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Open3DSG: Open-Vocabulary 3D Scene Graphs from Point Clouds with Queryable Objects and Open-Set Relationships
by: Koch, Sebastian, et al.
Published: (2024)
by: Koch, Sebastian, et al.
Published: (2024)
The Scene Language: Representing Scenes with Programs, Words, and Embeddings
by: Zhang, Yunzhi, et al.
Published: (2024)
by: Zhang, Yunzhi, et al.
Published: (2024)
Facial Expression Recognition with Controlled Privacy Preservation and Feature Compensation
by: Xu, Feng, et al.
Published: (2024)
by: Xu, Feng, et al.
Published: (2024)
Interpretable Zero-shot Referring Expression Comprehension with Query-driven Scene Graphs
by: Wu, Yike, et al.
Published: (2026)
by: Wu, Yike, et al.
Published: (2026)
Infini-News: Efficiently Queryable Access to 1.3 Billion Processed Common Crawl News Articles
by: Lazzaroni, Ruggero Marino, et al.
Published: (2026)
by: Lazzaroni, Ruggero Marino, et al.
Published: (2026)
Superpixel Semantics Representation and Pre-training for Vision-Language Task
by: Zhang, Siyu, et al.
Published: (2023)
by: Zhang, Siyu, et al.
Published: (2023)
Planning with the Views via Scene Self-Exploration
by: Wang, Kangrui, et al.
Published: (2026)
by: Wang, Kangrui, et al.
Published: (2026)
The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
by: Fan, Weichen, et al.
Published: (2025)
by: Fan, Weichen, et al.
Published: (2025)
Queryable LoRA: Instruction-Regularized Routing Over Shared Low-Rank Update Atoms
by: Vaidya, Omatharv Bharat, et al.
Published: (2026)
by: Vaidya, Omatharv Bharat, et al.
Published: (2026)
SemanticFormer: Holistic and Semantic Traffic Scene Representation for Trajectory Prediction using Knowledge Graphs
by: Sun, Zhigang, et al.
Published: (2024)
by: Sun, Zhigang, et al.
Published: (2024)
Knowledge Priors for Identity-Disentangled Open-Set Privacy-Preserving Video FER
by: Xu, Feng, et al.
Published: (2026)
by: Xu, Feng, et al.
Published: (2026)
Cross-Modal Scene Semantic Alignment for Image Complexity Assessment
by: Luo, Yuqing, et al.
Published: (2025)
by: Luo, Yuqing, et al.
Published: (2025)
Queryable Prototype Multiple Instance Learning with Vision-Language Models for Incremental Whole Slide Image Classification
by: Gou, Jiaxiang, et al.
Published: (2024)
by: Gou, Jiaxiang, et al.
Published: (2024)
Scene Abstraction for Lexical Semantics: Structured Representations of Situated Meaning
by: Cho, Yejin, et al.
Published: (2026)
by: Cho, Yejin, et al.
Published: (2026)
Vision-Language Models Create Cross-Modal Task Representations
by: Luo, Grace, et al.
Published: (2024)
by: Luo, Grace, et al.
Published: (2024)
Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion
by: Li, Bohan, et al.
Published: (2023)
by: Li, Bohan, et al.
Published: (2023)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
Visual Superordinate Abstraction for Robust Concept Learning
by: Zheng, Qi, et al.
Published: (2022)
by: Zheng, Qi, et al.
Published: (2022)
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
by: Bi, Jinhe, et al.
Published: (2024)
by: Bi, Jinhe, et al.
Published: (2024)
What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
PFSD: A Multi-Modal Pedestrian-Focus Scene Dataset for Rich Tasks in Semi-Structured Environments
by: Liu, Yueting, et al.
Published: (2025)
by: Liu, Yueting, et al.
Published: (2025)
The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities
by: Wu, Zhaofeng, et al.
Published: (2024)
by: Wu, Zhaofeng, et al.
Published: (2024)
Aquatic-GS: A Hybrid 3D Representation for Underwater Scenes
by: Liu, Shaohua, et al.
Published: (2024)
by: Liu, Shaohua, et al.
Published: (2024)
Vision-based 3D Semantic Scene Completion via Capture Dynamic Representations
by: Wang, Meng, et al.
Published: (2025)
by: Wang, Meng, et al.
Published: (2025)
General Scene Adaptation for Vision-and-Language Navigation
by: Hong, Haodong, et al.
Published: (2025)
by: Hong, Haodong, et al.
Published: (2025)
Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
Estimating Pasture Biomass from Top-View Images: A Dataset for Precision Agriculture
by: Liao, Qiyu, et al.
Published: (2025)
by: Liao, Qiyu, et al.
Published: (2025)
Two-Stream Interactive Joint Learning of Scene Parsing and Geometric Vision Tasks
by: Tang, Guanfeng, et al.
Published: (2026)
by: Tang, Guanfeng, et al.
Published: (2026)
Manipulable Semantic Components: a Computational Representation of Data Visualization Scenes
by: Liu, Zhicheng, et al.
Published: (2024)
by: Liu, Zhicheng, et al.
Published: (2024)
Global-Recent Semantic Reasoning on Dynamic Text-Attributed Graphs with Large Language Models
by: Wang, Yunan, et al.
Published: (2025)
by: Wang, Yunan, et al.
Published: (2025)
Similar Items
-
An Appraisal-Based Approach to Human-Centred Explanations
by: Somarathna, Rukshani, et al.
Published: (2025) -
LBurst: Learning-Based Robotic Burst Feature Extraction for 3D Reconstruction in Low Light
by: Ravendran, Ahalya, et al.
Published: (2024) -
QueSTMaps: Queryable Semantic Topological Maps for 3D Scene Understanding
by: Mehan, Yash, et al.
Published: (2024) -
SODIUM: From Open Web Data to Queryable Databases
by: Hu, Chuxuan, et al.
Published: (2026) -
Representation-Centric Survey of Supervised Skeletal Action Recognition and the New Benchmark
by: Liu, Yang, et al.
Published: (2022)