3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Chang, Chun-Peng, Pagani, Alain, Stricker, Didier |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding
by: Chang, Chun-Peng, et al.
Published: (2024)
by: Chang, Chun-Peng, et al.
Published: (2024)
SG-PGM: Partial Graph Matching Network with Semantic Geometric Fusion for 3D Scene Graph Alignment and Its Downstream Tasks
by: Xie, Yaxu, et al.
Published: (2024)
by: Xie, Yaxu, et al.
Published: (2024)
G3FA: Geometry-guided GAN for Face Animation
by: Javanmardi, Alireza, et al.
Published: (2024)
by: Javanmardi, Alireza, et al.
Published: (2024)
Uni-SLAM: Uncertainty-Aware Neural Implicit SLAM for Real-Time Dense Indoor Scene Reconstruction
by: Wang, Shaoxiang, et al.
Published: (2024)
by: Wang, Shaoxiang, et al.
Published: (2024)
Inpaint360GS: Efficient Object-Aware 3D Inpainting via Gaussian Splatting for 360° Scenes
by: Wang, Shaoxiang, et al.
Published: (2025)
by: Wang, Shaoxiang, et al.
Published: (2025)
Beyond Averages: Open-Vocabulary 3D Scene Understanding with Gaussian Splatting and Bag of Embeddings
by: Arafa, Abdalla, et al.
Published: (2025)
by: Arafa, Abdalla, et al.
Published: (2025)
EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera
by: Millerdurai, Christen, et al.
Published: (2026)
by: Millerdurai, Christen, et al.
Published: (2026)
EventEgo3D++: 3D Human Motion Capture from a Head-Mounted Event Camera
by: Millerdurai, Christen, et al.
Published: (2025)
by: Millerdurai, Christen, et al.
Published: (2025)
Spatial Preference Rewarding for MLLMs Spatial Understanding
by: Qiu, Han, et al.
Published: (2025)
by: Qiu, Han, et al.
Published: (2025)
Object-Centric 2D Gaussian Splatting: Background Removal and Occlusion-Aware Pruning for Compact Object Models
by: Rogge, Marcel, et al.
Published: (2025)
by: Rogge, Marcel, et al.
Published: (2025)
SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking
by: Khan, Muhammad Saif Ullah, et al.
Published: (2026)
by: Khan, Muhammad Saif Ullah, et al.
Published: (2026)
TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model
by: Javanmardi, Alireza, et al.
Published: (2025)
by: Javanmardi, Alireza, et al.
Published: (2025)
Probing the Reliability of Driving VLMs: From Inconsistent Responses to Grounded Temporal Reasoning
by: Chang, Chun-Peng, et al.
Published: (2026)
by: Chang, Chun-Peng, et al.
Published: (2026)
TinyIceNet: Low-Power SAR Sea Ice Segmentation for On-Board FPGA Inference
by: Koutayni, Mhd Rashed Al, et al.
Published: (2026)
by: Koutayni, Mhd Rashed Al, et al.
Published: (2026)
Invaria: Learning Scale and Density Invariance in Point Clouds via Next-Resolution Prediction
by: Chang, Chun-Peng, et al.
Published: (2026)
by: Chang, Chun-Peng, et al.
Published: (2026)
ReLaGS: Relational Language Gaussian Splatting
by: Xie, Yaxu, et al.
Published: (2026)
by: Xie, Yaxu, et al.
Published: (2026)
Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation
by: Chang, Chun-Peng, et al.
Published: (2025)
by: Chang, Chun-Peng, et al.
Published: (2025)
Towards Unconstrained 2D Pose Estimation of the Human Spine
by: Khan, Muhammad Saif Ullah, et al.
Published: (2025)
by: Khan, Muhammad Saif Ullah, et al.
Published: (2025)
Disambiguating Monocular Reconstruction of 3D Clothed Human with Spatial-Temporal Transformer
by: Deng, Yong, et al.
Published: (2024)
by: Deng, Yong, et al.
Published: (2024)
SF3D-RGB: Scene Flow Estimation from Monocular Camera and Sparse LiDAR
by: Alhimdiat, Rajai, et al.
Published: (2026)
by: Alhimdiat, Rajai, et al.
Published: (2026)
RMS-FlowNet++: Efficient and Robust Multi-Scale Scene Flow Estimation for Large-Scale Point Clouds
by: Battrawy, Ramy, et al.
Published: (2024)
by: Battrawy, Ramy, et al.
Published: (2024)
ShapeAug: Occlusion Augmentation for Event Camera Data
by: Bendig, Katharina, et al.
Published: (2024)
by: Bendig, Katharina, et al.
Published: (2024)
ShapeAug++: More Realistic Shape Augmentation for Event Data
by: Bendig, Katharina, et al.
Published: (2024)
by: Bendig, Katharina, et al.
Published: (2024)
EgoFlowNet: Non-Rigid Scene Flow from Point Clouds with Ego-Motion Support
by: Battrawy, Ramy, et al.
Published: (2024)
by: Battrawy, Ramy, et al.
Published: (2024)
PoseAdapt: Sustainable Human Pose Estimation via Continual Learning Benchmarks and Toolkit
by: Khan, Muhammad Saif Ullah, et al.
Published: (2024)
by: Khan, Muhammad Saif Ullah, et al.
Published: (2024)
MILE: Mixture of Incremental LoRA Experts for Continual Semantic Segmentation across Domains and Modalities
by: Muralidhara, Shishir, et al.
Published: (2026)
by: Muralidhara, Shishir, et al.
Published: (2026)
Sensor Generalization for Adaptive Sensing in Event-based Object Detection via Joint Distribution Training
by: Saha, Aheli, et al.
Published: (2026)
by: Saha, Aheli, et al.
Published: (2026)
Domain-Incremental Semantic Segmentation for Autonomous Driving under Adverse Driving Conditions
by: Muralidhara, Shishir, et al.
Published: (2025)
by: Muralidhara, Shishir, et al.
Published: (2025)
PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion
by: Chamseddine, Mahdi, et al.
Published: (2026)
by: Chamseddine, Mahdi, et al.
Published: (2026)
SAILS: Segment Anything with Incrementally Learned Semantics for Task-Invariant and Training-Free Continual Learning
by: Muralidhara, Shishir, et al.
Published: (2026)
by: Muralidhara, Shishir, et al.
Published: (2026)
DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance
by: Govil, Shreedhar, et al.
Published: (2025)
by: Govil, Shreedhar, et al.
Published: (2025)
ReConText3D: Replay-based Continual Text-to-3D Generation
by: Khan, Muhammad Ahmed Ullah, et al.
Published: (2026)
by: Khan, Muhammad Ahmed Ullah, et al.
Published: (2026)
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
by: Li, Yun, et al.
Published: (2025)
by: Li, Yun, et al.
Published: (2025)
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
SpatialTree: How Spatial Abilities Branch Out in MLLMs
by: Xiao, Yuxi, et al.
Published: (2025)
by: Xiao, Yuxi, et al.
Published: (2025)
A Hybrid Approach for Document Layout Analysis in Document images
by: Shehzadi, Tahira, et al.
Published: (2024)
by: Shehzadi, Tahira, et al.
Published: (2024)
BIMStruct3D: A Fully Automated Hybrid Learning Scan-to-BIM Pipeline with Integrated Topology Refinement
by: Chamseddine, Mahdi, et al.
Published: (2026)
by: Chamseddine, Mahdi, et al.
Published: (2026)
3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
by: Huang, Xiaohu, et al.
Published: (2025)
by: Huang, Xiaohu, et al.
Published: (2025)
JOPP-3D: Joint Open Vocabulary Semantic Segmentation on Point Clouds and Panoramas
by: Inuganti, Sandeep, et al.
Published: (2026)
by: Inuganti, Sandeep, et al.
Published: (2026)
Modality-Incremental Learning with Disjoint Relevance Mapping Networks for Image-based Semantic Segmentation
by: Hegde, Niharika, et al.
Published: (2024)
by: Hegde, Niharika, et al.
Published: (2024)
Similar Items
-
MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding
by: Chang, Chun-Peng, et al.
Published: (2024) -
SG-PGM: Partial Graph Matching Network with Semantic Geometric Fusion for 3D Scene Graph Alignment and Its Downstream Tasks
by: Xie, Yaxu, et al.
Published: (2024) -
G3FA: Geometry-guided GAN for Face Animation
by: Javanmardi, Alireza, et al.
Published: (2024) -
Uni-SLAM: Uncertainty-Aware Neural Implicit SLAM for Real-Time Dense Indoor Scene Reconstruction
by: Wang, Shaoxiang, et al.
Published: (2024) -
Inpaint360GS: Efficient Object-Aware 3D Inpainting via Gaussian Splatting for 360° Scenes
by: Wang, Shaoxiang, et al.
Published: (2025)