Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Jeon, Yerim, Lee, Miso, Moon, WonJun, Heo, Jae-Pil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Disambiguating 2D-3D Correspondences in Gaussian Splatting-based Feature Fields for Visual Localization
by: Lee, Miso, et al.
Published: (2026)
by: Lee, Miso, et al.
Published: (2026)
Temporally Consistent Long-Term Memory for 3D Single Object Tracking
by: Yoo, Jaejoon, et al.
Published: (2026)
by: Yoo, Jaejoon, et al.
Published: (2026)
From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning
by: Seong, Hyun Seok, et al.
Published: (2026)
by: Seong, Hyun Seok, et al.
Published: (2026)
Selective Contrastive Learning for Weakly Supervised Affordance Grounding
by: Moon, WonJun, et al.
Published: (2025)
by: Moon, WonJun, et al.
Published: (2025)
Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
by: Moon, WonJun, et al.
Published: (2026)
by: Moon, WonJun, et al.
Published: (2026)
Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding
by: Moon, WonJun, et al.
Published: (2023)
by: Moon, WonJun, et al.
Published: (2023)
VLCounter: Text-aware Visual Representation for Zero-Shot Object Counting
by: Kang, Seunggu, et al.
Published: (2023)
by: Kang, Seunggu, et al.
Published: (2023)
Activating Self-Attention for Multi-Scene Absolute Pose Regression
by: Lee, Miso, et al.
Published: (2024)
by: Lee, Miso, et al.
Published: (2024)
Progressive Proxy Anchor Propagation for Unsupervised Semantic Segmentation
by: Seong, Hyun Seok, et al.
Published: (2024)
by: Seong, Hyun Seok, et al.
Published: (2024)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
Auxiliary Descriptive Knowledge for Few-Shot Adaptation of Vision-Language Model
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
Mitigating Background Shift in Class-Incremental Semantic Segmentation
by: Park, Gilhan, et al.
Published: (2024)
by: Park, Gilhan, et al.
Published: (2024)
Mutually-Aware Feature Learning for Few-Shot Object Counting
by: Jeon, Yerim, et al.
Published: (2024)
by: Jeon, Yerim, et al.
Published: (2024)
Long-term Pre-training for Temporal Action Detection with Transformers
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Boundary-Recovering Network for Temporal Action Detection
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
White Aggregation and Restoration for Few-shot 3D Point Cloud Semantic Segmentation
by: Im, Jiyun, et al.
Published: (2025)
by: Im, Jiyun, et al.
Published: (2025)
Looking Beyond the Window: Global-Local Aligned CLIP for Training-free Open-Vocabulary Semantic Segmentation
by: Lee, ByeongCheol, et al.
Published: (2026)
by: Lee, ByeongCheol, et al.
Published: (2026)
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
by: Moon, WonJun, et al.
Published: (2025)
by: Moon, WonJun, et al.
Published: (2025)
Prediction-Feedback DETR for Temporal Action Detection
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Mitigating Semantic Collapse in Partially Relevant Video Retrieval
by: Moon, WonJun, et al.
Published: (2025)
by: Moon, WonJun, et al.
Published: (2025)
GSGAN: Adversarial Learning for Hierarchical Generation of 3D Gaussian Splats
by: Hyun, Sangeek, et al.
Published: (2024)
by: Hyun, Sangeek, et al.
Published: (2024)
Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-Guidance
by: Hur, Jiwan, et al.
Published: (2024)
by: Hur, Jiwan, et al.
Published: (2024)
Noise-free Optimization in Early Training Steps for Image Super-Resolution
by: Lee, MinKyu, et al.
Published: (2023)
by: Lee, MinKyu, et al.
Published: (2023)
SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning
by: Hwang, Chan Yeong, et al.
Published: (2026)
by: Hwang, Chan Yeong, et al.
Published: (2026)
Style Injection in Diffusion: A Training-free Approach for Adapting Large-scale Diffusion Models for Style Transfer
by: Chung, Jiwoo, et al.
Published: (2023)
by: Chung, Jiwoo, et al.
Published: (2023)
Cross-scale Aligned Supervision for Training GANs
by: Hyun, Sangeek, et al.
Published: (2026)
by: Hyun, Sangeek, et al.
Published: (2026)
Auto-Encoded Supervision for Perceptual Image Super-Resolution
by: Lee, MinKyu, et al.
Published: (2024)
by: Lee, MinKyu, et al.
Published: (2024)
PDF-GS: Progressive Distractor Filtering for Robust 3D Gaussian Splatting
by: Seo, Kangmin, et al.
Published: (2026)
by: Seo, Kangmin, et al.
Published: (2026)
Scalable GANs with Transformers
by: Hyun, Sangeek, et al.
Published: (2025)
by: Hyun, Sangeek, et al.
Published: (2025)
SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding
by: Zheng, Hongpei, et al.
Published: (2025)
by: Zheng, Hongpei, et al.
Published: (2025)
Masked Spatial Propagation Network for Sparsity-Adaptive Depth Refinement
by: Jun, Jinyoung, et al.
Published: (2024)
by: Jun, Jinyoung, et al.
Published: (2024)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
by: Liu, Tianhui, et al.
Published: (2026)
by: Liu, Tianhui, et al.
Published: (2026)
SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
by: Huang, Jiaxin, et al.
Published: (2025)
by: Huang, Jiaxin, et al.
Published: (2025)
Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
by: Li, Qiming, et al.
Published: (2025)
by: Li, Qiming, et al.
Published: (2025)
Analyzing the Training Dynamics of Image Restoration Transformers: A Revisit to Layer Normalization
by: Lee, MinKyu, et al.
Published: (2025)
by: Lee, MinKyu, et al.
Published: (2025)
SAM-Guided Masked Token Prediction for 3D Scene Understanding
by: Chen, Zhimin, et al.
Published: (2024)
by: Chen, Zhimin, et al.
Published: (2024)
Diversity-aware Channel Pruning for StyleGAN Compression
by: Chung, Jiwoo, et al.
Published: (2024)
by: Chung, Jiwoo, et al.
Published: (2024)
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
by: Chen, Jiahua, et al.
Published: (2026)
by: Chen, Jiahua, et al.
Published: (2026)
Foreground-Covering Prototype Generation and Matching for SAM-Aided Few-Shot Segmentation
by: Park, Suho, et al.
Published: (2025)
by: Park, Suho, et al.
Published: (2025)
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
by: Ma, Chenyang, et al.
Published: (2024)
by: Ma, Chenyang, et al.
Published: (2024)
Similar Items
-
Disambiguating 2D-3D Correspondences in Gaussian Splatting-based Feature Fields for Visual Localization
by: Lee, Miso, et al.
Published: (2026) -
Temporally Consistent Long-Term Memory for 3D Single Object Tracking
by: Yoo, Jaejoon, et al.
Published: (2026) -
From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning
by: Seong, Hyun Seok, et al.
Published: (2026) -
Selective Contrastive Learning for Weakly Supervised Affordance Grounding
by: Moon, WonJun, et al.
Published: (2025) -
Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
by: Moon, WonJun, et al.
Published: (2026)