Cog3DMap: Multi-View Vision-Language Reasoning with 3D Cognitive Maps
Fuente:
arXiv
Saved in:
| Main Authors: | Gwak, Chanyoung, Jeong, Yoonwoo, Jeon, Byungwoo, Lee, Hyunseok, Shin, Jinwoo, Cho, Minsu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision-aligned Latent Reasoning for Multi-modal Large Language Model
by: Jeon, Byungwoo, et al.
Published: (2026)
by: Jeon, Byungwoo, et al.
Published: (2026)
NVS-Adapter: Plug-and-Play Novel View Synthesis from a Single Image
by: Jeong, Yoonwoo, et al.
Published: (2023)
by: Jeong, Yoonwoo, et al.
Published: (2023)
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
by: Jeon, Byungwoo, et al.
Published: (2026)
by: Jeon, Byungwoo, et al.
Published: (2026)
RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos
by: Lee, Junmyeong, et al.
Published: (2024)
by: Lee, Junmyeong, et al.
Published: (2024)
Quantile Rendering: Efficiently Embedding High-dimensional Feature on 3D Gaussian Splatting
by: Jeong, Yoonwoo, et al.
Published: (2025)
by: Jeong, Yoonwoo, et al.
Published: (2025)
MV-SAM: Multi-view Promptable Segmentation using Pointmap Guidance
by: Jeong, Yoonwoo, et al.
Published: (2026)
by: Jeong, Yoonwoo, et al.
Published: (2026)
M3DMap: Object-aware Multimodal 3D Mapping for Dynamic Environments
by: Yudin, Dmitry
Published: (2025)
by: Yudin, Dmitry
Published: (2025)
CogME: A Cognition-Inspired Multi-Dimensional Evaluation Metric for Story Understanding
by: Shin, Minjung, et al.
Published: (2021)
by: Shin, Minjung, et al.
Published: (2021)
EVT: Efficient View Transformation for Multi-Modal 3D Object Detection
by: Lee, Yongjin, et al.
Published: (2024)
by: Lee, Yongjin, et al.
Published: (2024)
PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
by: Jung, Seongmin, et al.
Published: (2025)
by: Jung, Seongmin, et al.
Published: (2025)
Affostruction: 3D Affordance Grounding with Generative Reconstruction
by: Park, Chunghyun, et al.
Published: (2026)
by: Park, Chunghyun, et al.
Published: (2026)
Leveraging 3D Geometric Priors in 2D Rotation Symmetry Detection
by: Seo, Ahyun, et al.
Published: (2025)
by: Seo, Ahyun, et al.
Published: (2025)
DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning
by: Lee, Junha, et al.
Published: (2026)
by: Lee, Junha, et al.
Published: (2026)
3D Equivariant Pose Regression via Direct Wigner-D Harmonics Prediction
by: Lee, Jongmin, et al.
Published: (2024)
by: Lee, Jongmin, et al.
Published: (2024)
Thickness-aware E(3)-Equivariant 3D Mesh Neural Networks
by: Kim, Sungwon, et al.
Published: (2025)
by: Kim, Sungwon, et al.
Published: (2025)
Harnessing the Power of Training-Free Techniques in Text-to-2D Generation for Text-to-3D Generation via Score Distillation Sampling
by: Lee, Junhong, et al.
Published: (2025)
by: Lee, Junhong, et al.
Published: (2025)
Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors
by: Jeon, Subin, et al.
Published: (2025)
by: Jeon, Subin, et al.
Published: (2025)
CorrespondentDream: Enhancing 3D Fidelity of Text-to-3D using Cross-View Correspondences
by: Kim, Seungwook, et al.
Published: (2024)
by: Kim, Seungwook, et al.
Published: (2024)
ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
DreamFlow: High-Quality Text-to-3D Generation by Approximating Probability Flow
by: Lee, Kyungmin, et al.
Published: (2024)
by: Lee, Kyungmin, et al.
Published: (2024)
uCLIP: Parameter-Efficient Multilingual Extension of Vision-Language Models with Unpaired Data
by: Chung, Dahyun, et al.
Published: (2025)
by: Chung, Dahyun, et al.
Published: (2025)
Cog2Gen3D: Sculpturing 3D Semantic-Geometric Cognition for 3D Generation
by: Wang, Haonan, et al.
Published: (2026)
by: Wang, Haonan, et al.
Published: (2026)
Multi-view Image Prompted Multi-view Diffusion for Improved 3D Generation
by: Kim, Seungwook, et al.
Published: (2024)
by: Kim, Seungwook, et al.
Published: (2024)
Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
by: Lee, Kyungmin, et al.
Published: (2025)
by: Lee, Kyungmin, et al.
Published: (2025)
Learning Multi-View Spatial Reasoning from Cross-View Relations
by: Jeong, Suchae, et al.
Published: (2026)
by: Jeong, Suchae, et al.
Published: (2026)
Point2Pose: A Generative Framework for 3D Human Pose Estimation with Multi-View Point Cloud Dataset
by: Lee, Hyunsoo, et al.
Published: (2025)
by: Lee, Hyunsoo, et al.
Published: (2025)
Video Summarization with Large Language Models
by: Lee, Min Jung, et al.
Published: (2025)
by: Lee, Min Jung, et al.
Published: (2025)
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis
by: Kim, Seungwook, et al.
Published: (2025)
by: Kim, Seungwook, et al.
Published: (2025)
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
by: Won, John, et al.
Published: (2025)
by: Won, John, et al.
Published: (2025)
TouchMap-OR: Multi-View 3D Mapping of Hand-Surface Contacts
by: Ktistakis, Sophokles, et al.
Published: (2026)
by: Ktistakis, Sophokles, et al.
Published: (2026)
SpaCeFormer: Fast Proposal-Free Open-Vocabulary 3D Instance Segmentation
by: Choy, Chris, et al.
Published: (2026)
by: Choy, Chris, et al.
Published: (2026)
CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion
by: Zheng, Wendi, et al.
Published: (2024)
by: Zheng, Wendi, et al.
Published: (2024)
Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation
by: Lee, Junha, et al.
Published: (2025)
by: Lee, Junha, et al.
Published: (2025)
Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
by: Kim, Dongkeun, et al.
Published: (2025)
by: Kim, Dongkeun, et al.
Published: (2025)
CogniMap3D: Cognitive 3D Mapping and Rapid Retrieval
by: Wang, Feiran, et al.
Published: (2026)
by: Wang, Feiran, et al.
Published: (2026)
Hierarchically Structured Neural Bones for Reconstructing Animatable Objects from Casual Videos
by: Jeon, Subin, et al.
Published: (2024)
by: Jeon, Subin, et al.
Published: (2024)
CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition
by: Yang, Hongji, et al.
Published: (2026)
by: Yang, Hongji, et al.
Published: (2026)
View-on-Graph: Zero-shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs
by: Liu, Yuanyuan, et al.
Published: (2025)
by: Liu, Yuanyuan, et al.
Published: (2025)
3D Geometric Shape Assembly via Efficient Point Cloud Matching
by: Lee, Nahyuk, et al.
Published: (2024)
by: Lee, Nahyuk, et al.
Published: (2024)
Similar Items
-
Vision-aligned Latent Reasoning for Multi-modal Large Language Model
by: Jeon, Byungwoo, et al.
Published: (2026) -
NVS-Adapter: Plug-and-Play Novel View Synthesis from a Single Image
by: Jeong, Yoonwoo, et al.
Published: (2023) -
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
by: Jeon, Byungwoo, et al.
Published: (2026) -
RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos
by: Lee, Junmyeong, et al.
Published: (2024) -
Quantile Rendering: Efficiently Embedding High-dimensional Feature on 3D Gaussian Splatting
by: Jeong, Yoonwoo, et al.
Published: (2025)