WorldComp2D: Spatio-semantic Representations of Object Identity and Location from Local Views
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, SeongMin, Jeong, Doo Seok |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning
by: Seong, Hyun Seok, et al.
Published: (2026)
by: Seong, Hyun Seok, et al.
Published: (2026)
IterL2Norm: Fast Iterative L2-Normalization
by: Ye, ChangMin, et al.
Published: (2024)
by: Ye, ChangMin, et al.
Published: (2024)
Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
by: Moon, WonJun, et al.
Published: (2026)
by: Moon, WonJun, et al.
Published: (2026)
Compositional Semantics for Open Vocabulary Spatio-semantic Representations
by: Karlsson, Robin, et al.
Published: (2023)
by: Karlsson, Robin, et al.
Published: (2023)
Unsupervised Machine Learning for Detecting and Locating Human-Made Objects in 3D Point Cloud
by: Zhao, Hong, et al.
Published: (2024)
by: Zhao, Hong, et al.
Published: (2024)
GTA: Guided Transfer of Spatial Attention from Object-Centric Representations
by: Seo, SeokHyun, et al.
Published: (2024)
by: Seo, SeokHyun, et al.
Published: (2024)
Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
by: Jun, Youngjun, et al.
Published: (2026)
by: Jun, Youngjun, et al.
Published: (2026)
MVDiff: Scalable and Flexible Multi-View Diffusion for 3D Object Reconstruction from Single-View
by: Bourigault, Emmanuelle, et al.
Published: (2024)
by: Bourigault, Emmanuelle, et al.
Published: (2024)
UOD: Unseen Object Detection in 3D Point Cloud
by: Choi, Hyunjun, et al.
Published: (2024)
by: Choi, Hyunjun, et al.
Published: (2024)
Cross-View World Models
by: Sharma, Rishabh, et al.
Published: (2026)
by: Sharma, Rishabh, et al.
Published: (2026)
3D Reconstruction of Objects in Hands without Real World 3D Supervision
by: Prakash, Aditya, et al.
Published: (2023)
by: Prakash, Aditya, et al.
Published: (2023)
Interacted Object Grounding in Spatio-Temporal Human-Object Interactions
by: Liu, Xiaoyang, et al.
Published: (2024)
by: Liu, Xiaoyang, et al.
Published: (2024)
TeD-Loc: Text Distillation for Weakly Supervised Object Localization
by: Murtaza, Shakeeb, et al.
Published: (2025)
by: Murtaza, Shakeeb, et al.
Published: (2025)
CINA: Conditional Implicit Neural Atlas for Spatio-Temporal Representation of Fetal Brains
by: Dannecker, Maik, et al.
Published: (2024)
by: Dannecker, Maik, et al.
Published: (2024)
DreamSat: Towards a General 3D Model for Novel View Synthesis of Space Objects
by: Mathihalli, Nidhi, et al.
Published: (2024)
by: Mathihalli, Nidhi, et al.
Published: (2024)
DPAC: Distribution-Preserving Adversarial Control for Diffusion Sampling
by: Lee, Han-Jin, et al.
Published: (2025)
by: Lee, Han-Jin, et al.
Published: (2025)
Boosting Object Representation Learning via Motion and Object Continuity
by: Delfosse, Quentin, et al.
Published: (2022)
by: Delfosse, Quentin, et al.
Published: (2022)
Towards Gradient-based Time-Series Explanations through a SpatioTemporal Attention Network
by: Lee, Min Hun
Published: (2024)
by: Lee, Min Hun
Published: (2024)
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
by: Uselis, Arnas, et al.
Published: (2026)
by: Uselis, Arnas, et al.
Published: (2026)
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
by: Lee, Jaewoo, et al.
Published: (2023)
by: Lee, Jaewoo, et al.
Published: (2023)
What Matters in Range View 3D Object Detection
by: Wilson, Benjamin, et al.
Published: (2024)
by: Wilson, Benjamin, et al.
Published: (2024)
Anchors Aweigh! Sail for Optimal Unified Multi-Modal Representations
by: Jeong, Minoh, et al.
Published: (2024)
by: Jeong, Minoh, et al.
Published: (2024)
Local Curvature Smoothing with Stein's Identity for Efficient Score Matching
by: Osada, Genki, et al.
Published: (2024)
by: Osada, Genki, et al.
Published: (2024)
TrackRef3D: Multi-View Consistent Track-then-Label for Open-World Referring Segmentation in 3D Gaussian Splatting
by: Tan, Yuyang, et al.
Published: (2026)
by: Tan, Yuyang, et al.
Published: (2026)
Grouped Discrete Representation for Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2024)
by: Zhao, Rongzhen, et al.
Published: (2024)
Are Object-Centric Representations Better At Compositional Generalization?
by: Kapl, Ferdinand, et al.
Published: (2026)
by: Kapl, Ferdinand, et al.
Published: (2026)
Zero-Shot Object-Centric Representation Learning
by: Didolkar, Aniket, et al.
Published: (2024)
by: Didolkar, Aniket, et al.
Published: (2024)
Identifiable Object Representations under Spatial Ambiguities
by: Kori, Avinash, et al.
Published: (2025)
by: Kori, Avinash, et al.
Published: (2025)
Object-Centric Relational Representations for Image Generation
by: Butera, Luca, et al.
Published: (2023)
by: Butera, Luca, et al.
Published: (2023)
Cross-View Graph Consistency Learning for Invariant Graph Representations
by: Chen, Jie, et al.
Published: (2023)
by: Chen, Jie, et al.
Published: (2023)
Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
by: Burgess, James, et al.
Published: (2023)
by: Burgess, James, et al.
Published: (2023)
Decoupling Spatio-Temporal Adapter for Fine-Grained Badminton Action Localization
by: Wang, Tianyu, et al.
Published: (2026)
by: Wang, Tianyu, et al.
Published: (2026)
Human Gaze Boosts Object-Centered Representation Learning
by: Schaumlöffel, Timothy, et al.
Published: (2025)
by: Schaumlöffel, Timothy, et al.
Published: (2025)
A Geometric View of SRC: Learning Representations for Stable Residual Inference
by: Oikonomou, Vangelis P.
Published: (2026)
by: Oikonomou, Vangelis P.
Published: (2026)
Generalized Open-World Semi-Supervised Object Detection
by: Allabadi, Garvita, et al.
Published: (2023)
by: Allabadi, Garvita, et al.
Published: (2023)
Next day fire prediction via semantic segmentation
by: Alexis, Konstantinos, et al.
Published: (2024)
by: Alexis, Konstantinos, et al.
Published: (2024)
CINeMA: Conditional Implicit Neural Multi-Modal Atlas for a Spatio-Temporal Representation of the Perinatal Brain
by: Dannecker, Maik, et al.
Published: (2025)
by: Dannecker, Maik, et al.
Published: (2025)
Deep Models for Multi-View 3D Object Recognition: A Review
by: Alzahrani, Mona, et al.
Published: (2024)
by: Alzahrani, Mona, et al.
Published: (2024)
Spatio-Temporal Multi-Subgraph GCN for 3D Human Motion Prediction
by: Wang, Jiexin, et al.
Published: (2024)
by: Wang, Jiexin, et al.
Published: (2024)
FMC-DETR: Frequency-Decoupled Multi-Domain Coordination for Aerial-View Object Detection
by: Liang, Ben, et al.
Published: (2025)
by: Liang, Ben, et al.
Published: (2025)
Similar Items
-
From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning
by: Seong, Hyun Seok, et al.
Published: (2026) -
IterL2Norm: Fast Iterative L2-Normalization
by: Ye, ChangMin, et al.
Published: (2024) -
Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
by: Moon, WonJun, et al.
Published: (2026) -
Compositional Semantics for Open Vocabulary Spatio-semantic Representations
by: Karlsson, Robin, et al.
Published: (2023) -
Unsupervised Machine Learning for Detecting and Locating Human-Made Objects in 3D Point Cloud
by: Zhao, Hong, et al.
Published: (2024)