Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Gandhi, Sanket, Atul, Mahajan, Samanyu, Sharma, Vishal, Gupta, Rushil, Mondal, Arnab Kumar, Singla, Parag |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Global Object-Centric Representations via Disentangled Slot Attention
by: Chen, Tonglin, et al.
Published: (2024)
by: Chen, Tonglin, et al.
Published: (2024)
Explicitly Disentangled Representations in Object-Centric Learning
by: Majellaro, Riccardo, et al.
Published: (2024)
by: Majellaro, Riccardo, et al.
Published: (2024)
Disentangled Object-Centric Image Representation for Robotic Manipulation
by: Emukpere, David, et al.
Published: (2025)
by: Emukpere, David, et al.
Published: (2025)
DOCTR: Disentangled Object-Centric Transformer for Point Scene Understanding
by: Yu, Xiaoxuan, et al.
Published: (2024)
by: Yu, Xiaoxuan, et al.
Published: (2024)
DRISHTIKON: Visual Grounding at Multiple Granularities in Documents
by: Kasuba, Badri Vishal, et al.
Published: (2025)
by: Kasuba, Badri Vishal, et al.
Published: (2025)
Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders
by: Kumar, Amandeep, et al.
Published: (2026)
by: Kumar, Amandeep, et al.
Published: (2026)
Fine-Tuning Without Forgetting: Adaptation of YOLOv8 Preserves COCO Performance
by: Gandhi, Vishal, et al.
Published: (2025)
by: Gandhi, Vishal, et al.
Published: (2025)
Personalized Image Generation from an Author Writing Style
by: Gandhi, Sagar, et al.
Published: (2025)
by: Gandhi, Sagar, et al.
Published: (2025)
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
Mimicking Human Visual Development for Learning Robust Image Representations
by: Raj, Ankita, et al.
Published: (2025)
by: Raj, Ankita, et al.
Published: (2025)
Context Matters: Learning Global Semantics via Object-Centric Representation
by: Zhong, Jike, et al.
Published: (2025)
by: Zhong, Jike, et al.
Published: (2025)
Grouped Discrete Representation for Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2024)
by: Zhao, Rongzhen, et al.
Published: (2024)
Zero-Shot Object-Centric Representation Learning
by: Didolkar, Aniket, et al.
Published: (2024)
by: Didolkar, Aniket, et al.
Published: (2024)
Object-Centric Representation Learning for Enhanced 3D Semantic Scene Graph Prediction
by: Heo, KunHo, et al.
Published: (2025)
by: Heo, KunHo, et al.
Published: (2025)
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
by: Le, Huy, et al.
Published: (2025)
by: Le, Huy, et al.
Published: (2025)
Top-Down Guidance for Learning Object-Centric Representations
by: Zou, Junhong, et al.
Published: (2024)
by: Zou, Junhong, et al.
Published: (2024)
Organized Grouped Discrete Representation for Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2024)
by: Zhao, Rongzhen, et al.
Published: (2024)
LatentGAN Autoencoder: Learning Disentangled Latent Distribution
by: Kalwar, Sanket, et al.
Published: (2022)
by: Kalwar, Sanket, et al.
Published: (2022)
YOLO-LAN: Precise Polyp Detection via Optimized Loss, Augmentations and Negatives
by: Gupta, Siddharth, et al.
Published: (2025)
by: Gupta, Siddharth, et al.
Published: (2025)
Crisp Attention: Regularizing Transformers via Structured Sparsity
by: Gandhi, Sagar, et al.
Published: (2025)
by: Gandhi, Sagar, et al.
Published: (2025)
Grouped Discrete Representation Guides Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2024)
by: Zhao, Rongzhen, et al.
Published: (2024)
Spectral State Space Model for Rotation-Invariant Visual Representation Learning
by: Dastani, Sahar, et al.
Published: (2025)
by: Dastani, Sahar, et al.
Published: (2025)
Sequential Representation Learning via Static-Dynamic Conditional Disentanglement
by: Simon, Mathieu Cyrille, et al.
Published: (2024)
by: Simon, Mathieu Cyrille, et al.
Published: (2024)
HeightFormer: Learning Height Prediction in Voxel Features for Roadside Vision Centric 3D Object Detection via Transformer
by: Zhang, Zhang, et al.
Published: (2025)
by: Zhang, Zhang, et al.
Published: (2025)
OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
by: Song, Yeon-Ji, et al.
Published: (2024)
by: Song, Yeon-Ji, et al.
Published: (2024)
Learning Object-Centric Representations Based on Slots in Real World Scenarios
by: Akan, Adil Kaan
Published: (2025)
by: Akan, Adil Kaan
Published: (2025)
Transparent Visual Reasoning via Object-Centric Agent Collaboration
by: Teoh, Benjamin, et al.
Published: (2025)
by: Teoh, Benjamin, et al.
Published: (2025)
CarFormer: Self-Driving with Learned Object-Centric Representations
by: Hamdan, Shadi, et al.
Published: (2024)
by: Hamdan, Shadi, et al.
Published: (2024)
BayesSDF: Surface-Based Laplacian Uncertainty Estimation for 3D Geometry with Neural Signed Distance Fields
by: Desai, Rushil
Published: (2025)
by: Desai, Rushil
Published: (2025)
Learning Physical Dynamics for Object-centric Visual Prediction
by: Xu, Huilin, et al.
Published: (2024)
by: Xu, Huilin, et al.
Published: (2024)
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
by: Villar-Corrales, Angel, et al.
Published: (2025)
by: Villar-Corrales, Angel, et al.
Published: (2025)
An Investigation into Pre-Training Object-Centric Representations for Reinforcement Learning
by: Yoon, Jaesik, et al.
Published: (2023)
by: Yoon, Jaesik, et al.
Published: (2023)
What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization
by: Yoshihashi, Ryota, et al.
Published: (2026)
by: Yoshihashi, Ryota, et al.
Published: (2026)
Are Object-Centric Representations Better At Compositional Generalization?
by: Kapl, Ferdinand, et al.
Published: (2026)
by: Kapl, Ferdinand, et al.
Published: (2026)
Object-Centric Relational Representations for Image Generation
by: Butera, Luca, et al.
Published: (2023)
by: Butera, Luca, et al.
Published: (2023)
Learning Object-Centric Representations in SAR Images with Multi-Level Feature Fusion
by: Jang, Oh-Tae, et al.
Published: (2025)
by: Jang, Oh-Tae, et al.
Published: (2025)
Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting
by: Hsu, Tsuheng, et al.
Published: (2026)
by: Hsu, Tsuheng, et al.
Published: (2026)
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
by: Das, Swadhin, et al.
Published: (2025)
by: Das, Swadhin, et al.
Published: (2025)
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models
by: Mamaghan, Amir Mohammad Karimi, et al.
Published: (2024)
by: Mamaghan, Amir Mohammad Karimi, et al.
Published: (2024)
Object-Centric Action-Enhanced Representations for Robot Visuo-Motor Policy Learning
by: Giannakakis, Nikos, et al.
Published: (2025)
by: Giannakakis, Nikos, et al.
Published: (2025)
Similar Items
-
Learning Global Object-Centric Representations via Disentangled Slot Attention
by: Chen, Tonglin, et al.
Published: (2024) -
Explicitly Disentangled Representations in Object-Centric Learning
by: Majellaro, Riccardo, et al.
Published: (2024) -
Disentangled Object-Centric Image Representation for Robotic Manipulation
by: Emukpere, David, et al.
Published: (2025) -
DOCTR: Disentangled Object-Centric Transformer for Point Scene Understanding
by: Yu, Xiaoxuan, et al.
Published: (2024) -
DRISHTIKON: Visual Grounding at Multiple Granularities in Documents
by: Kasuba, Badri Vishal, et al.
Published: (2025)