CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Didolkar, Aniket, Zadaianchuk, Andrii, Awal, Rabiul, Seitzer, Maximilian, Gavves, Efstratios, Agrawal, Aishwarya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-Shot Object-Centric Representation Learning
by: Didolkar, Aniket, et al.
Published: (2024)
by: Didolkar, Aniket, et al.
Published: (2024)
Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities
by: Zadaianchuk, Andrii, et al.
Published: (2023)
by: Zadaianchuk, Andrii, et al.
Published: (2023)
Temporally Consistent Object-Centric Learning by Contrasting Slots
by: Manasyan, Anna, et al.
Published: (2024)
by: Manasyan, Anna, et al.
Published: (2024)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
by: Awal, Rabiul, et al.
Published: (2023)
by: Awal, Rabiul, et al.
Published: (2023)
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
by: Zhang, Le, et al.
Published: (2023)
by: Zhang, Le, et al.
Published: (2023)
Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination
by: Barcellona, Leonardo, et al.
Published: (2024)
by: Barcellona, Leonardo, et al.
Published: (2024)
VisMin: Visual Minimal-Change Understanding
by: Awal, Rabiul, et al.
Published: (2024)
by: Awal, Rabiul, et al.
Published: (2024)
Benchmarking Vision Language Models for Cultural Understanding
by: Nayak, Shravan, et al.
Published: (2024)
by: Nayak, Shravan, et al.
Published: (2024)
Are Object-Centric Representations Better At Compositional Generalization?
by: Kapl, Ferdinand, et al.
Published: (2026)
by: Kapl, Ferdinand, et al.
Published: (2026)
Reconstruction by Generation: 3D Multi-Object Scene Reconstruction from Sparse Observations
by: Zadaianchuk, Andrii, et al.
Published: (2026)
by: Zadaianchuk, Andrii, et al.
Published: (2026)
Morpheus: Benchmarking Physical Reasoning of Video Generative Models with Real Physical Experiments
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
by: Gandhi, Sanket, et al.
Published: (2024)
by: Gandhi, Sanket, et al.
Published: (2024)
LoTUS: Large-Scale Machine Unlearning with a Taste of Uncertainty
by: Spartalis, Christoforos N., et al.
Published: (2025)
by: Spartalis, Christoforos N., et al.
Published: (2025)
Unleashing Uncertainty: Efficient Machine Unlearning for Generative AI
by: Spartalis, Christoforos N., et al.
Published: (2025)
by: Spartalis, Christoforos N., et al.
Published: (2025)
Object-Centric Temporal Consistency via Conditional Autoregressive Inductive Biases
by: Meo, Cristian, et al.
Published: (2024)
by: Meo, Cristian, et al.
Published: (2024)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
by: Le, Huy, et al.
Published: (2025)
by: Le, Huy, et al.
Published: (2025)
Explicitly Disentangled Representations in Object-Centric Learning
by: Majellaro, Riccardo, et al.
Published: (2024)
by: Majellaro, Riccardo, et al.
Published: (2024)
GAP3D: Generative Alignment of VLM Latents to Patch-Level Embeddings for 3D Generation
by: Gkotsi, Polytimi Anna, et al.
Published: (2026)
by: Gkotsi, Polytimi Anna, et al.
Published: (2026)
Redefining Normal: A Novel Object-Level Approach for Multi-Object Novelty Detection
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Grounding Continuous Representations in Geometry: Equivariant Neural Fields
by: Wessels, David R, et al.
Published: (2024)
by: Wessels, David R, et al.
Published: (2024)
CarFormer: Self-Driving with Learned Object-Centric Representations
by: Hamdan, Shadi, et al.
Published: (2024)
by: Hamdan, Shadi, et al.
Published: (2024)
From Explainable to Explained AI: Ideas for Falsifying and Quantifying Explanations
by: Schirris, Yoni, et al.
Published: (2025)
by: Schirris, Yoni, et al.
Published: (2025)
VISA: Reasoning Video Object Segmentation via Large Language Models
by: Yan, Cilin, et al.
Published: (2024)
by: Yan, Cilin, et al.
Published: (2024)
An Investigation into Pre-Training Object-Centric Representations for Reinforcement Learning
by: Yoon, Jaesik, et al.
Published: (2023)
by: Yoon, Jaesik, et al.
Published: (2023)
DyST: Towards Dynamic Neural Scene Representations on Real-World Videos
by: Seitzer, Maximilian, et al.
Published: (2023)
by: Seitzer, Maximilian, et al.
Published: (2023)
Object-Centric Latent Action Learning
by: Klepach, Albina, et al.
Published: (2025)
by: Klepach, Albina, et al.
Published: (2025)
Transparent Visual Reasoning via Object-Centric Agent Collaboration
by: Teoh, Benjamin, et al.
Published: (2025)
by: Teoh, Benjamin, et al.
Published: (2025)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
by: Chi, Donghwan, et al.
Published: (2025)
by: Chi, Donghwan, et al.
Published: (2025)
Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-shot Semantic Segmentation
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
Assessing and Learning Alignment of Unimodal Vision and Language Models
by: Zhang, Le, et al.
Published: (2024)
by: Zhang, Le, et al.
Published: (2024)
Scaling Language-Centric Omnimodal Representation Learning
by: Xiao, Chenghao, et al.
Published: (2025)
by: Xiao, Chenghao, et al.
Published: (2025)
SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
Hierarchical Object-Centric Learning with Capsule Networks
by: Renzulli, Riccardo
Published: (2024)
by: Renzulli, Riccardo
Published: (2024)
RiT: Vanilla Diffusion Transformers Suffice in Representation Space
by: Zhang, Le, et al.
Published: (2026)
by: Zhang, Le, et al.
Published: (2026)
Improving Automatic VQA Evaluation Using Large Language Models
by: Mañas, Oscar, et al.
Published: (2023)
by: Mañas, Oscar, et al.
Published: (2023)
Discovering Failure Modes in Vision-Language Models using RL
by: Jain, Kanishk, et al.
Published: (2026)
by: Jain, Kanishk, et al.
Published: (2026)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
by: Ahmadi, Saba, et al.
Published: (2023)
by: Ahmadi, Saba, et al.
Published: (2023)
Grounded Object Centric Learning
by: Kori, Avinash, et al.
Published: (2023)
by: Kori, Avinash, et al.
Published: (2023)
Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection
by: Chandhok, Shivam, et al.
Published: (2025)
by: Chandhok, Shivam, et al.
Published: (2025)
Similar Items
-
Zero-Shot Object-Centric Representation Learning
by: Didolkar, Aniket, et al.
Published: (2024) -
Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities
by: Zadaianchuk, Andrii, et al.
Published: (2023) -
Temporally Consistent Object-Centric Learning by Contrasting Slots
by: Manasyan, Anna, et al.
Published: (2024) -
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
by: Awal, Rabiul, et al.
Published: (2023) -
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
by: Zhang, Le, et al.
Published: (2023)