Successes and Limitations of Object-centric Models at Compositional Generalisation
Fuente:
arXiv
Saved in:
| Main Authors: | Montero, Milton L., Bowers, Jeffrey S., Malhotra, Gaurav |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lost in Latent Space: Disentangled Models and the Challenge of Combinatorial Generalisation
by: Montero, Milton L., et al.
Published: (2022)
by: Montero, Milton L., et al.
Published: (2022)
SlotPi: Physics-informed Object-centric Reasoning Models
by: Li, Jian, et al.
Published: (2025)
by: Li, Jian, et al.
Published: (2025)
DORSal: Diffusion for Object-centric Representations of Scenes et al
by: Jabri, Allan, et al.
Published: (2023)
by: Jabri, Allan, et al.
Published: (2023)
Weakly Supervised Concept Learning for Object-centric Visual Reasoning
by: Tiwari, Sparsh, et al.
Published: (2026)
by: Tiwari, Sparsh, et al.
Published: (2026)
Edge-case Synthesis for Fisheye Object Detection: A Data-centric Perspective
by: Kim, Seunghyeon, et al.
Published: (2025)
by: Kim, Seunghyeon, et al.
Published: (2025)
SlotLifter: Slot-guided Feature Lifting for Learning Object-centric Radiance Fields
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
by: Bendikas, Rokas, et al.
Published: (2025)
by: Bendikas, Rokas, et al.
Published: (2025)
Generative AI in Vision: A Survey on Models, Metrics and Applications
by: Raut, Gaurav, et al.
Published: (2024)
by: Raut, Gaurav, et al.
Published: (2024)
MindSet: Vision. A toolbox for testing DNNs on key psychological experiments
by: Biscione, Valerio, et al.
Published: (2024)
by: Biscione, Valerio, et al.
Published: (2024)
VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models
by: Barreto, Jesimon, et al.
Published: (2025)
by: Barreto, Jesimon, et al.
Published: (2025)
EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene Supervision
by: Chen, Jiahao, et al.
Published: (2026)
by: Chen, Jiahao, et al.
Published: (2026)
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Vehicle-centric Perception via Multimodal Structured Pre-training
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
Region-centric Image-Language Pretraining for Open-Vocabulary Detection
by: Kim, Dahun, et al.
Published: (2023)
by: Kim, Dahun, et al.
Published: (2023)
Towards Generalisable Time Series Understanding Across Domains
by: Turgut, Özgün, et al.
Published: (2024)
by: Turgut, Özgün, et al.
Published: (2024)
Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation Data
by: Zhu, Xun, et al.
Published: (2025)
by: Zhu, Xun, et al.
Published: (2025)
Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs
by: Gautam, Somraj, et al.
Published: (2025)
by: Gautam, Somraj, et al.
Published: (2025)
LucidPPN: Unambiguous Prototypical Parts Network for User-centric Interpretable Computer Vision
by: Pach, Mateusz, et al.
Published: (2024)
by: Pach, Mateusz, et al.
Published: (2024)
A Self-Supervised Framework for Improved Generalisability in Ultrasound B-mode Image Segmentation
by: Ellis, Edward, et al.
Published: (2025)
by: Ellis, Edward, et al.
Published: (2025)
MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts
by: Zhu, Jie, et al.
Published: (2024)
by: Zhu, Jie, et al.
Published: (2024)
Object-centric 3D Motion Field for Robot Learning from Human Videos
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
XEdgeAI: A Human-centered Industrial Inspection Framework with Data-centric Explainable Edge AI Approach
by: Nguyen, Truong Thanh Hung, et al.
Published: (2024)
by: Nguyen, Truong Thanh Hung, et al.
Published: (2024)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
by: Lin, Ronghao, et al.
Published: (2024)
by: Lin, Ronghao, et al.
Published: (2024)
Green Screen Augmentation Enables Scene Generalisation in Robotic Manipulation
by: Teoh, Eugene, et al.
Published: (2024)
by: Teoh, Eugene, et al.
Published: (2024)
LAuReL: Learned Augmented Residual Layer
by: Menghani, Gaurav, et al.
Published: (2024)
by: Menghani, Gaurav, et al.
Published: (2024)
PlantDiseaseNet-RT50: A Fine-tuned ResNet50 Architecture for High-Accuracy Plant Disease Detection Beyond Standard CNNs
by: Sagnika, Santwana, et al.
Published: (2025)
by: Sagnika, Santwana, et al.
Published: (2025)
Exploring Perceptual Limitation of Multimodal Large Language Models
by: Zhang, Jiarui, et al.
Published: (2024)
by: Zhang, Jiarui, et al.
Published: (2024)
Interaction Field Matching: Overcoming Limitations of Electrostatic Models
by: Manukhov, Stepan I., et al.
Published: (2025)
by: Manukhov, Stepan I., et al.
Published: (2025)
From Fewer Samples to Fewer Bits: Reframing Dataset Distillation as Joint Optimization of Precision and Compactness
by: Dinh, My H., et al.
Published: (2026)
by: Dinh, My H., et al.
Published: (2026)
Spatial Transformer Network YOLO Model for Agricultural Object Detection
by: Zambre, Yash, et al.
Published: (2024)
by: Zambre, Yash, et al.
Published: (2024)
Enhancing Novel Object Detection via Cooperative Foundational Models
by: Bharadwaj, Rohit, et al.
Published: (2023)
by: Bharadwaj, Rohit, et al.
Published: (2023)
Compositional Entailment Learning for Hyperbolic Vision-Language Models
by: Pal, Avik, et al.
Published: (2024)
by: Pal, Avik, et al.
Published: (2024)
Dreamweaver: Learning Compositional World Models from Pixels
by: Baek, Junyeob, et al.
Published: (2025)
by: Baek, Junyeob, et al.
Published: (2025)
Multi-Object Tracking Consistently Improves Wildlife Inference
by: Muthivhi, Mufhumudzi, et al.
Published: (2026)
by: Muthivhi, Mufhumudzi, et al.
Published: (2026)
Zero-Shot Generalization of Vision-Based RL Without Data Augmentation
by: Batra, Sumeet, et al.
Published: (2024)
by: Batra, Sumeet, et al.
Published: (2024)
Automated Model Evaluation for Object Detection via Prediction Consistency and Reliability
by: Yoo, Seungju, et al.
Published: (2025)
by: Yoo, Seungju, et al.
Published: (2025)
A Foundation Model for General Moving Object Segmentation in Medical Images
by: Yan, Zhongnuo, et al.
Published: (2023)
by: Yan, Zhongnuo, et al.
Published: (2023)
Interacted Object Grounding in Spatio-Temporal Human-Object Interactions
by: Liu, Xiaoyang, et al.
Published: (2024)
by: Liu, Xiaoyang, et al.
Published: (2024)
Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
Cross-Domain Generalization Limits of Vision Foundation Models in Facial Deepfake Detection
by: Delibasoglu, Ibrahim
Published: (2026)
by: Delibasoglu, Ibrahim
Published: (2026)
Similar Items
-
Lost in Latent Space: Disentangled Models and the Challenge of Combinatorial Generalisation
by: Montero, Milton L., et al.
Published: (2022) -
SlotPi: Physics-informed Object-centric Reasoning Models
by: Li, Jian, et al.
Published: (2025) -
DORSal: Diffusion for Object-centric Representations of Scenes et al
by: Jabri, Allan, et al.
Published: (2023) -
Weakly Supervised Concept Learning for Object-centric Visual Reasoning
by: Tiwari, Sparsh, et al.
Published: (2026) -
Edge-case Synthesis for Fisheye Object Detection: A Data-centric Perspective
by: Kim, Seunghyeon, et al.
Published: (2025)