Divided Attention: Unsupervised Multi-Object Discovery with Contextually Separated Slots
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lao, Dong, Hu, Zhengyang, Locatello, Francesco, Yang, Yanchao, Soatto, Stefano |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Masked Multi-Query Slot Attention for Unsupervised Object Discovery
von: Pramanik, Rishav, et al.
Veröffentlicht: (2024)
von: Pramanik, Rishav, et al.
Veröffentlicht: (2024)
Adaptive Slot Attention: Object Discovery with Dynamic Slot Number
von: Fan, Ke, et al.
Veröffentlicht: (2024)
von: Fan, Ke, et al.
Veröffentlicht: (2024)
Sub-token ViT Embedding via Stochastic Resonance Transformers
von: Lao, Dong, et al.
Veröffentlicht: (2023)
von: Lao, Dong, et al.
Veröffentlicht: (2023)
SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
von: Grigore, Diana-Nicoleta, et al.
Veröffentlicht: (2025)
von: Grigore, Diana-Nicoleta, et al.
Veröffentlicht: (2025)
Neural Slot Interpreters: Grounding Object Semantics in Emergent Slot Representations
von: Dedhia, Bhishma, et al.
Veröffentlicht: (2024)
von: Dedhia, Bhishma, et al.
Veröffentlicht: (2024)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
von: Pani, Anupam, et al.
Veröffentlicht: (2025)
von: Pani, Anupam, et al.
Veröffentlicht: (2025)
Test-Time Defense Against Adversarial Attacks via Stochastic Resonance of Latent Ensembles
von: Lao, Dong, et al.
Veröffentlicht: (2025)
von: Lao, Dong, et al.
Veröffentlicht: (2025)
Anatomy-Slot: Unsupervised Anatomical Factorization for Homologous Bilateral Reasoning in Retinal Diagnosis
von: Ma, Yingzhe, et al.
Veröffentlicht: (2026)
von: Ma, Yingzhe, et al.
Veröffentlicht: (2026)
Non-autoregressive Sequence-to-Sequence Vision-Language Models
von: Shi, Kunyu, et al.
Veröffentlicht: (2024)
von: Shi, Kunyu, et al.
Veröffentlicht: (2024)
SFA-UNet: More Attention to Multi-Scale Contrast and Contextual Information in Infrared Small Object Segmentation
von: Shah, Imad Ali, et al.
Veröffentlicht: (2024)
von: Shah, Imad Ali, et al.
Veröffentlicht: (2024)
Grounded Object Centric Learning
von: Kori, Avinash, et al.
Veröffentlicht: (2023)
von: Kori, Avinash, et al.
Veröffentlicht: (2023)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
von: Chi, Donghwan, et al.
Veröffentlicht: (2025)
von: Chi, Donghwan, et al.
Veröffentlicht: (2025)
Guided Slot Attention for Unsupervised Video Object Segmentation
von: Lee, Minhyeok, et al.
Veröffentlicht: (2023)
von: Lee, Minhyeok, et al.
Veröffentlicht: (2023)
Multi-View Slot Attention Using Paraphrased Texts for Face Anti-Spoofing
von: Yu, Jeongmin, et al.
Veröffentlicht: (2025)
von: Yu, Jeongmin, et al.
Veröffentlicht: (2025)
CGSA: Class-Guided Slot-Aware Adaptation for Source-Free Object Detection
von: Dai, Boyang, et al.
Veröffentlicht: (2026)
von: Dai, Boyang, et al.
Veröffentlicht: (2026)
Object-Centric Learning with Slot Mixture Module
von: Kirilenko, Daniil, et al.
Veröffentlicht: (2023)
von: Kirilenko, Daniil, et al.
Veröffentlicht: (2023)
Future Slot Prediction for Unsupervised Object Discovery in Surgical Video
von: Liao, Guiqiu, et al.
Veröffentlicht: (2025)
von: Liao, Guiqiu, et al.
Veröffentlicht: (2025)
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
von: Kaul, Prannay, et al.
Veröffentlicht: (2024)
von: Kaul, Prannay, et al.
Veröffentlicht: (2024)
SlotLifter: Slot-guided Feature Lifting for Learning Object-centric Radiance Fields
von: Liu, Yu, et al.
Veröffentlicht: (2024)
von: Liu, Yu, et al.
Veröffentlicht: (2024)
Mutual Information guided Visual Contrastive Learning
von: Chen, Hanyang, et al.
Veröffentlicht: (2025)
von: Chen, Hanyang, et al.
Veröffentlicht: (2025)
NeRF-Insert: 3D Local Editing with Multimodal Control Signals
von: Sabat, Benet Oriol, et al.
Veröffentlicht: (2024)
von: Sabat, Benet Oriol, et al.
Veröffentlicht: (2024)
Near, far: Patch-ordering enhances vision foundation models' scene understanding
von: Pariza, Valentinos, et al.
Veröffentlicht: (2024)
von: Pariza, Valentinos, et al.
Veröffentlicht: (2024)
Leveraging Unsupervised Learning for Cost-Effective Visual Anomaly Detection
von: Long, Yunbo, et al.
Veröffentlicht: (2024)
von: Long, Yunbo, et al.
Veröffentlicht: (2024)
SlotPi: Physics-informed Object-centric Reasoning Models
von: Li, Jian, et al.
Veröffentlicht: (2025)
von: Li, Jian, et al.
Veröffentlicht: (2025)
Diffeomorphic Template Registration for Atmospheric Turbulence Mitigation
von: Lao, Dong, et al.
Veröffentlicht: (2024)
von: Lao, Dong, et al.
Veröffentlicht: (2024)
QASA: Quality-Guided K-Adaptive Slot Attention for Unsupervised Object-Centric Learning
von: Ouyang, Tianran, et al.
Veröffentlicht: (2026)
von: Ouyang, Tianran, et al.
Veröffentlicht: (2026)
Unsupervised Object Detection with Theoretical Guarantees
von: Longa, Marian, et al.
Veröffentlicht: (2024)
von: Longa, Marian, et al.
Veröffentlicht: (2024)
Binding Dynamics in Rotating Features
von: Löwe, Sindy, et al.
Veröffentlicht: (2024)
von: Löwe, Sindy, et al.
Veröffentlicht: (2024)
Temporally Consistent Object-Centric Learning by Contrasting Slots
von: Manasyan, Anna, et al.
Veröffentlicht: (2024)
von: Manasyan, Anna, et al.
Veröffentlicht: (2024)
Transformers and Slot Encoding for Sample Efficient Physical World Modelling
von: Petri, Francesco, et al.
Veröffentlicht: (2024)
von: Petri, Francesco, et al.
Veröffentlicht: (2024)
Contextual Object Detection with Multimodal Large Language Models
von: Zang, Yuhang, et al.
Veröffentlicht: (2023)
von: Zang, Yuhang, et al.
Veröffentlicht: (2023)
Divide and Conquer: Object Co-occurrence Helps Mitigate Simplicity Bias in OOD Detection
von: Dai, Boyang, et al.
Veröffentlicht: (2026)
von: Dai, Boyang, et al.
Veröffentlicht: (2026)
Slot-ID: Identity-Preserving Video Generation from Reference Videos via Slot-Based Temporal Identity Encoding
von: Lai, Yixuan, et al.
Veröffentlicht: (2026)
von: Lai, Yixuan, et al.
Veröffentlicht: (2026)
Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model
von: Li, Xiaolong, et al.
Veröffentlicht: (2024)
von: Li, Xiaolong, et al.
Veröffentlicht: (2024)
Toddlers' Active Gaze Behavior Supports Self-Supervised Object Learning
von: Yu, Zhengyang, et al.
Veröffentlicht: (2024)
von: Yu, Zhengyang, et al.
Veröffentlicht: (2024)
Towards Robust Unsupervised Attention Prediction in Autonomous Driving
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
unMORE: Unsupervised Multi-Object Segmentation via Center-Boundary Reasoning
von: Yang, Yafei, et al.
Veröffentlicht: (2025)
von: Yang, Yafei, et al.
Veröffentlicht: (2025)
WorDepth: Variational Language Prior for Monocular Depth Estimation
von: Zeng, Ziyao, et al.
Veröffentlicht: (2024)
von: Zeng, Ziyao, et al.
Veröffentlicht: (2024)
Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration
von: Zhu, Younan, et al.
Veröffentlicht: (2025)
von: Zhu, Younan, et al.
Veröffentlicht: (2025)
Dynamic Modality-Camera Invariant Clustering for Unsupervised Visible-Infrared Person Re-identification
von: Yang, Yiming, et al.
Veröffentlicht: (2024)
von: Yang, Yiming, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Masked Multi-Query Slot Attention for Unsupervised Object Discovery
von: Pramanik, Rishav, et al.
Veröffentlicht: (2024) -
Adaptive Slot Attention: Object Discovery with Dynamic Slot Number
von: Fan, Ke, et al.
Veröffentlicht: (2024) -
Sub-token ViT Embedding via Stochastic Resonance Transformers
von: Lao, Dong, et al.
Veröffentlicht: (2023) -
SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
von: Grigore, Diana-Nicoleta, et al.
Veröffentlicht: (2025) -
Neural Slot Interpreters: Grounding Object Semantics in Emergent Slot Representations
von: Dedhia, Bhishma, et al.
Veröffentlicht: (2024)