Unsupervised Open-Vocabulary Object Localization in Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Ke, Bai, Zechen, Xiao, Tianjun, Zietlow, Dominik, Horn, Max, Zhao, Zixu, Simon-Gabriel, Carl-Johann, Shou, Mike Zheng, Locatello, Francesco, Schiele, Bernt, Brox, Thomas, Zhang, Zheng, Fu, Yanwei, He, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Slot Attention: Object Discovery with Dynamic Slot Number
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations
by: Gairola, Siddhartha, et al.
Published: (2025)
by: Gairola, Siddhartha, et al.
Published: (2025)
Hallucination of Multimodal Large Language Models: A Survey
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
by: Bai, Zechen, et al.
Published: (2025)
by: Bai, Zechen, et al.
Published: (2025)
Impossible Videos
by: Bai, Zechen, et al.
Published: (2025)
by: Bai, Zechen, et al.
Published: (2025)
Factorized Visual Tokenization and Generation
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
by: Lin, Kevin Qinghong, et al.
Published: (2025)
by: Lin, Kevin Qinghong, et al.
Published: (2025)
World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy
by: Liu, Xiaokang, et al.
Published: (2026)
by: Liu, Xiaokang, et al.
Published: (2026)
LOVA3: Learning to Visual Question Answering, Asking and Assessment
by: Zhao, Henry Hengyuan, et al.
Published: (2024)
by: Zhao, Henry Hengyuan, et al.
Published: (2024)
Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance
by: Lin, Yiqi, et al.
Published: (2026)
by: Lin, Yiqi, et al.
Published: (2026)
Skip \n: A Simple Method to Reduce Hallucination in Large Vision-Language Models
by: Han, Zongbo, et al.
Published: (2024)
by: Han, Zongbo, et al.
Published: (2024)
Divided Attention: Unsupervised Multi-Object Discovery with Contextually Separated Slots
by: Lao, Dong, et al.
Published: (2023)
by: Lao, Dong, et al.
Published: (2023)
Boosting Unsupervised Segmentation Learning
by: Sari, Alp Eren, et al.
Published: (2024)
by: Sari, Alp Eren, et al.
Published: (2024)
One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
VideoSAM: Open-World Video Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
PersonaHOI: Effortlessly Improving Personalized Face with Human-Object Interaction Generation
by: Hu, Xinting, et al.
Published: (2025)
by: Hu, Xinting, et al.
Published: (2025)
Bring Your Own Character: A Holistic Solution for Automatic Facial Animation Generation of Customized Characters
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
Training-free Boost for Open-Vocabulary Object Detection with Confidence Aggregation
by: Zheng, Yanhao, et al.
Published: (2024)
by: Zheng, Yanhao, et al.
Published: (2024)
Binding Dynamics in Rotating Features
by: Löwe, Sindy, et al.
Published: (2024)
by: Löwe, Sindy, et al.
Published: (2024)
Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation
by: Cheng, Jiaxin, et al.
Published: (2024)
by: Cheng, Jiaxin, et al.
Published: (2024)
Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning
by: Pham, Nhi, et al.
Published: (2025)
by: Pham, Nhi, et al.
Published: (2025)
Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation
by: Chen, Yanzhe, et al.
Published: (2026)
by: Chen, Yanzhe, et al.
Published: (2026)
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
by: Ahmed, Noor, et al.
Published: (2024)
by: Ahmed, Noor, et al.
Published: (2024)
Convolutional Dynamic Alignment Networks for Interpretable Classifications
by: Böhle, Moritz, et al.
Published: (2021)
by: Böhle, Moritz, et al.
Published: (2021)
Better Understanding Differences in Attribution Methods via Systematic Evaluations
by: Rao, Sukrut, et al.
Published: (2023)
by: Rao, Sukrut, et al.
Published: (2023)
DWDN: Deep Wiener Deconvolution Network for Non-Blind Image Deblurring
by: Dong, Jiangxin, et al.
Published: (2021)
by: Dong, Jiangxin, et al.
Published: (2021)
Optimising for Interpretability: Convolutional Dynamic Alignment Networks
by: Böhle, Moritz, et al.
Published: (2021)
by: Böhle, Moritz, et al.
Published: (2021)
VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow
by: Gorgun, Ada, et al.
Published: (2025)
by: Gorgun, Ada, et al.
Published: (2025)
Towards Better Understanding Attribution Methods
by: Rao, Sukrut, et al.
Published: (2022)
by: Rao, Sukrut, et al.
Published: (2022)
Adversarial Training against Location-Optimized Adversarial Patches
by: Rao, Sukrut, et al.
Published: (2020)
by: Rao, Sukrut, et al.
Published: (2020)
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
by: Xie, Jiahao, et al.
Published: (2026)
by: Xie, Jiahao, et al.
Published: (2026)
Causal Learning with the Invariance Principle
by: Montagna, Francesco, et al.
Published: (2026)
by: Montagna, Francesco, et al.
Published: (2026)
WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
OpenObject-NAV: Open-Vocabulary Object-Oriented Navigation Based on Dynamic Carrier-Relationship Scene Graph
by: Tang, Yujie, et al.
Published: (2024)
by: Tang, Yujie, et al.
Published: (2024)
TPDiff: Temporal Pyramid Video Diffusion Model
by: Ran, Lingmin, et al.
Published: (2025)
by: Ran, Lingmin, et al.
Published: (2025)
P-Flow: Prompting Visual Effects Generation
by: Zhao, Rui, et al.
Published: (2026)
by: Zhao, Rui, et al.
Published: (2026)
D-AR: Diffusion via Autoregressive Models
by: Gao, Ziteng, et al.
Published: (2025)
by: Gao, Ziteng, et al.
Published: (2025)
Ego-centric Predictive Model Conditioned on Hand Trajectories
by: Zhang, Binjie, et al.
Published: (2025)
by: Zhang, Binjie, et al.
Published: (2025)
Demystifying amortized causal discovery with transformers
by: Montagna, Francesco, et al.
Published: (2024)
by: Montagna, Francesco, et al.
Published: (2024)
Similar Items
-
Adaptive Slot Attention: Object Discovery with Dynamic Slot Number
by: Fan, Ke, et al.
Published: (2024) -
Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach
by: Bai, Zechen, et al.
Published: (2024) -
How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations
by: Gairola, Siddhartha, et al.
Published: (2025) -
Hallucination of Multimodal Large Language Models: A Survey
by: Bai, Zechen, et al.
Published: (2024) -
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
by: Bai, Zechen, et al.
Published: (2025)