SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Jiwook, Ahn, Geo, Kim, Youngrae, Choi, Jinwoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
DEVIAS: Learning Disentangled Video Representations of Action and Scene
von: Bae, Kyungho, et al.
Veröffentlicht: (2023)
von: Bae, Kyungho, et al.
Veröffentlicht: (2023)
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
von: An, Joungbin, et al.
Veröffentlicht: (2026)
von: An, Joungbin, et al.
Veröffentlicht: (2026)
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
von: Lee, Jongseo, et al.
Veröffentlicht: (2024)
von: Lee, Jongseo, et al.
Veröffentlicht: (2024)
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT
von: Xu, Yifang, et al.
Veröffentlicht: (2024)
von: Xu, Yifang, et al.
Veröffentlicht: (2024)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
von: Chi, Donghwan, et al.
Veröffentlicht: (2025)
von: Chi, Donghwan, et al.
Veröffentlicht: (2025)
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2026)
von: Zheng, Minghang, et al.
Veröffentlicht: (2026)
When Slots Compete: Slot Merging in Object-Centric Learning
von: Chatzisavvas, Christos, et al.
Veröffentlicht: (2026)
von: Chatzisavvas, Christos, et al.
Veröffentlicht: (2026)
Slot-VAE: Object-Centric Scene Generation with Slot Attention
von: Wang, Yanbo, et al.
Veröffentlicht: (2023)
von: Wang, Yanbo, et al.
Veröffentlicht: (2023)
FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding
von: Cao, Zhuo, et al.
Veröffentlicht: (2024)
von: Cao, Zhuo, et al.
Veröffentlicht: (2024)
Leveraging Image Augmentation for Object Manipulation: Towards Interpretable Controllability in Object-Centric Learning
von: Kim, Jinwoo, et al.
Veröffentlicht: (2023)
von: Kim, Jinwoo, et al.
Veröffentlicht: (2023)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
von: Dou, Weijia, et al.
Veröffentlicht: (2026)
von: Dou, Weijia, et al.
Veröffentlicht: (2026)
Temporally Consistent Object-Centric Learning by Contrasting Slots
von: Manasyan, Anna, et al.
Veröffentlicht: (2024)
von: Manasyan, Anna, et al.
Veröffentlicht: (2024)
SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
von: Grigore, Diana-Nicoleta, et al.
Veröffentlicht: (2025)
von: Grigore, Diana-Nicoleta, et al.
Veröffentlicht: (2025)
Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
von: Moon, WonJun, et al.
Veröffentlicht: (2026)
von: Moon, WonJun, et al.
Veröffentlicht: (2026)
MetaSlot: Break Through the Fixed Number of Slots in Object-Centric Learning
von: Liu, Hongjia, et al.
Veröffentlicht: (2025)
von: Liu, Hongjia, et al.
Veröffentlicht: (2025)
Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
von: Lee, Seunghun, et al.
Veröffentlicht: (2025)
von: Lee, Seunghun, et al.
Veröffentlicht: (2025)
Learning Global Object-Centric Representations via Disentangled Slot Attention
von: Chen, Tonglin, et al.
Veröffentlicht: (2024)
von: Chen, Tonglin, et al.
Veröffentlicht: (2024)
Learning Object-Centric Representations Based on Slots in Real World Scenarios
von: Akan, Adil Kaan
Veröffentlicht: (2025)
von: Akan, Adil Kaan
Veröffentlicht: (2025)
OpenSlot: Mixed Open-Set Recognition with Object-Centric Learning
von: Yin, Xu, et al.
Veröffentlicht: (2024)
von: Yin, Xu, et al.
Veröffentlicht: (2024)
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
von: Villar-Corrales, Angel, et al.
Veröffentlicht: (2025)
von: Villar-Corrales, Angel, et al.
Veröffentlicht: (2025)
Compositional Video Synthesis by Temporal Object-Centric Learning
von: Akan, Adil Kaan, et al.
Veröffentlicht: (2025)
von: Akan, Adil Kaan, et al.
Veröffentlicht: (2025)
Neural Slot Interpreters: Grounding Object Semantics in Emergent Slot Representations
von: Dedhia, Bhishma, et al.
Veröffentlicht: (2024)
von: Dedhia, Bhishma, et al.
Veröffentlicht: (2024)
TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding
von: Lee, Jin-Seop, et al.
Veröffentlicht: (2025)
von: Lee, Jin-Seop, et al.
Veröffentlicht: (2025)
Object-Centric Learning with Slot Mixture Module
von: Kirilenko, Daniil, et al.
Veröffentlicht: (2023)
von: Kirilenko, Daniil, et al.
Veröffentlicht: (2023)
GLASS: Guided Latent Slot Diffusion for Object-Centric Learning
von: Singh, Krishnakant, et al.
Veröffentlicht: (2024)
von: Singh, Krishnakant, et al.
Veröffentlicht: (2024)
Guided Slot Attention for Unsupervised Video Object Segmentation
von: Lee, Minhyeok, et al.
Veröffentlicht: (2023)
von: Lee, Minhyeok, et al.
Veröffentlicht: (2023)
QASA: Quality-Guided K-Adaptive Slot Attention for Unsupervised Object-Centric Learning
von: Ouyang, Tianran, et al.
Veröffentlicht: (2026)
von: Ouyang, Tianran, et al.
Veröffentlicht: (2026)
Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2024)
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2024)
Class-Continuous Conditional Generative Neural Radiance Field
von: Kim, Jiwook, et al.
Veröffentlicht: (2023)
von: Kim, Jiwook, et al.
Veröffentlicht: (2023)
ORIDa: Object-centric Real-world Image Composition Dataset
von: Kim, Jinwoo, et al.
Veröffentlicht: (2025)
von: Kim, Jinwoo, et al.
Veröffentlicht: (2025)
OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
von: Song, Yeon-Ji, et al.
Veröffentlicht: (2024)
von: Song, Yeon-Ji, et al.
Veröffentlicht: (2024)
Future Slot Prediction for Unsupervised Object Discovery in Surgical Video
von: Liao, Guiqiu, et al.
Veröffentlicht: (2025)
von: Liao, Guiqiu, et al.
Veröffentlicht: (2025)
ContextFusion and Bootstrap: An Effective Approach to Improve Slot Attention-Based Object-Centric Learning
von: Tian, Pinzhuo, et al.
Veröffentlicht: (2025)
von: Tian, Pinzhuo, et al.
Veröffentlicht: (2025)
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
von: Bae, Kyungho, et al.
Veröffentlicht: (2025)
von: Bae, Kyungho, et al.
Veröffentlicht: (2025)
Infusing Environmental Captions for Long-Form Video Language Grounding
von: Lee, Hyogun, et al.
Veröffentlicht: (2024)
von: Lee, Hyogun, et al.
Veröffentlicht: (2024)
Scribble-Guided Diffusion for Training-free Text-to-Image Generation
von: Lee, Seonho, et al.
Veröffentlicht: (2024)
von: Lee, Seonho, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
von: Ahn, Geo, et al.
Veröffentlicht: (2026) -
DEVIAS: Learning Disentangled Video Representations of Action and Scene
von: Bae, Kyungho, et al.
Veröffentlicht: (2023) -
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
von: Guo, Yongxin, et al.
Veröffentlicht: (2024) -
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
von: An, Joungbin, et al.
Veröffentlicht: (2026) -
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
von: Lee, Jongseo, et al.
Veröffentlicht: (2024)