Slot-VLM: SlowFast Slots for Video-Language Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Jiaqi, Lan, Cuiling, Xie, Wenxuan, Chen, Xuejin, Lu, Yan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Long Video Understanding with Learnable Retrieval in Video-Language Models
by: Xu, Jiaqi, et al.
Published: (2023)
by: Xu, Jiaqi, et al.
Published: (2023)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
by: Xu, Mingze, et al.
Published: (2024)
by: Xu, Mingze, et al.
Published: (2024)
SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging
by: Zeng, Haijin, et al.
Published: (2025)
by: Zeng, Haijin, et al.
Published: (2025)
Deciphering Functions of Neurons in Vision-Language Models
by: Xu, Jiaqi, et al.
Published: (2025)
by: Xu, Jiaqi, et al.
Published: (2025)
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
by: Xu, Mingze, et al.
Published: (2025)
by: Xu, Mingze, et al.
Published: (2025)
Predicting Video Slot Attention Queries from Random Slot-Feature Pairs
by: Zhao, Rongzhen, et al.
Published: (2025)
by: Zhao, Rongzhen, et al.
Published: (2025)
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
by: Wei, Meng, et al.
Published: (2025)
by: Wei, Meng, et al.
Published: (2025)
PhysMamba: Efficient Remote Physiological Measurement with SlowFast Temporal Difference Mamba
by: Luo, Chaoqi, et al.
Published: (2024)
by: Luo, Chaoqi, et al.
Published: (2024)
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
by: Hong, Yining, et al.
Published: (2024)
by: Hong, Yining, et al.
Published: (2024)
SFMViT: SlowFast Meet ViT in Chaotic World
by: Lin, Jiaying, et al.
Published: (2024)
by: Lin, Jiaying, et al.
Published: (2024)
SF-TMN: SlowFast Temporal Modeling Network for Surgical Phase Recognition
by: Zhang, Bokai, et al.
Published: (2023)
by: Zhang, Bokai, et al.
Published: (2023)
Slot-VAE: Object-Centric Scene Generation with Slot Attention
by: Wang, Yanbo, et al.
Published: (2023)
by: Wang, Yanbo, et al.
Published: (2023)
When Slots Compete: Slot Merging in Object-Centric Learning
by: Chatzisavvas, Christos, et al.
Published: (2026)
by: Chatzisavvas, Christos, et al.
Published: (2026)
Guided Slot Attention for Unsupervised Video Object Segmentation
by: Lee, Minhyeok, et al.
Published: (2023)
by: Lee, Minhyeok, et al.
Published: (2023)
Adaptive Slot Attention: Object Discovery with Dynamic Slot Number
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
MetaSlot: Break Through the Fixed Number of Slots in Object-Centric Learning
by: Liu, Hongjia, et al.
Published: (2025)
by: Liu, Hongjia, et al.
Published: (2025)
Slot-ID: Identity-Preserving Video Generation from Reference Videos via Slot-Based Temporal Identity Encoding
by: Lai, Yixuan, et al.
Published: (2026)
by: Lai, Yixuan, et al.
Published: (2026)
Slot Structured World Models
by: Collu, Jonathan, et al.
Published: (2024)
by: Collu, Jonathan, et al.
Published: (2024)
Future Slot Prediction for Unsupervised Object Discovery in Surgical Video
by: Liao, Guiqiu, et al.
Published: (2025)
by: Liao, Guiqiu, et al.
Published: (2025)
Neural Slot Interpreters: Grounding Object Semantics in Emergent Slot Representations
by: Dedhia, Bhishma, et al.
Published: (2024)
by: Dedhia, Bhishma, et al.
Published: (2024)
WhisperNetV2: SlowFast Siamese Network For Lip-Based Biometrics
by: Zakeri, Abdollah, et al.
Published: (2024)
by: Zakeri, Abdollah, et al.
Published: (2024)
OpenSlot: Mixed Open-Set Recognition with Object-Centric Learning
by: Yin, Xu, et al.
Published: (2024)
by: Yin, Xu, et al.
Published: (2024)
Smoothing Slot Attention Iterations and Recurrences
by: Zhao, Rongzhen, et al.
Published: (2025)
by: Zhao, Rongzhen, et al.
Published: (2025)
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
by: Han, Jiwook, et al.
Published: (2026)
by: Han, Jiwook, et al.
Published: (2026)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
Slot-guided Volumetric Object Radiance Fields
by: Qi, Di, et al.
Published: (2024)
by: Qi, Di, et al.
Published: (2024)
Slot Attention with Re-Initialization and Self-Distillation
by: Zhao, Rongzhen, et al.
Published: (2025)
by: Zhao, Rongzhen, et al.
Published: (2025)
Text Grouping Adapter: Adapting Pre-trained Text Detector for Layout Analysis
by: Bi, Tianci, et al.
Published: (2024)
by: Bi, Tianci, et al.
Published: (2024)
SlotLifter: Slot-guided Feature Lifting for Learning Object-centric Radiance Fields
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Attention Normalization Impacts Cardinality Generalization in Slot Attention
by: Krimmel, Markus, et al.
Published: (2024)
by: Krimmel, Markus, et al.
Published: (2024)
MUFASA: A Multi-Layer Framework for Slot Attention
by: Bock, Sebastian, et al.
Published: (2026)
by: Bock, Sebastian, et al.
Published: (2026)
Slot-BERT: Self-supervised Object Discovery in Surgical Video
by: Liao, Guiqiu, et al.
Published: (2025)
by: Liao, Guiqiu, et al.
Published: (2025)
SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
Learning Global Object-Centric Representations via Disentangled Slot Attention
by: Chen, Tonglin, et al.
Published: (2024)
by: Chen, Tonglin, et al.
Published: (2024)
Slot-Level Robotic Placement via Visual Imitation from Single Human Video
by: Shan, Dandan, et al.
Published: (2025)
by: Shan, Dandan, et al.
Published: (2025)
Slow-Fast Architecture for Video Multi-Modal Large Language Models
by: Shi, Min, et al.
Published: (2025)
by: Shi, Min, et al.
Published: (2025)
PRISM: Progressive Reasoning through Iterative Slot Memory for Vision
by: Wang, Ziyu, et al.
Published: (2026)
by: Wang, Ziyu, et al.
Published: (2026)
Slot Attention-based Feature Filtering for Few-Shot Learning
by: Rodenas, Javier, et al.
Published: (2025)
by: Rodenas, Javier, et al.
Published: (2025)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
by: Yang, Jianxuan, et al.
Published: (2025)
by: Yang, Jianxuan, et al.
Published: (2025)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
by: Hanyu, Taisei, et al.
Published: (2025)
by: Hanyu, Taisei, et al.
Published: (2025)
Similar Items
-
Long Video Understanding with Learnable Retrieval in Video-Language Models
by: Xu, Jiaqi, et al.
Published: (2023) -
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
by: Xu, Mingze, et al.
Published: (2024) -
SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging
by: Zeng, Haijin, et al.
Published: (2025) -
Deciphering Functions of Neurons in Vision-Language Models
by: Xu, Jiaqi, et al.
Published: (2025) -
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
by: Xu, Mingze, et al.
Published: (2025)