Slow-Fast Architecture for Video Multi-Modal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Min, Wang, Shihao, Chen, Chieh-Yun, Jain, Jitesh, Wang, Kai, Xiong, Junjun, Liu, Guilin, Yu, Zhiding, Shi, Humphrey |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
von: Jain, Jitesh, et al.
Veröffentlicht: (2024)
von: Jain, Jitesh, et al.
Veröffentlicht: (2024)
T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation
von: Chen, Chieh-Yun, et al.
Veröffentlicht: (2025)
von: Chen, Chieh-Yun, et al.
Veröffentlicht: (2025)
Benchmarking Object Detectors with COCO: A New Path Forward
von: Singh, Shweta, et al.
Veröffentlicht: (2024)
von: Singh, Shweta, et al.
Veröffentlicht: (2024)
PhyCritic: Multimodal Critic Models for Physical AI
von: Xiong, Tianyi, et al.
Veröffentlicht: (2026)
von: Xiong, Tianyi, et al.
Veröffentlicht: (2026)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
von: Jain, Jitesh, et al.
Veröffentlicht: (2025)
von: Jain, Jitesh, et al.
Veröffentlicht: (2025)
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
von: Wang, Shihao, et al.
Veröffentlicht: (2024)
von: Wang, Shihao, et al.
Veröffentlicht: (2024)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
von: Xu, Mingze, et al.
Veröffentlicht: (2024)
von: Xu, Mingze, et al.
Veröffentlicht: (2024)
Geo-LLaVA: A Large Multi-Modal Model for Solving Geometry Math Problems with Meta In-Context Learning
von: Xu, Shihao, et al.
Veröffentlicht: (2024)
von: Xu, Shihao, et al.
Veröffentlicht: (2024)
Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
von: Shi, Min, et al.
Veröffentlicht: (2024)
von: Shi, Min, et al.
Veröffentlicht: (2024)
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
von: Wang, Shihao, et al.
Veröffentlicht: (2026)
von: Wang, Shihao, et al.
Veröffentlicht: (2026)
Slot-VLM: SlowFast Slots for Video-Language Modeling
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
StreamChat: Chatting with Streaming Video
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
TrajSV: A Trajectory-based Model for Sports Video Representations and Applications
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
von: Xu, Mingze, et al.
Veröffentlicht: (2025)
von: Xu, Mingze, et al.
Veröffentlicht: (2025)
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
von: Hong, Yining, et al.
Veröffentlicht: (2024)
von: Hong, Yining, et al.
Veröffentlicht: (2024)
Versatile Diffusion: Text, Images and Variations All in One Diffusion Model
von: Xu, Xingqian, et al.
Veröffentlicht: (2022)
von: Xu, Xingqian, et al.
Veröffentlicht: (2022)
Multi-Modal Video Dialog State Tracking in the Wild
von: Abdessaied, Adnen, et al.
Veröffentlicht: (2024)
von: Abdessaied, Adnen, et al.
Veröffentlicht: (2024)
LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
von: Man, Yunze, et al.
Veröffentlicht: (2025)
von: Man, Yunze, et al.
Veröffentlicht: (2025)
Seeing Fast and Slow: Learning the Flow of Time in Videos
von: Wu, Yen-Siang, et al.
Veröffentlicht: (2026)
von: Wu, Yen-Siang, et al.
Veröffentlicht: (2026)
Jailbreak Large Vision-Language Models Through Multi-Modal Linkage
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
von: Man, Yunze, et al.
Veröffentlicht: (2025)
von: Man, Yunze, et al.
Veröffentlicht: (2025)
Visual-Oriented Fine-Grained Knowledge Editing for MultiModal Large Language Models
von: Zeng, Zhen, et al.
Veröffentlicht: (2024)
von: Zeng, Zhen, et al.
Veröffentlicht: (2024)
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
von: Li, Chenxuan, et al.
Veröffentlicht: (2024)
von: Li, Chenxuan, et al.
Veröffentlicht: (2024)
Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
von: Zhao, Han, et al.
Veröffentlicht: (2024)
von: Zhao, Han, et al.
Veröffentlicht: (2024)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
von: Peruzzo, Elia, et al.
Veröffentlicht: (2025)
von: Peruzzo, Elia, et al.
Veröffentlicht: (2025)
SlowFastVAD: Video Anomaly Detection via Integrating Simple Detector and RAG-Enhanced Vision-Language Model
von: Ding, Zongcan, et al.
Veröffentlicht: (2025)
von: Ding, Zongcan, et al.
Veröffentlicht: (2025)
Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression Inference
von: Shi, Zhiceng, et al.
Veröffentlicht: (2026)
von: Shi, Zhiceng, et al.
Veröffentlicht: (2026)
Generating, Fast and Slow: Scalable Parallel Video Generation with Video Interface Networks
von: Dedhia, Bhishma, et al.
Veröffentlicht: (2025)
von: Dedhia, Bhishma, et al.
Veröffentlicht: (2025)
From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
von: Yin, Tianwei, et al.
Veröffentlicht: (2024)
von: Yin, Tianwei, et al.
Veröffentlicht: (2024)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop Training
von: Li, Zhenxin, et al.
Veröffentlicht: (2025)
von: Li, Zhenxin, et al.
Veröffentlicht: (2025)
MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models
von: Chen, Chieh-Yun, et al.
Veröffentlicht: (2025)
von: Chen, Chieh-Yun, et al.
Veröffentlicht: (2025)
Fast-Slow Efficient Training for Multimodal Large Language Models via Visual Token Pruning
von: Zhang, Dingkun, et al.
Veröffentlicht: (2026)
von: Zhang, Dingkun, et al.
Veröffentlicht: (2026)
Hypergraph Multi-modal Large Language Model: Exploiting EEG and Eye-tracking Modalities to Evaluate Heterogeneous Responses for Video Understanding
von: Wu, Minghui, et al.
Veröffentlicht: (2024)
von: Wu, Minghui, et al.
Veröffentlicht: (2024)
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
von: Xu, Wenhao, et al.
Veröffentlicht: (2025)
von: Xu, Wenhao, et al.
Veröffentlicht: (2025)
GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentation
von: Lin, Lang, et al.
Veröffentlicht: (2025)
von: Lin, Lang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
von: Jain, Jitesh, et al.
Veröffentlicht: (2024) -
T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation
von: Chen, Chieh-Yun, et al.
Veröffentlicht: (2025) -
Benchmarking Object Detectors with COCO: A New Path Forward
von: Singh, Shweta, et al.
Veröffentlicht: (2024) -
PhyCritic: Multimodal Critic Models for Physical AI
von: Xiong, Tianyi, et al.
Veröffentlicht: (2026) -
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
von: Wang, Shihao, et al.
Veröffentlicht: (2025)