SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Mingze, Gao, Mingfei, Gan, Zhe, Chen, Hong-You, Lai, Zhengfeng, Gang, Haiming, Kang, Kai, Dehghan, Afshin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
di: Xu, Mingze, et al.
Pubblicazione: (2025)
di: Xu, Mingze, et al.
Pubblicazione: (2025)
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
di: Hong, Yining, et al.
Pubblicazione: (2024)
di: Hong, Yining, et al.
Pubblicazione: (2024)
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs
di: Shen, Leqi, et al.
Pubblicazione: (2025)
di: Shen, Leqi, et al.
Pubblicazione: (2025)
UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning
di: Tian, Rui, et al.
Pubblicazione: (2025)
di: Tian, Rui, et al.
Pubblicazione: (2025)
TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
di: Gao, Mingze, et al.
Pubblicazione: (2024)
di: Gao, Mingze, et al.
Pubblicazione: (2024)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
di: Zeer, Ahmed, et al.
Pubblicazione: (2024)
di: Zeer, Ahmed, et al.
Pubblicazione: (2024)
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
di: Yan, Dawei, et al.
Pubblicazione: (2024)
di: Yan, Dawei, et al.
Pubblicazione: (2024)
Slot-VLM: SlowFast Slots for Video-Language Modeling
di: Xu, Jiaqi, et al.
Pubblicazione: (2024)
di: Xu, Jiaqi, et al.
Pubblicazione: (2024)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
di: Zhang, Yuanhan, et al.
Pubblicazione: (2024)
di: Zhang, Yuanhan, et al.
Pubblicazione: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
di: Zhang, Shaolei, et al.
Pubblicazione: (2025)
di: Zhang, Shaolei, et al.
Pubblicazione: (2025)
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
di: Qu, Tingyu, et al.
Pubblicazione: (2024)
di: Qu, Tingyu, et al.
Pubblicazione: (2024)
Using SlowFast Networks for Near-Miss Incident Analysis in Dashcam Videos
di: Zhang, Yucheng, et al.
Pubblicazione: (2024)
di: Zhang, Yucheng, et al.
Pubblicazione: (2024)
SFMViT: SlowFast Meet ViT in Chaotic World
di: Lin, Jiaying, et al.
Pubblicazione: (2024)
di: Lin, Jiaying, et al.
Pubblicazione: (2024)
MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning
di: Zhang, Jianyi, et al.
Pubblicazione: (2024)
di: Zhang, Jianyi, et al.
Pubblicazione: (2024)
SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging
di: Zeng, Haijin, et al.
Pubblicazione: (2025)
di: Zeng, Haijin, et al.
Pubblicazione: (2025)
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
di: Wang, Xiyao, et al.
Pubblicazione: (2025)
di: Wang, Xiyao, et al.
Pubblicazione: (2025)
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
di: Ye, Xubing, et al.
Pubblicazione: (2024)
di: Ye, Xubing, et al.
Pubblicazione: (2024)
Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge
di: Yang, Yi, et al.
Pubblicazione: (2025)
di: Yang, Yi, et al.
Pubblicazione: (2025)
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
di: Wang, Haibo, et al.
Pubblicazione: (2025)
di: Wang, Haibo, et al.
Pubblicazione: (2025)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
di: Xu, Lin, et al.
Pubblicazione: (2024)
di: Xu, Lin, et al.
Pubblicazione: (2024)
Spatial Reasoning is Not a Free Lunch: A Controlled Study on LLaVA
di: Alam, Nahid, et al.
Pubblicazione: (2026)
di: Alam, Nahid, et al.
Pubblicazione: (2026)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
di: Lin, Bin, et al.
Pubblicazione: (2023)
di: Lin, Bin, et al.
Pubblicazione: (2023)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
di: Lu, Weiheng, et al.
Pubblicazione: (2024)
di: Lu, Weiheng, et al.
Pubblicazione: (2024)
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
di: Sun, Boyuan, et al.
Pubblicazione: (2025)
di: Sun, Boyuan, et al.
Pubblicazione: (2025)
LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information
di: Wang, Ke, et al.
Pubblicazione: (2024)
di: Wang, Ke, et al.
Pubblicazione: (2024)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
di: Zhao, Xiangyu, et al.
Pubblicazione: (2024)
di: Zhao, Xiangyu, et al.
Pubblicazione: (2024)
LLaVA-Critic: Learning to Evaluate Multimodal Models
di: Xiong, Tianyi, et al.
Pubblicazione: (2024)
di: Xiong, Tianyi, et al.
Pubblicazione: (2024)
LLaVAC: Fine-tuning LLaVA as a Multimodal Sentiment Classifier
di: Chay-intr, T., et al.
Pubblicazione: (2025)
di: Chay-intr, T., et al.
Pubblicazione: (2025)
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages
di: Andersland, Michael
Pubblicazione: (2024)
di: Andersland, Michael
Pubblicazione: (2024)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
di: Jin, Yizhang, et al.
Pubblicazione: (2024)
di: Jin, Yizhang, et al.
Pubblicazione: (2024)
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
di: Cai, Yuxuan, et al.
Pubblicazione: (2024)
di: Cai, Yuxuan, et al.
Pubblicazione: (2024)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
di: An, Ruichuan, et al.
Pubblicazione: (2025)
di: An, Ruichuan, et al.
Pubblicazione: (2025)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
di: Yuan, Haobo, et al.
Pubblicazione: (2025)
di: Yuan, Haobo, et al.
Pubblicazione: (2025)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
di: An, Ruichuan, et al.
Pubblicazione: (2024)
di: An, Ruichuan, et al.
Pubblicazione: (2024)
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
di: Tian, Rui, et al.
Pubblicazione: (2025)
di: Tian, Rui, et al.
Pubblicazione: (2025)
R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest
di: Chen, Xupeng, et al.
Pubblicazione: (2024)
di: Chen, Xupeng, et al.
Pubblicazione: (2024)
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
di: Sun, Boyuan, et al.
Pubblicazione: (2025)
di: Sun, Boyuan, et al.
Pubblicazione: (2025)
Can Sound Replace Vision in LLaVA With Token Substitution?
di: Vosoughi, Ali, et al.
Pubblicazione: (2025)
di: Vosoughi, Ali, et al.
Pubblicazione: (2025)
Enhance Image-to-Image Generation with LLaVA-generated Prompts
di: Ding, Zhicheng, et al.
Pubblicazione: (2024)
di: Ding, Zhicheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
di: Xu, Mingze, et al.
Pubblicazione: (2025) -
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
di: Hong, Yining, et al.
Pubblicazione: (2024) -
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs
di: Shen, Leqi, et al.
Pubblicazione: (2025) -
UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning
di: Tian, Rui, et al.
Pubblicazione: (2025) -
TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
di: Gao, Mingze, et al.
Pubblicazione: (2024)