LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Sun, Boyuan, Zhao, Jiaxing, Chen, Xiang, Wei, Xihan, Hou, Qibin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
di: Sun, Boyuan, et al.
Pubblicazione: (2025)
di: Sun, Boyuan, et al.
Pubblicazione: (2025)
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness
di: Zhao, Jiaxing, et al.
Pubblicazione: (2025)
di: Zhao, Jiaxing, et al.
Pubblicazione: (2025)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
di: Zhang, Yuanhan, et al.
Pubblicazione: (2024)
di: Zhang, Yuanhan, et al.
Pubblicazione: (2024)
LLaVA-c: Continual Improved Visual Instruction Tuning
di: Liu, Wenzhuo, et al.
Pubblicazione: (2025)
di: Liu, Wenzhuo, et al.
Pubblicazione: (2025)
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
di: Yan, Dawei, et al.
Pubblicazione: (2024)
di: Yan, Dawei, et al.
Pubblicazione: (2024)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
di: Zhao, Xiangyu, et al.
Pubblicazione: (2024)
di: Zhao, Xiangyu, et al.
Pubblicazione: (2024)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
di: Sun, Boyuan, et al.
Pubblicazione: (2026)
di: Sun, Boyuan, et al.
Pubblicazione: (2026)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
di: Yuan, Haobo, et al.
Pubblicazione: (2025)
di: Yuan, Haobo, et al.
Pubblicazione: (2025)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
di: Gao, Mingze, et al.
Pubblicazione: (2024)
di: Gao, Mingze, et al.
Pubblicazione: (2024)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
di: Xu, Lin, et al.
Pubblicazione: (2024)
di: Xu, Lin, et al.
Pubblicazione: (2024)
LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning
di: Zhang, Dewen, et al.
Pubblicazione: (2025)
di: Zhang, Dewen, et al.
Pubblicazione: (2025)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
di: Cao, Meng, et al.
Pubblicazione: (2024)
di: Cao, Meng, et al.
Pubblicazione: (2024)
Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
di: Chaubey, Ashutosh, et al.
Pubblicazione: (2025)
di: Chaubey, Ashutosh, et al.
Pubblicazione: (2025)
DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-Reflection
di: Shen, Zhuokang, et al.
Pubblicazione: (2025)
di: Shen, Zhuokang, et al.
Pubblicazione: (2025)
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
di: Li, Hongyu, et al.
Pubblicazione: (2025)
di: Li, Hongyu, et al.
Pubblicazione: (2025)
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs
di: Shen, Leqi, et al.
Pubblicazione: (2025)
di: Shen, Leqi, et al.
Pubblicazione: (2025)
Dr-LLaVA: Visual Instruction Tuning with Symbolic Clinical Grounding
di: Sun, Shenghuan, et al.
Pubblicazione: (2024)
di: Sun, Shenghuan, et al.
Pubblicazione: (2024)
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
di: Ye, Xubing, et al.
Pubblicazione: (2024)
di: Ye, Xubing, et al.
Pubblicazione: (2024)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
di: Lin, Bin, et al.
Pubblicazione: (2023)
di: Lin, Bin, et al.
Pubblicazione: (2023)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
di: Lu, Weiheng, et al.
Pubblicazione: (2024)
di: Lu, Weiheng, et al.
Pubblicazione: (2024)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
di: Zeer, Ahmed, et al.
Pubblicazione: (2024)
di: Zeer, Ahmed, et al.
Pubblicazione: (2024)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
di: Zhang, Tao, et al.
Pubblicazione: (2024)
di: Zhang, Tao, et al.
Pubblicazione: (2024)
LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models
di: Zhang, Ruiyi, et al.
Pubblicazione: (2024)
di: Zhang, Ruiyi, et al.
Pubblicazione: (2024)
R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest
di: Chen, Xupeng, et al.
Pubblicazione: (2024)
di: Chen, Xupeng, et al.
Pubblicazione: (2024)
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
di: An, Xiang, et al.
Pubblicazione: (2026)
di: An, Xiang, et al.
Pubblicazione: (2026)
Enhance Image-to-Image Generation with LLaVA-generated Prompts
di: Ding, Zhicheng, et al.
Pubblicazione: (2024)
di: Ding, Zhicheng, et al.
Pubblicazione: (2024)
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
di: Huang, Runhui, et al.
Pubblicazione: (2024)
di: Huang, Runhui, et al.
Pubblicazione: (2024)
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
di: Cai, Yuxuan, et al.
Pubblicazione: (2024)
di: Cai, Yuxuan, et al.
Pubblicazione: (2024)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
di: Liang, Han, et al.
Pubblicazione: (2024)
di: Liang, Han, et al.
Pubblicazione: (2024)
Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA
di: Kanjula, Karthik Reddy, et al.
Pubblicazione: (2025)
di: Kanjula, Karthik Reddy, et al.
Pubblicazione: (2025)
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
di: Xu, Guowei, et al.
Pubblicazione: (2024)
di: Xu, Guowei, et al.
Pubblicazione: (2024)
LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs
di: Chen, Shaoxiang, et al.
Pubblicazione: (2024)
di: Chen, Shaoxiang, et al.
Pubblicazione: (2024)
LLaVA-Critic: Learning to Evaluate Multimodal Models
di: Xiong, Tianyi, et al.
Pubblicazione: (2024)
di: Xiong, Tianyi, et al.
Pubblicazione: (2024)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
di: Xu, Mingze, et al.
Pubblicazione: (2024)
di: Xu, Mingze, et al.
Pubblicazione: (2024)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
di: Zhao, Jiaxing, et al.
Pubblicazione: (2025)
di: Zhao, Jiaxing, et al.
Pubblicazione: (2025)
u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model
di: Xu, Jinjin, et al.
Pubblicazione: (2023)
di: Xu, Jinjin, et al.
Pubblicazione: (2023)
LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration
di: Inal, Gokce, et al.
Pubblicazione: (2026)
di: Inal, Gokce, et al.
Pubblicazione: (2026)
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
di: Xu, Mingze, et al.
Pubblicazione: (2025)
di: Xu, Mingze, et al.
Pubblicazione: (2025)
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
di: Zhou, Hanyu, et al.
Pubblicazione: (2025)
di: Zhou, Hanyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
di: Sun, Boyuan, et al.
Pubblicazione: (2025) -
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness
di: Zhao, Jiaxing, et al.
Pubblicazione: (2025) -
LLaVA-Video: Video Instruction Tuning With Synthetic Data
di: Zhang, Yuanhan, et al.
Pubblicazione: (2024) -
LLaVA-c: Continual Improved Visual Instruction Tuning
di: Liu, Wenzhuo, et al.
Pubblicazione: (2025) -
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
di: Yan, Dawei, et al.
Pubblicazione: (2024)