PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Lin, Zhao, Yilin, Zhou, Daquan, Lin, Zhijie, Ng, See Kiong, Feng, Jiashi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
von: Xiong, Tianwei, et al.
Veröffentlicht: (2024)
von: Xiong, Tianwei, et al.
Veröffentlicht: (2024)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
von: Lin, Bin, et al.
Veröffentlicht: (2023)
von: Lin, Bin, et al.
Veröffentlicht: (2023)
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
von: Yan, Dawei, et al.
Veröffentlicht: (2024)
von: Yan, Dawei, et al.
Veröffentlicht: (2024)
TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
von: Gao, Mingze, et al.
Veröffentlicht: (2024)
von: Gao, Mingze, et al.
Veröffentlicht: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
von: Shu, Fangxun, et al.
Veröffentlicht: (2024)
von: Shu, Fangxun, et al.
Veröffentlicht: (2024)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
von: Zeer, Ahmed, et al.
Veröffentlicht: (2024)
von: Zeer, Ahmed, et al.
Veröffentlicht: (2024)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
von: Lu, Weiheng, et al.
Veröffentlicht: (2024)
von: Lu, Weiheng, et al.
Veröffentlicht: (2024)
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs
von: Shen, Leqi, et al.
Veröffentlicht: (2025)
von: Shen, Leqi, et al.
Veröffentlicht: (2025)
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
Enhance Image-to-Image Generation with LLaVA-generated Prompts
von: Ding, Zhicheng, et al.
Veröffentlicht: (2024)
von: Ding, Zhicheng, et al.
Veröffentlicht: (2024)
LLaVA-c: Continual Improved Visual Instruction Tuning
von: Liu, Wenzhuo, et al.
Veröffentlicht: (2025)
von: Liu, Wenzhuo, et al.
Veröffentlicht: (2025)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
von: Shi, Wenhao, et al.
Veröffentlicht: (2024)
von: Shi, Wenhao, et al.
Veröffentlicht: (2024)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
von: Xu, Mingze, et al.
Veröffentlicht: (2024)
von: Xu, Mingze, et al.
Veröffentlicht: (2024)
LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning
von: Li, Jiajie, et al.
Veröffentlicht: (2024)
von: Li, Jiajie, et al.
Veröffentlicht: (2024)
LLaVA-Critic: Learning to Evaluate Multimodal Models
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
von: Liang, Han, et al.
Veröffentlicht: (2024)
von: Liang, Han, et al.
Veröffentlicht: (2024)
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
LLaVA-OneVision: Easy Visual Task Transfer
von: Li, Bo, et al.
Veröffentlicht: (2024)
von: Li, Bo, et al.
Veröffentlicht: (2024)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
von: Li, Feng, et al.
Veröffentlicht: (2024)
von: Li, Feng, et al.
Veröffentlicht: (2024)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
von: Lin, Bin, et al.
Veröffentlicht: (2024)
von: Lin, Bin, et al.
Veröffentlicht: (2024)
LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2024)
Why do LLaVA Vision-Language Models Reply to Images in English?
von: Hinck, Musashi, et al.
Veröffentlicht: (2024)
von: Hinck, Musashi, et al.
Veröffentlicht: (2024)
ViDRiP-LLaVA: A Dataset and Benchmark for Diagnostic Reasoning from Pathology Videos
von: Vuong, Trinh T. L., et al.
Veröffentlicht: (2025)
von: Vuong, Trinh T. L., et al.
Veröffentlicht: (2025)
DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-Reflection
von: Shen, Zhuokang, et al.
Veröffentlicht: (2025)
von: Shen, Zhuokang, et al.
Veröffentlicht: (2025)
INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
von: Xu, Ruyi, et al.
Veröffentlicht: (2024)
von: Xu, Ruyi, et al.
Veröffentlicht: (2024)
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
von: An, Xiang, et al.
Veröffentlicht: (2026)
von: An, Xiang, et al.
Veröffentlicht: (2026)
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
von: Xu, Mingze, et al.
Veröffentlicht: (2025)
von: Xu, Mingze, et al.
Veröffentlicht: (2025)
LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information
von: Wang, Ke, et al.
Veröffentlicht: (2024)
von: Wang, Ke, et al.
Veröffentlicht: (2024)
Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
von: Xu, Guowei, et al.
Veröffentlicht: (2024)
von: Xu, Guowei, et al.
Veröffentlicht: (2024)
LLaVA-ReID: Selective Multi-image Questioner for Interactive Person Re-Identification
von: Lu, Yiding, et al.
Veröffentlicht: (2025)
von: Lu, Yiding, et al.
Veröffentlicht: (2025)
LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration
von: Inal, Gokce, et al.
Veröffentlicht: (2026)
von: Inal, Gokce, et al.
Veröffentlicht: (2026)
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
von: Zhou, Yupeng, et al.
Veröffentlicht: (2024)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
von: Yuan, Haobo, et al.
Veröffentlicht: (2025) -
LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
von: Xiong, Tianwei, et al.
Veröffentlicht: (2024) -
LLaVA-Video: Video Instruction Tuning With Synthetic Data
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024) -
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
von: Lin, Bin, et al.
Veröffentlicht: (2023) -
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
von: Yan, Dawei, et al.
Veröffentlicht: (2024)