TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Xingjian, Weng, Xi, Yue, Yihao, Fan, Zhaoxin, Wu, Wenjun, Huang, Lei |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
par: Zhang, Xingjian, et autres
Publié: (2025)
par: Zhang, Xingjian, et autres
Publié: (2025)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
par: Shu, Fangxun, et autres
Publié: (2024)
par: Shu, Fangxun, et autres
Publié: (2024)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
par: Zhang, Zicheng, et autres
Publié: (2024)
par: Zhang, Zicheng, et autres
Publié: (2024)
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
par: Wen, Siwei, et autres
Publié: (2026)
par: Wen, Siwei, et autres
Publié: (2026)
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
par: Deng, Jiajun, et autres
Publié: (2025)
par: Deng, Jiajun, et autres
Publié: (2025)
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
par: Zhou, Hanyu, et autres
Publié: (2025)
par: Zhou, Hanyu, et autres
Publié: (2025)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
par: Yuan, Haobo, et autres
Publié: (2025)
par: Yuan, Haobo, et autres
Publié: (2025)
TinyLLaVA: A Framework of Small-scale Large Multimodal Models
par: Zhou, Baichuan, et autres
Publié: (2024)
par: Zhou, Baichuan, et autres
Publié: (2024)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
par: Zhang, Yuanhan, et autres
Publié: (2024)
par: Zhang, Yuanhan, et autres
Publié: (2024)
TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
par: Gao, Mingze, et autres
Publié: (2024)
par: Gao, Mingze, et autres
Publié: (2024)
TinyLLaVA Factory: A Modularized Codebase for Small-scale Large Multimodal Models
par: Jia, Junlong, et autres
Publié: (2024)
par: Jia, Junlong, et autres
Publié: (2024)
Text-Conditioned Resampler For Long Form Video Understanding
par: Korbar, Bruno, et autres
Publié: (2023)
par: Korbar, Bruno, et autres
Publié: (2023)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
par: Xu, Lin, et autres
Publié: (2024)
par: Xu, Lin, et autres
Publié: (2024)
LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMs
par: Zhou, Hanyu, et autres
Publié: (2025)
par: Zhou, Hanyu, et autres
Publié: (2025)
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
par: Zhu, Chenming, et autres
Publié: (2024)
par: Zhu, Chenming, et autres
Publié: (2024)
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs
par: Shen, Leqi, et autres
Publié: (2025)
par: Shen, Leqi, et autres
Publié: (2025)
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
par: Sun, Boyuan, et autres
Publié: (2025)
par: Sun, Boyuan, et autres
Publié: (2025)
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
par: Zhang, Boqiang, et autres
Publié: (2025)
par: Zhang, Boqiang, et autres
Publié: (2025)
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
par: Xu, Mingze, et autres
Publié: (2025)
par: Xu, Mingze, et autres
Publié: (2025)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
par: Lin, Bin, et autres
Publié: (2023)
par: Lin, Bin, et autres
Publié: (2023)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
par: Lu, Weiheng, et autres
Publié: (2024)
par: Lu, Weiheng, et autres
Publié: (2024)
LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning
par: Li, Jiajie, et autres
Publié: (2024)
par: Li, Jiajie, et autres
Publié: (2024)
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs
par: Yamao, Sosuke, et autres
Publié: (2024)
par: Yamao, Sosuke, et autres
Publié: (2024)
TennisExpert: Towards Expert-Level Analytical Sports Video Understanding
par: Liu, Zhaoyu, et autres
Publié: (2026)
par: Liu, Zhaoyu, et autres
Publié: (2026)
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
par: Khattak, Muhammad Uzair, et autres
Publié: (2024)
par: Khattak, Muhammad Uzair, et autres
Publié: (2024)
CityLLaVA: Efficient Fine-Tuning for VLMs in City Scenario
par: Duan, Zhizhao, et autres
Publié: (2024)
par: Duan, Zhizhao, et autres
Publié: (2024)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
par: Bharadwaj, Rohit, et autres
Publié: (2024)
par: Bharadwaj, Rohit, et autres
Publié: (2024)
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
par: Fan, Yue, et autres
Publié: (2024)
par: Fan, Yue, et autres
Publié: (2024)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
par: Wang, Yuxuan, et autres
Publié: (2024)
par: Wang, Yuxuan, et autres
Publié: (2024)
ViLLa: Video Reasoning Segmentation with Large Language Model
par: Zheng, Rongkun, et autres
Publié: (2024)
par: Zheng, Rongkun, et autres
Publié: (2024)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
par: Zhao, Xiangyu, et autres
Publié: (2024)
par: Zhao, Xiangyu, et autres
Publié: (2024)
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
par: Guo, Yuwei, et autres
Publié: (2025)
par: Guo, Yuwei, et autres
Publié: (2025)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
par: Fan, Yue, et autres
Publié: (2024)
par: Fan, Yue, et autres
Publié: (2024)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
par: Xu, Mingze, et autres
Publié: (2024)
par: Xu, Mingze, et autres
Publié: (2024)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
par: Cheng, Zesen, et autres
Publié: (2024)
par: Cheng, Zesen, et autres
Publié: (2024)
TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs
par: Wang, Juntong, et autres
Publié: (2025)
par: Wang, Juntong, et autres
Publié: (2025)
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
par: Lou, Haoran, et autres
Publié: (2025)
par: Lou, Haoran, et autres
Publié: (2025)
Towards Universal Soccer Video Understanding
par: Rao, Jiayuan, et autres
Publié: (2024)
par: Rao, Jiayuan, et autres
Publié: (2024)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
par: Zhang, Tao, et autres
Publié: (2024)
par: Zhang, Tao, et autres
Publié: (2024)
AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs
par: Xia, Shuhan, et autres
Publié: (2025)
par: Xia, Shuhan, et autres
Publié: (2025)
Documents similaires
-
TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
par: Zhang, Xingjian, et autres
Publié: (2025) -
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
par: Shu, Fangxun, et autres
Publié: (2024) -
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
par: Zhang, Zicheng, et autres
Publié: (2024) -
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
par: Wen, Siwei, et autres
Publié: (2026) -
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
par: Deng, Jiajun, et autres
Publié: (2025)