TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Xingjian, Wen, Siwei, Wu, Wenjun, Huang, Lei |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
par: Zhang, Xingjian, et autres
Publié: (2025)
par: Zhang, Xingjian, et autres
Publié: (2025)
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
par: Wen, Siwei, et autres
Publié: (2026)
par: Wen, Siwei, et autres
Publié: (2026)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
par: Shu, Fangxun, et autres
Publié: (2024)
par: Shu, Fangxun, et autres
Publié: (2024)
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
par: Deng, Jiajun, et autres
Publié: (2025)
par: Deng, Jiajun, et autres
Publié: (2025)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
par: Zhang, Yuanhan, et autres
Publié: (2024)
par: Zhang, Yuanhan, et autres
Publié: (2024)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
par: Zhang, Zicheng, et autres
Publié: (2024)
par: Zhang, Zicheng, et autres
Publié: (2024)
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
par: Zhu, Chenming, et autres
Publié: (2024)
par: Zhu, Chenming, et autres
Publié: (2024)
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
par: Khattak, Muhammad Uzair, et autres
Publié: (2024)
par: Khattak, Muhammad Uzair, et autres
Publié: (2024)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
par: Xu, Lin, et autres
Publié: (2024)
par: Xu, Lin, et autres
Publié: (2024)
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
par: Zhou, Hanyu, et autres
Publié: (2025)
par: Zhou, Hanyu, et autres
Publié: (2025)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
par: Yuan, Haobo, et autres
Publié: (2025)
par: Yuan, Haobo, et autres
Publié: (2025)
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs
par: Shen, Leqi, et autres
Publié: (2025)
par: Shen, Leqi, et autres
Publié: (2025)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
par: Lin, Bin, et autres
Publié: (2023)
par: Lin, Bin, et autres
Publié: (2023)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
par: Lu, Weiheng, et autres
Publié: (2024)
par: Lu, Weiheng, et autres
Publié: (2024)
Video-R1: Reinforcing Video Reasoning in MLLMs
par: Feng, Kaituo, et autres
Publié: (2025)
par: Feng, Kaituo, et autres
Publié: (2025)
ViLLa: Video Reasoning Segmentation with Large Language Model
par: Zheng, Rongkun, et autres
Publié: (2024)
par: Zheng, Rongkun, et autres
Publié: (2024)
LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning
par: Li, Jiajie, et autres
Publié: (2024)
par: Li, Jiajie, et autres
Publié: (2024)
Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought
par: Huang, Chao, et autres
Publié: (2025)
par: Huang, Chao, et autres
Publié: (2025)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
par: Bharadwaj, Rohit, et autres
Publié: (2024)
par: Bharadwaj, Rohit, et autres
Publié: (2024)
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
par: Zhang, Jianrui, et autres
Publié: (2024)
par: Zhang, Jianrui, et autres
Publié: (2024)
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
par: Xu, Mingze, et autres
Publié: (2025)
par: Xu, Mingze, et autres
Publié: (2025)
MMSearch-R1: Incentivizing LMMs to Search
par: Wu, Jinming, et autres
Publié: (2025)
par: Wu, Jinming, et autres
Publié: (2025)
LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMs
par: Zhou, Hanyu, et autres
Publié: (2025)
par: Zhou, Hanyu, et autres
Publié: (2025)
ViDRiP-LLaVA: A Dataset and Benchmark for Diagnostic Reasoning from Pathology Videos
par: Vuong, Trinh T. L., et autres
Publié: (2025)
par: Vuong, Trinh T. L., et autres
Publié: (2025)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
par: Zhao, Xiangyu, et autres
Publié: (2024)
par: Zhao, Xiangyu, et autres
Publié: (2024)
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
par: Xu, Guowei, et autres
Publié: (2024)
par: Xu, Guowei, et autres
Publié: (2024)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
par: Xu, Mingze, et autres
Publié: (2024)
par: Xu, Mingze, et autres
Publié: (2024)
AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs
par: Xia, Shuhan, et autres
Publié: (2025)
par: Xia, Shuhan, et autres
Publié: (2025)
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
par: An, Xiang, et autres
Publié: (2026)
par: An, Xiang, et autres
Publié: (2026)
TinyLLaVA: A Framework of Small-scale Large Multimodal Models
par: Zhou, Baichuan, et autres
Publié: (2024)
par: Zhou, Baichuan, et autres
Publié: (2024)
Interpolating Video-LLMs: Toward Longer-sequence LMMs in a Training-free Manner
par: Shang, Yuzhang, et autres
Publié: (2024)
par: Shang, Yuzhang, et autres
Publié: (2024)
TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs
par: Wang, Juntong, et autres
Publié: (2025)
par: Wang, Juntong, et autres
Publié: (2025)
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
par: Zhang, Boqiang, et autres
Publié: (2025)
par: Zhang, Boqiang, et autres
Publié: (2025)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
par: Zhang, Tao, et autres
Publié: (2024)
par: Zhang, Tao, et autres
Publié: (2024)
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
par: Sun, Boyuan, et autres
Publié: (2025)
par: Sun, Boyuan, et autres
Publié: (2025)
TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
par: Gao, Mingze, et autres
Publié: (2024)
par: Gao, Mingze, et autres
Publié: (2024)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
par: Zhang, Shaolei, et autres
Publié: (2025)
par: Zhang, Shaolei, et autres
Publié: (2025)
VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
par: Liu, Shuming, et autres
Publié: (2026)
par: Liu, Shuming, et autres
Publié: (2026)
Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
par: Zhang, Yi-Fan, et autres
Publié: (2024)
par: Zhang, Yi-Fan, et autres
Publié: (2024)
NOVO: Bridging LLaVA and SAM with Visual-only Prompts for Reasoning Segmentation
par: Yoon, Kyung-Yoon, et autres
Publié: (2025)
par: Yoon, Kyung-Yoon, et autres
Publié: (2025)
Documents similaires
-
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
par: Zhang, Xingjian, et autres
Publié: (2025) -
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
par: Wen, Siwei, et autres
Publié: (2026) -
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
par: Shu, Fangxun, et autres
Publié: (2024) -
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
par: Deng, Jiajun, et autres
Publié: (2025) -
LLaVA-Video: Video Instruction Tuning With Synthetic Data
par: Zhang, Yuanhan, et autres
Publié: (2024)