TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Shicheng, Li, Lei, Ouyang, Kun, Ren, Shuhuai, Liu, Yuanxin, Zhang, Yuanxing, Zhang, Fuzheng, Kong, Lingpeng, Liu, Qi, Sun, Xu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
von: Li, Shicheng, et al.
Veröffentlicht: (2023)
von: Li, Shicheng, et al.
Veröffentlicht: (2023)
TempCompass: Do Video LLMs Really Understand Videos?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2024)
von: Liu, Yuanxin, et al.
Veröffentlicht: (2024)
Temporal Reasoning Transfer from Text to Video
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
von: Ren, Shuhuai, et al.
Veröffentlicht: (2023)
von: Ren, Shuhuai, et al.
Veröffentlicht: (2023)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
von: Yao, Linli, et al.
Veröffentlicht: (2026)
von: Yao, Linli, et al.
Veröffentlicht: (2026)
DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
von: Yao, Linli, et al.
Veröffentlicht: (2024)
von: Yao, Linli, et al.
Veröffentlicht: (2024)
Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality
von: Chen, Sishuo, et al.
Veröffentlicht: (2024)
von: Chen, Sishuo, et al.
Veröffentlicht: (2024)
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
von: Wang, Yuchi, et al.
Veröffentlicht: (2025)
von: Wang, Yuchi, et al.
Veröffentlicht: (2025)
Generative Frame Sampler for Long Video Understanding
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
von: Zhang, Hongzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Hongzhi, et al.
Veröffentlicht: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
von: Xu, Zhiyang, et al.
Veröffentlicht: (2026)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2026)
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
von: Li, Yun, et al.
Veröffentlicht: (2025)
von: Li, Yun, et al.
Veröffentlicht: (2025)
Clapper: Compact Learning and Video Representation in VLMs
von: Kong, Lingyu, et al.
Veröffentlicht: (2025)
von: Kong, Lingyu, et al.
Veröffentlicht: (2025)
VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
von: Tao, Xijia, et al.
Veröffentlicht: (2024)
von: Tao, Xijia, et al.
Veröffentlicht: (2024)
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models
von: Yi, Hao, et al.
Veröffentlicht: (2024)
von: Yi, Hao, et al.
Veröffentlicht: (2024)
Understanding the Mechanism of Altruism in Large Language Models
von: Zhang, Shuhuai, et al.
Veröffentlicht: (2026)
von: Zhang, Shuhuai, et al.
Veröffentlicht: (2026)
Jailbreaking as a Reward Misspecification Problem
von: Xie, Zhihui, et al.
Veröffentlicht: (2024)
von: Xie, Zhihui, et al.
Veröffentlicht: (2024)
Crowd-SFT: Crowdsourcing for LLM Alignment
von: Sotiropoulos, Alex, et al.
Veröffentlicht: (2025)
von: Sotiropoulos, Alex, et al.
Veröffentlicht: (2025)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
BBA: Bi-Modal Behavioral Alignment for Reasoning with Large Vision-Language Models
von: Zhao, Xueliang, et al.
Veröffentlicht: (2024)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2024)
V-CORE: Temporally Consistent Video Understanding for Video-LLM
von: Kang, Zhengjian, et al.
Veröffentlicht: (2026)
von: Kang, Zhengjian, et al.
Veröffentlicht: (2026)
Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
Incentive Effects of a Cut-Off Score: Optimal Contest Design with Transparent Pre-Selection
von: Liu, Hanbing, et al.
Veröffentlicht: (2026)
von: Liu, Hanbing, et al.
Veröffentlicht: (2026)
Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
von: Qiu, Haibo, et al.
Veröffentlicht: (2025)
von: Qiu, Haibo, et al.
Veröffentlicht: (2025)
Vibe AIGC: A New Paradigm for Content Generation via Agentic Orchestration
von: Liu, Jiaheng, et al.
Veröffentlicht: (2026)
von: Liu, Jiaheng, et al.
Veröffentlicht: (2026)
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
Wasserstein Dependent Graph Attention Network for Collaborative Filtering with Uncertainty
von: Li, Haoxuan, et al.
Veröffentlicht: (2024)
von: Li, Haoxuan, et al.
Veröffentlicht: (2024)
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
von: Shi, Yang, et al.
Veröffentlicht: (2025)
von: Shi, Yang, et al.
Veröffentlicht: (2025)
ProCQA: A Large-scale Community-based Programming Question Answering Dataset for Code Search
von: Li, Zehan, et al.
Veröffentlicht: (2024)
von: Li, Zehan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
von: Yao, Linli, et al.
Veröffentlicht: (2025) -
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
von: Li, Shicheng, et al.
Veröffentlicht: (2023) -
TempCompass: Do Video LLMs Really Understand Videos?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2024) -
Temporal Reasoning Transfer from Text to Video
von: Li, Lei, et al.
Veröffentlicht: (2024) -
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
von: Ren, Shuhuai, et al.
Veröffentlicht: (2023)