Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Yuxuan, Wang, Yueqian, Wu, Pengfei, Liang, Jianxin, Zhao, Dongyan, Liu, Yang, Zheng, Zilong |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
par: Wang, Yueqian, et autres
Publié: (2024)
par: Wang, Yueqian, et autres
Publié: (2024)
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
par: Wang, Yueqian, et autres
Publié: (2024)
par: Wang, Yueqian, et autres
Publié: (2024)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
par: Liang, Jianxin, et autres
Publié: (2025)
par: Liang, Jianxin, et autres
Publié: (2025)
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
par: Wang, Yuxuan, et autres
Publié: (2024)
par: Wang, Yuxuan, et autres
Publié: (2024)
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
par: Wang, Yueqian, et autres
Publié: (2024)
par: Wang, Yueqian, et autres
Publié: (2024)
Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
par: Wang, Yueqian, et autres
Publié: (2024)
par: Wang, Yueqian, et autres
Publié: (2024)
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
par: Liang, Jianxin, et autres
Publié: (2024)
par: Liang, Jianxin, et autres
Publié: (2024)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
par: Liang, Jianxin, et autres
Publié: (2025)
par: Liang, Jianxin, et autres
Publié: (2025)
OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts
par: Wang, Yuxuan, et autres
Publié: (2025)
par: Wang, Yuxuan, et autres
Publié: (2025)
MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning
par: Wang, Jianghui, et autres
Publié: (2023)
par: Wang, Jianghui, et autres
Publié: (2023)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
par: Chen, Jiaxing, et autres
Publié: (2024)
par: Chen, Jiaxing, et autres
Publié: (2024)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
par: Wang, Yuxuan, et autres
Publié: (2024)
par: Wang, Yuxuan, et autres
Publié: (2024)
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
par: Wang, Ye, et autres
Publié: (2025)
par: Wang, Ye, et autres
Publié: (2025)
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
par: Wang, Yueqian, et autres
Publié: (2025)
par: Wang, Yueqian, et autres
Publié: (2025)
MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
par: Wang, Yueqian, et autres
Publié: (2025)
par: Wang, Yueqian, et autres
Publié: (2025)
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
par: Shangguan, Ziyao, et autres
Publié: (2024)
par: Shangguan, Ziyao, et autres
Publié: (2024)
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
par: Fan, Rong, et autres
Publié: (2026)
par: Fan, Rong, et autres
Publié: (2026)
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
par: Li, Yun, et autres
Publié: (2025)
par: Li, Yun, et autres
Publié: (2025)
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
par: Yin, Zhihan, et autres
Publié: (2026)
par: Yin, Zhihan, et autres
Publié: (2026)
Factorized Learning for Temporally Grounded Video-Language Models
par: Zeng, Wenzheng, et autres
Publié: (2025)
par: Zeng, Wenzheng, et autres
Publié: (2025)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
par: Zhang, Jun, et autres
Publié: (2025)
par: Zhang, Jun, et autres
Publié: (2025)
Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding
par: Wang, Yueqian, et autres
Publié: (2024)
par: Wang, Yueqian, et autres
Publié: (2024)
Teaching Text-to-Image Models to Communicate in Dialog
par: Sun, Xiaowen, et autres
Publié: (2023)
par: Sun, Xiaowen, et autres
Publié: (2023)
HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
par: HyperAI Team, et autres
Publié: (2025)
par: HyperAI Team, et autres
Publié: (2025)
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
par: Li, Jinyuan, et autres
Publié: (2024)
par: Li, Jinyuan, et autres
Publié: (2024)
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
par: Nguyen, Thong, et autres
Publié: (2023)
par: Nguyen, Thong, et autres
Publié: (2023)
A Survey on Video Temporal Grounding with Multimodal Large Language Model
par: Wu, Jianlong, et autres
Publié: (2025)
par: Wu, Jianlong, et autres
Publié: (2025)
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
par: Liang, Yiming, et autres
Publié: (2026)
par: Liang, Yiming, et autres
Publié: (2026)
AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
par: Huang, Yipo, et autres
Publié: (2024)
par: Huang, Yipo, et autres
Publié: (2024)
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
par: Hannan, Tanveer, et autres
Publié: (2024)
par: Hannan, Tanveer, et autres
Publié: (2024)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
par: Gao, Shida, et autres
Publié: (2025)
par: Gao, Shida, et autres
Publié: (2025)
Detecting Offensive Memes with Social Biases in Singapore Context Using Multimodal Large Language Models
par: Yuxuan, Cao, et autres
Publié: (2025)
par: Yuxuan, Cao, et autres
Publié: (2025)
DeepSight: Bridging Depth Maps and Language with a Depth-Driven Multimodal Model
par: Yang, Hao, et autres
Publié: (2026)
par: Yang, Hao, et autres
Publié: (2026)
MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models
par: Wang, Zhongxi, et autres
Publié: (2026)
par: Wang, Zhongxi, et autres
Publié: (2026)
Probing and Inducing Combinational Creativity in Vision-Language Models
par: Peng, Yongqian, et autres
Publié: (2025)
par: Peng, Yongqian, et autres
Publié: (2025)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
par: Wang, Jiankang, et autres
Publié: (2025)
par: Wang, Jiankang, et autres
Publié: (2025)
Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging
par: Fu, Zihang, et autres
Publié: (2026)
par: Fu, Zihang, et autres
Publié: (2026)
Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
par: Cui, Yiming, et autres
Publié: (2025)
par: Cui, Yiming, et autres
Publié: (2025)
Unhackable Temporal Rewarding for Scalable Video MLLMs
par: Yu, En, et autres
Publié: (2025)
par: Yu, En, et autres
Publié: (2025)
EFUF: Efficient Fine-grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models
par: Xing, Shangyu, et autres
Publié: (2024)
par: Xing, Shangyu, et autres
Publié: (2024)
Documents similaires
-
STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
par: Wang, Yueqian, et autres
Publié: (2024) -
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
par: Wang, Yueqian, et autres
Publié: (2024) -
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
par: Liang, Jianxin, et autres
Publié: (2025) -
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
par: Wang, Yuxuan, et autres
Publié: (2024) -
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
par: Wang, Yueqian, et autres
Publié: (2024)