DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Peng, Zhang, Guanghao, He, Wanggui, Zhang, Longxiang, Liu, Mushui, Xia, Yan, Peng, Zhenhao, Dai, Weilong, Liu, Jinlong, Tang, Haobing, Zhang, Le, Jiang, Hao, Huang, Pipei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture
por: Zhang, Longxiang, et al.
Publicado: (2026)
por: Zhang, Longxiang, et al.
Publicado: (2026)
PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning
por: Liu, Jinlong, et al.
Publicado: (2026)
por: Liu, Jinlong, et al.
Publicado: (2026)
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO
por: Tong, Yunze, et al.
Publicado: (2026)
por: Tong, Yunze, et al.
Publicado: (2026)
Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
por: Wang, Yi, et al.
Publicado: (2025)
por: Wang, Yi, et al.
Publicado: (2025)
CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
por: Zhang, Guanghao, et al.
Publicado: (2025)
por: Zhang, Guanghao, et al.
Publicado: (2025)
Adaptive Greedy Frame Selection for Long Video Understanding
por: Huang, Yuning, et al.
Publicado: (2026)
por: Huang, Yuning, et al.
Publicado: (2026)
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
por: Ghazanfari, Sara, et al.
Publicado: (2025)
por: Ghazanfari, Sara, et al.
Publicado: (2025)
CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers
por: She, D., et al.
Publicado: (2025)
por: She, D., et al.
Publicado: (2025)
GIFT: Global Irreplaceability Frame Targeting for Efficient Video Understanding
por: Ma, Junpeng, et al.
Publicado: (2026)
por: Ma, Junpeng, et al.
Publicado: (2026)
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
por: Bao, Xiaoyi, et al.
Publicado: (2025)
por: Bao, Xiaoyi, et al.
Publicado: (2025)
Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition
por: Li, Bozheng, et al.
Publicado: (2024)
por: Li, Bozheng, et al.
Publicado: (2024)
Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
por: Huang, Qihan, et al.
Publicado: (2025)
por: Huang, Qihan, et al.
Publicado: (2025)
Shot-Aware Frame Sampling for Video Understanding
por: Zhao, Mengyu, et al.
Publicado: (2026)
por: Zhao, Mengyu, et al.
Publicado: (2026)
PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation
por: Huang, Qihan, et al.
Publicado: (2024)
por: Huang, Qihan, et al.
Publicado: (2024)
Generative Frame Sampler for Long Video Understanding
por: Yao, Linli, et al.
Publicado: (2025)
por: Yao, Linli, et al.
Publicado: (2025)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
por: Zhang, Hongzhi, et al.
Publicado: (2025)
por: Zhang, Hongzhi, et al.
Publicado: (2025)
Spatial-ORMLLM: Improve Spatial Relation Understanding in the Operating Room with Multimodal Large Language Model
por: He, Peiqi, et al.
Publicado: (2025)
por: He, Peiqi, et al.
Publicado: (2025)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
por: Rahman, Aimon, et al.
Publicado: (2024)
por: Rahman, Aimon, et al.
Publicado: (2024)
SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix
por: Dai, Peng, et al.
Publicado: (2024)
por: Dai, Peng, et al.
Publicado: (2024)
Frame-Voyager: Learning to Query Frames for Video Large Language Models
por: Yu, Sicheng, et al.
Publicado: (2024)
por: Yu, Sicheng, et al.
Publicado: (2024)
Improving LLM Video Understanding with 16 Frames Per Second
por: Li, Yixuan, et al.
Publicado: (2025)
por: Li, Yixuan, et al.
Publicado: (2025)
M-LLM Based Video Frame Selection for Efficient Video Understanding
por: Hu, Kai, et al.
Publicado: (2025)
por: Hu, Kai, et al.
Publicado: (2025)
Sparse Global Matching for Video Frame Interpolation with Large Motion
por: Liu, Chunxu, et al.
Publicado: (2024)
por: Liu, Chunxu, et al.
Publicado: (2024)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
por: Ge, Haonan, et al.
Publicado: (2025)
por: Ge, Haonan, et al.
Publicado: (2025)
Motion-Aware Video Frame Interpolation
por: Han, Pengfei, et al.
Publicado: (2024)
por: Han, Pengfei, et al.
Publicado: (2024)
FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing
por: Huang, Xijie, et al.
Publicado: (2026)
por: Huang, Xijie, et al.
Publicado: (2026)
VFIMamba: Video Frame Interpolation with State Space Models
por: Zhang, Guozhen, et al.
Publicado: (2024)
por: Zhang, Guozhen, et al.
Publicado: (2024)
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
por: Zhang, Deyu, et al.
Publicado: (2025)
por: Zhang, Deyu, et al.
Publicado: (2025)
Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
por: Zhang, Lvmin, et al.
Publicado: (2025)
por: Zhang, Lvmin, et al.
Publicado: (2025)
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
por: Zhang, Shaojie, et al.
Publicado: (2025)
por: Zhang, Shaojie, et al.
Publicado: (2025)
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
por: Huang, De-An, et al.
Publicado: (2025)
por: Huang, De-An, et al.
Publicado: (2025)
Learned Rate Control for Frame-Level Adaptive Neural Video Compression via Dynamic Neural Network
por: Zhang, Chenhao, et al.
Publicado: (2025)
por: Zhang, Chenhao, et al.
Publicado: (2025)
S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix
por: Dai, Peng, et al.
Publicado: (2025)
por: Dai, Peng, et al.
Publicado: (2025)
Follow the Clues, Frame the Truth: Hybrid-evidential Deductive Reasoning in Open-Vocabulary Multimodal Emotion Recognition
por: Liu, Yu, et al.
Publicado: (2026)
por: Liu, Yu, et al.
Publicado: (2026)
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
por: Chen, Chen, et al.
Publicado: (2025)
por: Chen, Chen, et al.
Publicado: (2025)
LDMVFI: Video Frame Interpolation with Latent Diffusion Models
por: Danier, Duolikun, et al.
Publicado: (2023)
por: Danier, Duolikun, et al.
Publicado: (2023)
A Subjective Quality Study for Video Frame Interpolation
por: Danier, Duolikun, et al.
Publicado: (2022)
por: Danier, Duolikun, et al.
Publicado: (2022)
BVI-VFI: A Video Quality Database for Video Frame Interpolation
por: Danier, Duolikun, et al.
Publicado: (2022)
por: Danier, Duolikun, et al.
Publicado: (2022)
Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
por: Chen, Wang, et al.
Publicado: (2026)
por: Chen, Wang, et al.
Publicado: (2026)
Adaptify: A Refined Adaptation Scheme for Frame Classification in Atrophic Gastritis Videos
por: Xiong, Zinan, et al.
Publicado: (2024)
por: Xiong, Zinan, et al.
Publicado: (2024)
Ejemplares similares
-
Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture
por: Zhang, Longxiang, et al.
Publicado: (2026) -
PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning
por: Liu, Jinlong, et al.
Publicado: (2026) -
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO
por: Tong, Yunze, et al.
Publicado: (2026) -
Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
por: Wang, Yi, et al.
Publicado: (2025) -
CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
por: Zhang, Guanghao, et al.
Publicado: (2025)