DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Bao, Xiaoyi, Xie, Chenwei, Tang, Hao, Weng, Tingyu, Wang, Xiaofeng, Zheng, Yun, Wang, Xingang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
por: Tang, Hao, et al.
Publicado: (2025)
por: Tang, Hao, et al.
Publicado: (2025)
UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface
por: Tang, Hao, et al.
Publicado: (2025)
por: Tang, Hao, et al.
Publicado: (2025)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
por: Zhang, Peng, et al.
Publicado: (2026)
por: Zhang, Peng, et al.
Publicado: (2026)
DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation
por: Zhao, Guosheng, et al.
Publicado: (2024)
por: Zhao, Guosheng, et al.
Publicado: (2024)
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
por: Wang, Xiaofeng, et al.
Publicado: (2024)
por: Wang, Xiaofeng, et al.
Publicado: (2024)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
por: Zhang, Hongzhi, et al.
Publicado: (2025)
por: Zhang, Hongzhi, et al.
Publicado: (2025)
GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning
por: Bao, Xiaoyi, et al.
Publicado: (2025)
por: Bao, Xiaoyi, et al.
Publicado: (2025)
Aligned Better, Listen Better for Audio-Visual Large Language Models
por: Guo, Yuxin, et al.
Publicado: (2025)
por: Guo, Yuxin, et al.
Publicado: (2025)
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
por: Ma, Shuailei, et al.
Publicado: (2023)
por: Ma, Shuailei, et al.
Publicado: (2023)
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
por: Lee, Yeonkyung, et al.
Publicado: (2026)
por: Lee, Yeonkyung, et al.
Publicado: (2026)
CoReS: Orchestrating the Dance of Reasoning and Segmentation
por: Bao, Xiaoyi, et al.
Publicado: (2024)
por: Bao, Xiaoyi, et al.
Publicado: (2024)
Horizontal Versus Vertical: The Visual Balance Effects of Comparative Price Presentation
por: Xingang Wang, et al.
Publicado: (2024)
por: Xingang Wang, et al.
Publicado: (2024)
Dyn-Adapter: Towards Disentangled Representation for Efficient Visual Recognition
por: Zhang, Yurong, et al.
Publicado: (2024)
por: Zhang, Yurong, et al.
Publicado: (2024)
AG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal Models
por: Qu, Zhen, et al.
Publicado: (2026)
por: Qu, Zhen, et al.
Publicado: (2026)
Event-Anchored Frame Selection for Effective Long-Video Understanding
por: Chen, Wang, et al.
Publicado: (2026)
por: Chen, Wang, et al.
Publicado: (2026)
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
por: Han, Yudong, et al.
Publicado: (2024)
por: Han, Yudong, et al.
Publicado: (2024)
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
por: Xu, Runsen, et al.
Publicado: (2025)
por: Xu, Runsen, et al.
Publicado: (2025)
ReconDreamer++: Harmonizing Generative and Reconstructive Models for Driving Scene Representation
por: Zhao, Guosheng, et al.
Publicado: (2025)
por: Zhao, Guosheng, et al.
Publicado: (2025)
Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
por: Chen, Wang, et al.
Publicado: (2026)
por: Chen, Wang, et al.
Publicado: (2026)
Dyn-O: Building Structured World Models with Object-Centric Representations
por: Wang, Zizhao, et al.
Publicado: (2025)
por: Wang, Zizhao, et al.
Publicado: (2025)
KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding
por: Li, Zongyao, et al.
Publicado: (2025)
por: Li, Zongyao, et al.
Publicado: (2025)
Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo
por: Xu, Ninghui, et al.
Publicado: (2026)
por: Xu, Ninghui, et al.
Publicado: (2026)
I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models
por: Ouyang, Wenqi, et al.
Publicado: (2024)
por: Ouyang, Wenqi, et al.
Publicado: (2024)
Shot-Aware Frame Sampling for Video Understanding
por: Zhao, Mengyu, et al.
Publicado: (2026)
por: Zhao, Mengyu, et al.
Publicado: (2026)
A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech Recognition
por: Dai, Yusheng, et al.
Publicado: (2024)
por: Dai, Yusheng, et al.
Publicado: (2024)
From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding
por: Lin, Shih-Yao, et al.
Publicado: (2025)
por: Lin, Shih-Yao, et al.
Publicado: (2025)
Supervised Learning Model for Key Frame Identification from Cow Teat Videos
por: Wang, Minghao, et al.
Publicado: (2024)
por: Wang, Minghao, et al.
Publicado: (2024)
Towards Visual-Prompt Temporal Answering Grounding in Medical Instructional Video
por: Li, Bin, et al.
Publicado: (2022)
por: Li, Bin, et al.
Publicado: (2022)
Video Finetuning Improves Reasoning Between Frames
por: Yang, Ruiqi, et al.
Publicado: (2025)
por: Yang, Ruiqi, et al.
Publicado: (2025)
GIFT: Global Irreplaceability Frame Targeting for Efficient Video Understanding
por: Ma, Junpeng, et al.
Publicado: (2026)
por: Ma, Junpeng, et al.
Publicado: (2026)
ProxyImg: Towards Highly-Controllable Image Representation via Hierarchical Disentangled Proxy Embedding
por: Chen, Ye, et al.
Publicado: (2026)
por: Chen, Ye, et al.
Publicado: (2026)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
por: Li, Huilai, et al.
Publicado: (2026)
por: Li, Huilai, et al.
Publicado: (2026)
Metacognitive Prompting Improves Understanding in Large Language Models
por: Wang, Yuqing, et al.
Publicado: (2023)
por: Wang, Yuqing, et al.
Publicado: (2023)
Open-Ended Multi-Modal Relational Reasoning for Video Question Answering
por: Luo, Haozheng, et al.
Publicado: (2020)
por: Luo, Haozheng, et al.
Publicado: (2020)
Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding
por: Guo, Weiyu, et al.
Publicado: (2025)
por: Guo, Weiyu, et al.
Publicado: (2025)
Precise Action-to-Video Generation Through Visual Action Prompts
por: Wang, Yuang, et al.
Publicado: (2025)
por: Wang, Yuang, et al.
Publicado: (2025)
DynVFX: Augmenting Real Videos with Dynamic Content
por: Yatim, Danah, et al.
Publicado: (2025)
por: Yatim, Danah, et al.
Publicado: (2025)
Beyond Boundary Frames: Context-Centric Video Interpolation with Audio-Visual Semantics
por: Deng, Yuchen, et al.
Publicado: (2025)
por: Deng, Yuchen, et al.
Publicado: (2025)
SegImgNet: Segmentation-Guided Dual-Branch Network for Retinal Disease Diagnoses
por: Luo, Xinwei, et al.
Publicado: (2025)
por: Luo, Xinwei, et al.
Publicado: (2025)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
por: Rahman, Aimon, et al.
Publicado: (2024)
por: Rahman, Aimon, et al.
Publicado: (2024)
Ejemplares similares
-
UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
por: Tang, Hao, et al.
Publicado: (2025) -
UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface
por: Tang, Hao, et al.
Publicado: (2025) -
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
por: Zhang, Peng, et al.
Publicado: (2026) -
DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation
por: Zhao, Guosheng, et al.
Publicado: (2024) -
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
por: Wang, Xiaofeng, et al.
Publicado: (2024)