Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Shaojie, Yang, Jiahui, Yin, Jianqin, Luo, Zhenbo, Luan, Jian |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
par: Tan, Wenhui, et autres
Publié: (2026)
par: Tan, Wenhui, et autres
Publié: (2026)
Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models
par: Xu, Longwei, et autres
Publié: (2026)
par: Xu, Longwei, et autres
Publié: (2026)
Frame-Voyager: Learning to Query Frames for Video Large Language Models
par: Yu, Sicheng, et autres
Publié: (2024)
par: Yu, Sicheng, et autres
Publié: (2024)
HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
par: Yang, Yiqing, et autres
Publié: (2025)
par: Yang, Yiqing, et autres
Publié: (2025)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
par: Guan, Yiran, et autres
Publié: (2026)
par: Guan, Yiran, et autres
Publié: (2026)
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
par: Zhang, Shaojie, et autres
Publié: (2023)
par: Zhang, Shaojie, et autres
Publié: (2023)
IMTBench: A Multi-Scenario Cross-Modal Collaborative Evaluation Benchmark for In-Image Machine Translation
par: Lyu, Jiahao, et autres
Publié: (2026)
par: Lyu, Jiahao, et autres
Publié: (2026)
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
par: Kim, Jihwan, et autres
Publié: (2026)
par: Kim, Jihwan, et autres
Publié: (2026)
Motion-aware Latent Diffusion Models for Video Frame Interpolation
par: Huang, Zhilin, et autres
Publié: (2024)
par: Huang, Zhilin, et autres
Publié: (2024)
Kinematics Modeling Network for Video-based Human Pose Estimation
par: Dang, Yonghao, et autres
Publié: (2022)
par: Dang, Yonghao, et autres
Publié: (2022)
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
par: Zhang, Deyu, et autres
Publié: (2025)
par: Zhang, Deyu, et autres
Publié: (2025)
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
par: Ghazanfari, Sara, et autres
Publié: (2025)
par: Ghazanfari, Sara, et autres
Publié: (2025)
Burst Image Super-Resolution with Base Frame Selection
par: Kim, Sanghyun, et autres
Publié: (2024)
par: Kim, Sanghyun, et autres
Publié: (2024)
Event-Anchored Frame Selection for Effective Long-Video Understanding
par: Chen, Wang, et autres
Publié: (2026)
par: Chen, Wang, et autres
Publié: (2026)
Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning
par: Li, Yifan, et autres
Publié: (2025)
par: Li, Yifan, et autres
Publié: (2025)
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
par: Zhang, Yulin, et autres
Publié: (2026)
par: Zhang, Yulin, et autres
Publié: (2026)
MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
par: Wei, Wei, et autres
Publié: (2025)
par: Wei, Wei, et autres
Publié: (2025)
SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
par: Zhang, Shaojie, et autres
Publié: (2023)
par: Zhang, Shaojie, et autres
Publié: (2023)
Velocity Disambiguation for Video Frame Interpolation
par: Zhong, Zhihang, et autres
Publié: (2023)
par: Zhong, Zhihang, et autres
Publié: (2023)
Video Super-Resolution Transformer with Masked Inter&Intra-Frame Attention
par: Zhou, Xingyu, et autres
Publié: (2024)
par: Zhou, Xingyu, et autres
Publié: (2024)
ZeroSmooth: Training-free Diffuser Adaptation for High Frame Rate Video Generation
par: Yang, Shaoshu, et autres
Publié: (2024)
par: Yang, Shaoshu, et autres
Publié: (2024)
Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
par: Chen, Wang, et autres
Publié: (2026)
par: Chen, Wang, et autres
Publié: (2026)
MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs
par: Sun, Hui, et autres
Publié: (2025)
par: Sun, Hui, et autres
Publié: (2025)
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
par: He, Zefeng, et autres
Publié: (2025)
par: He, Zefeng, et autres
Publié: (2025)
PMQ-VE: Progressive Multi-Frame Quantization for Video Enhancement
par: Feng, ZhanFeng, et autres
Publié: (2025)
par: Feng, ZhanFeng, et autres
Publié: (2025)
A Single-Frame and Multi-Frame Cascaded Image Super-Resolution Method
par: Sun, Jing, et autres
Publié: (2024)
par: Sun, Jing, et autres
Publié: (2024)
Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
par: Zhang, Shaojie, et autres
Publié: (2025)
par: Zhang, Shaojie, et autres
Publié: (2025)
Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval
par: Yu, Jiaao, et autres
Publié: (2025)
par: Yu, Jiaao, et autres
Publié: (2025)
Adaptify: A Refined Adaptation Scheme for Frame Classification in Atrophic Gastritis Videos
par: Xiong, Zinan, et autres
Publié: (2024)
par: Xiong, Zinan, et autres
Publié: (2024)
Joint Video Enhancement with Deblurring, Super-Resolution, and Frame Interpolation Network
par: Choi, Giyong, et autres
Publié: (2025)
par: Choi, Giyong, et autres
Publié: (2025)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
par: Rahman, Aimon, et autres
Publié: (2024)
par: Rahman, Aimon, et autres
Publié: (2024)
Progress-Aware Video Frame Captioning
par: Xue, Zihui, et autres
Publié: (2024)
par: Xue, Zihui, et autres
Publié: (2024)
M-LLM Based Video Frame Selection for Efficient Video Understanding
par: Hu, Kai, et autres
Publié: (2025)
par: Hu, Kai, et autres
Publié: (2025)
Motion-Aware Video Frame Interpolation
par: Han, Pengfei, et autres
Publié: (2024)
par: Han, Pengfei, et autres
Publié: (2024)
Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
par: Zhang, Lvmin, et autres
Publié: (2025)
par: Zhang, Lvmin, et autres
Publié: (2025)
LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs
par: Chen, Jingfeng, et autres
Publié: (2026)
par: Chen, Jingfeng, et autres
Publié: (2026)
Benchmarking Video Frame Interpolation
par: Kiefhaber, Simon, et autres
Publié: (2024)
par: Kiefhaber, Simon, et autres
Publié: (2024)
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
par: Yang, Shaoshu, et autres
Publié: (2025)
par: Yang, Shaoshu, et autres
Publié: (2025)
Adaptive Greedy Frame Selection for Long Video Understanding
par: Huang, Yuning, et autres
Publié: (2026)
par: Huang, Yuning, et autres
Publié: (2026)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
par: Li, Jialuo, et autres
Publié: (2025)
par: Li, Jialuo, et autres
Publié: (2025)
Documents similaires
-
Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
par: Tan, Wenhui, et autres
Publié: (2026) -
Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models
par: Xu, Longwei, et autres
Publié: (2026) -
Frame-Voyager: Learning to Query Frames for Video Large Language Models
par: Yu, Sicheng, et autres
Publié: (2024) -
HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
par: Yang, Yiqing, et autres
Publié: (2025) -
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
par: Guan, Yiran, et autres
Publié: (2026)