AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Xian, Wu, Zexi, Li, Zinuo, Xu, Hongming, Gong, Luqi, Boussaid, Farid, Werghi, Naoufel, Bennamoun, Mohammed |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM
di: Li, Zinuo, et al.
Pubblicazione: (2025)
di: Li, Zinuo, et al.
Pubblicazione: (2025)
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
di: Taghipour, Ashkan, et al.
Pubblicazione: (2026)
di: Taghipour, Ashkan, et al.
Pubblicazione: (2026)
Generalized Closed-form Formulae for Feature-based Subpixel Alignment in Patch-based Matching
di: Jospin, Laurent Valentin, et al.
Pubblicazione: (2021)
di: Jospin, Laurent Valentin, et al.
Pubblicazione: (2021)
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
di: Taghipour, Ashkan, et al.
Pubblicazione: (2024)
di: Taghipour, Ashkan, et al.
Pubblicazione: (2024)
LatentMove: Towards Complex Human Movement Video Generation
di: Taghipour, Ashkan, et al.
Pubblicazione: (2025)
di: Taghipour, Ashkan, et al.
Pubblicazione: (2025)
3D Brain and Heart Volume Generative Models: A Survey
di: Liu, Yanbin, et al.
Pubblicazione: (2022)
di: Liu, Yanbin, et al.
Pubblicazione: (2022)
STEER: Structured Event Evidence for Video Reasoning via Multi-Objective Reinforcement Learning
di: Li, Zinuo, et al.
Pubblicazione: (2026)
di: Li, Zinuo, et al.
Pubblicazione: (2026)
AquaticCLIP: A Vision-Language Foundation Model for Underwater Scene Analysis
di: Alawode, Basit, et al.
Pubblicazione: (2025)
di: Alawode, Basit, et al.
Pubblicazione: (2025)
Adaptive Keyframe Sampling for Long Video Understanding
di: Tang, Xi, et al.
Pubblicazione: (2025)
di: Tang, Xi, et al.
Pubblicazione: (2025)
CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language Alignment
di: Javed, Sajid, et al.
Pubblicazione: (2024)
di: Javed, Sajid, et al.
Pubblicazione: (2024)
AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding
di: Yang, Xiao, et al.
Pubblicazione: (2026)
di: Yang, Xiao, et al.
Pubblicazione: (2026)
AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection
di: Zhang, Shuheng, et al.
Pubblicazione: (2025)
di: Zhang, Shuheng, et al.
Pubblicazione: (2025)
SVR-GS: Spatially Variant Regularization for Probabilistic Masks in 3D Gaussian Splatting
di: Taghipour, Ashkan, et al.
Pubblicazione: (2025)
di: Taghipour, Ashkan, et al.
Pubblicazione: (2025)
Dynamic Neural Surfaces for Elastic 4D Shape Representation and Analysis
di: Nizamani, Awais, et al.
Pubblicazione: (2025)
di: Nizamani, Awais, et al.
Pubblicazione: (2025)
Hybrid Transformer-Mamba Architecture for Weakly Supervised Volumetric Medical Segmentation
di: Lyu, Yiheng, et al.
Pubblicazione: (2025)
di: Lyu, Yiheng, et al.
Pubblicazione: (2025)
Auxiliary Tasks Enhanced Dual-affinity Learning for Weakly Supervised Semantic Segmentation
di: Xu, Lian, et al.
Pubblicazione: (2024)
di: Xu, Lian, et al.
Pubblicazione: (2024)
Multi-Modal Attention Networks for Enhanced Segmentation and Depth Estimation of Subsurface Defects in Pulse Thermography
di: Salah, Mohammed, et al.
Pubblicazione: (2025)
di: Salah, Mohammed, et al.
Pubblicazione: (2025)
Box It to Bind It: Unified Layout Control and Attribute Binding in T2I Diffusion Models
di: Taghipour, Ashkan, et al.
Pubblicazione: (2024)
di: Taghipour, Ashkan, et al.
Pubblicazione: (2024)
BENet: A Cross-domain Robust Network for Detecting Face Forgeries via Bias Expansion and Latent-space Attention
di: Liu, Weihua, et al.
Pubblicazione: (2024)
di: Liu, Weihua, et al.
Pubblicazione: (2024)
Advancing Histopathology with Deep Learning Under Data Scarcity: A Decade in Review
di: Obeid, Ahmad, et al.
Pubblicazione: (2024)
di: Obeid, Ahmad, et al.
Pubblicazione: (2024)
Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
di: Albastaki, Shahad, et al.
Pubblicazione: (2025)
di: Albastaki, Shahad, et al.
Pubblicazione: (2025)
DynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action Recognition
di: Zhu, Jingmin, et al.
Pubblicazione: (2025)
di: Zhu, Jingmin, et al.
Pubblicazione: (2025)
Fact or Fake? Assessing the Role of Deepfake Detectors in Multimodal Misinformation Detection
di: Sagar, A S M Sharifuzzaman, et al.
Pubblicazione: (2026)
di: Sagar, A S M Sharifuzzaman, et al.
Pubblicazione: (2026)
A Riemannian Approach for Spatiotemporal Analysis and Generation of 4D Tree-shaped Structures
di: Khanam, Tahmina, et al.
Pubblicazione: (2024)
di: Khanam, Tahmina, et al.
Pubblicazione: (2024)
A Riemannian Framework for the Elastic Analysis of the Spatiotemporal Variability in the Shape and Structure of Tree-like 4D Objects
di: Khanam, Tahmina, et al.
Pubblicazione: (2025)
di: Khanam, Tahmina, et al.
Pubblicazione: (2025)
SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition
di: Wang, Ning, et al.
Pubblicazione: (2026)
di: Wang, Ning, et al.
Pubblicazione: (2026)
Video Anomaly Detection in 10 Years: A Survey and Outlook
di: Abdalla, Moshira, et al.
Pubblicazione: (2024)
di: Abdalla, Moshira, et al.
Pubblicazione: (2024)
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
di: Alansari, Mohamad, et al.
Pubblicazione: (2026)
di: Alansari, Mohamad, et al.
Pubblicazione: (2026)
UIFormer: A Unified Transformer-based Framework for Incremental Few-Shot Object Detection and Instance Segmentation
di: Zhang, Chengyuan, et al.
Pubblicazione: (2024)
di: Zhang, Chengyuan, et al.
Pubblicazione: (2024)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
di: Wang, Yiheng, et al.
Pubblicazione: (2026)
di: Wang, Yiheng, et al.
Pubblicazione: (2026)
RobMOT: Robust 3D Multi-Object Tracking by Observational Noise and State Estimation Drift Mitigation on LiDAR PointCloud
di: Nagy, Mohamed, et al.
Pubblicazione: (2024)
di: Nagy, Mohamed, et al.
Pubblicazione: (2024)
Towards Accurate State Estimation: Kalman Filter Incorporating Motion Dynamics for 3D Multi-Object Tracking
di: Nagy, Mohamed, et al.
Pubblicazione: (2025)
di: Nagy, Mohamed, et al.
Pubblicazione: (2025)
Admitting Ignorance Helps the Video Question Answering Models to Answer
di: Li, Haopeng, et al.
Pubblicazione: (2025)
di: Li, Haopeng, et al.
Pubblicazione: (2025)
FOCUS: Efficient Keyframe Selection for Long Video Understanding
di: Zhu, Zirui, et al.
Pubblicazione: (2025)
di: Zhu, Zirui, et al.
Pubblicazione: (2025)
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
di: Li, Handong, et al.
Pubblicazione: (2026)
di: Li, Handong, et al.
Pubblicazione: (2026)
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
di: Xu, Chuanzhi, et al.
Pubblicazione: (2026)
di: Xu, Chuanzhi, et al.
Pubblicazione: (2026)
Rethinking Memory Design in SAM-Based Visual Object Tracking
di: Alansari, Mohamad, et al.
Pubblicazione: (2025)
di: Alansari, Mohamad, et al.
Pubblicazione: (2025)
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
di: Shlapentokh-Rothman, Michal, et al.
Pubblicazione: (2026)
di: Shlapentokh-Rothman, Michal, et al.
Pubblicazione: (2026)
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
di: Fang, Bo, et al.
Pubblicazione: (2025)
di: Fang, Bo, et al.
Pubblicazione: (2025)
A Robust Adversary Detection-Deactivation Method for Metaverse-oriented Collaborative Deep Learning
di: Li, Pengfei, et al.
Pubblicazione: (2023)
di: Li, Pengfei, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM
di: Li, Zinuo, et al.
Pubblicazione: (2025) -
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
di: Taghipour, Ashkan, et al.
Pubblicazione: (2026) -
Generalized Closed-form Formulae for Feature-based Subpixel Alignment in Patch-based Matching
di: Jospin, Laurent Valentin, et al.
Pubblicazione: (2021) -
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
di: Taghipour, Ashkan, et al.
Pubblicazione: (2024) -
LatentMove: Towards Complex Human Movement Video Generation
di: Taghipour, Ashkan, et al.
Pubblicazione: (2025)