Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Lin, Xiao, Zhihan, Xu, Haohang, Cong, Rong, Zhang, Zhibo, Zhang, Xiaopeng, Tian, Qi |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance
par: Xu, Haohang, et autres
Publié: (2026)
par: Xu, Haohang, et autres
Publié: (2026)
Reasoning to Align: Implicit Reasoning in Diffusion Transformers for Video Editing
par: Li, Yan, et autres
Publié: (2026)
par: Li, Yan, et autres
Publié: (2026)
AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection
par: Zhang, Shuheng, et autres
Publié: (2025)
par: Zhang, Shuheng, et autres
Publié: (2025)
LoVoRA: Text-guided and Mask-free Video Object Removal and Addition with Learnable Object-aware Localization
par: Xiao, Zhihan, et autres
Publié: (2025)
par: Xiao, Zhihan, et autres
Publié: (2025)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
par: He, Haibin, et autres
Publié: (2026)
par: He, Haibin, et autres
Publié: (2026)
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
par: Nie, Ming, et autres
Publié: (2025)
par: Nie, Ming, et autres
Publié: (2025)
FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions
par: Zhao, Peisen, et autres
Publié: (2026)
par: Zhao, Peisen, et autres
Publié: (2026)
KS-APR: Keyframe Selection for Robust Absolute Pose Regression
par: Liu, Changkun, et autres
Publié: (2023)
par: Liu, Changkun, et autres
Publié: (2023)
KeyframeFace: Language-Driven Facial Animation via Semantic Keyframes
par: Wu, Jingchao, et autres
Publié: (2025)
par: Wu, Jingchao, et autres
Publié: (2025)
Robust Video-Based Pothole Detection and Area Estimation for Intelligent Vehicles with Depth Map and Kalman Smoothing
par: Wang, Dehao, et autres
Publié: (2025)
par: Wang, Dehao, et autres
Publié: (2025)
O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
par: Chen, Yuqing, et autres
Publié: (2025)
par: Chen, Yuqing, et autres
Publié: (2025)
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
par: Song, Baiyang, et autres
Publié: (2026)
par: Song, Baiyang, et autres
Publié: (2026)
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
par: Kwon, Minchan, et autres
Publié: (2026)
par: Kwon, Minchan, et autres
Publié: (2026)
RASA: Replace Anyone, Say Anything -- A Training-Free Framework for Audio-Driven and Universal Portrait Video Editing
par: Pan, Tianrui, et autres
Publié: (2025)
par: Pan, Tianrui, et autres
Publié: (2025)
TOTNet: Occlusion-Aware Temporal Tracking for Robust Ball Detection in Sports Videos
par: Xu, Hao, et autres
Publié: (2025)
par: Xu, Hao, et autres
Publié: (2025)
FOCUS: Efficient Keyframe Selection for Long Video Understanding
par: Zhu, Zirui, et autres
Publié: (2025)
par: Zhu, Zirui, et autres
Publié: (2025)
KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
par: Wang, Xingrui, et autres
Publié: (2025)
par: Wang, Xingrui, et autres
Publié: (2025)
Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction
par: Jia, Mingda, et autres
Publié: (2025)
par: Jia, Mingda, et autres
Publié: (2025)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
par: Liang, Hao, et autres
Publié: (2024)
par: Liang, Hao, et autres
Publié: (2024)
CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editing
par: Li, Yan, et autres
Publié: (2025)
par: Li, Yan, et autres
Publié: (2025)
Occlusion-Aware SORT: Observing Occlusion for Robust Multi-Object Tracking
par: Li, Chunjiang, et autres
Publié: (2026)
par: Li, Chunjiang, et autres
Publié: (2026)
Global Occlusion-Aware Transformer for Robust Stereo Matching
par: Liu, Zihua, et autres
Publié: (2023)
par: Liu, Zihua, et autres
Publié: (2023)
GaussianEditor: Editing 3D Gaussians Delicately with Text Instructions
par: Wang, Junjie, et autres
Publié: (2023)
par: Wang, Junjie, et autres
Publié: (2023)
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
par: Lin, Hangyu, et autres
Publié: (2026)
par: Lin, Hangyu, et autres
Publié: (2026)
Agentic Keyframe Search for Video Question Answering
par: Fan, Sunqi, et autres
Publié: (2025)
par: Fan, Sunqi, et autres
Publié: (2025)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
par: Wang, Shaoguang, et autres
Publié: (2026)
par: Wang, Shaoguang, et autres
Publié: (2026)
Large Model based Sequential Keyframe Extraction for Video Summarization
par: Tan, Kailong, et autres
Publié: (2024)
par: Tan, Kailong, et autres
Publié: (2024)
Masked Autoencoders are Robust Data Augmentors
par: Xu, Haohang, et autres
Publié: (2022)
par: Xu, Haohang, et autres
Publié: (2022)
Betrayed by Attention: A Simple yet Effective Approach for Self-supervised Video Object Segmentation
par: Ding, Shuangrui, et autres
Publié: (2023)
par: Ding, Shuangrui, et autres
Publié: (2023)
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
par: Zhang, Xinyao, et autres
Publié: (2026)
par: Zhang, Xinyao, et autres
Publié: (2026)
The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection
par: He, Qingdong, et autres
Publié: (2025)
par: He, Qingdong, et autres
Publié: (2025)
SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing
par: Liang, Sen, et autres
Publié: (2026)
par: Liang, Sen, et autres
Publié: (2026)
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
par: Xu, Chuanzhi, et autres
Publié: (2026)
par: Xu, Chuanzhi, et autres
Publié: (2026)
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
par: Zhang, Xian, et autres
Publié: (2025)
par: Zhang, Xian, et autres
Publié: (2025)
Adaptive Keyframe Sampling for Long Video Understanding
par: Tang, Xi, et autres
Publié: (2025)
par: Tang, Xi, et autres
Publié: (2025)
GenEraser: Generalizable Video Object Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver
par: Chen, Yuqing, et autres
Publié: (2026)
par: Chen, Yuqing, et autres
Publié: (2026)
From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding
par: Lin, Shih-Yao, et autres
Publié: (2025)
par: Lin, Shih-Yao, et autres
Publié: (2025)
OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance
par: Wang, Cong, et autres
Publié: (2026)
par: Wang, Cong, et autres
Publié: (2026)
Rethinking Occlusion in FER: A Semantic-Aware Perspective and Go Beyond
par: Zhai, Huiyu, et autres
Publié: (2025)
par: Zhai, Huiyu, et autres
Publié: (2025)
EditCast3D: Single-Frame-Guided 3D Editing with Video Propagation and View Selection
par: Qu, Huaizhi, et autres
Publié: (2025)
par: Qu, Huaizhi, et autres
Publié: (2025)
Documents similaires
-
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance
par: Xu, Haohang, et autres
Publié: (2026) -
Reasoning to Align: Implicit Reasoning in Diffusion Transformers for Video Editing
par: Li, Yan, et autres
Publié: (2026) -
AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection
par: Zhang, Shuheng, et autres
Publié: (2025) -
LoVoRA: Text-guided and Mask-free Video Object Removal and Addition with Learnable Object-aware Localization
par: Xiao, Zhihan, et autres
Publié: (2025) -
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
par: He, Haibin, et autres
Publié: (2026)