Gespeichert in:
| Hauptverfasser: | Liu, Lin, Xiao, Zhihan, Xu, Haohang, Cong, Rong, Zhang, Zhibo, Zhang, Xiaopeng, Tian, Qi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2605.23192 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance
von: Xu, Haohang, et al.
Veröffentlicht: (2026)
von: Xu, Haohang, et al.
Veröffentlicht: (2026)
Reasoning to Align: Implicit Reasoning in Diffusion Transformers for Video Editing
von: Li, Yan, et al.
Veröffentlicht: (2026)
von: Li, Yan, et al.
Veröffentlicht: (2026)
LoVoRA: Text-guided and Mask-free Video Object Removal and Addition with Learnable Object-aware Localization
von: Xiao, Zhihan, et al.
Veröffentlicht: (2025)
von: Xiao, Zhihan, et al.
Veröffentlicht: (2025)
AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection
von: Zhang, Shuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Shuheng, et al.
Veröffentlicht: (2025)
FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions
von: Zhao, Peisen, et al.
Veröffentlicht: (2026)
von: Zhao, Peisen, et al.
Veröffentlicht: (2026)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2026)
von: He, Haibin, et al.
Veröffentlicht: (2026)
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
von: Nie, Ming, et al.
Veröffentlicht: (2025)
von: Nie, Ming, et al.
Veröffentlicht: (2025)
O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
von: Chen, Yuqing, et al.
Veröffentlicht: (2025)
von: Chen, Yuqing, et al.
Veröffentlicht: (2025)
Robust Video-Based Pothole Detection and Area Estimation for Intelligent Vehicles with Depth Map and Kalman Smoothing
von: Wang, Dehao, et al.
Veröffentlicht: (2025)
von: Wang, Dehao, et al.
Veröffentlicht: (2025)
KeyframeFace: Language-Driven Facial Animation via Semantic Keyframes
von: Wu, Jingchao, et al.
Veröffentlicht: (2025)
von: Wu, Jingchao, et al.
Veröffentlicht: (2025)
KS-APR: Keyframe Selection for Robust Absolute Pose Regression
von: Liu, Changkun, et al.
Veröffentlicht: (2023)
von: Liu, Changkun, et al.
Veröffentlicht: (2023)
RASA: Replace Anyone, Say Anything -- A Training-Free Framework for Audio-Driven and Universal Portrait Video Editing
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
FOCUS: Efficient Keyframe Selection for Long Video Understanding
von: Zhu, Zirui, et al.
Veröffentlicht: (2025)
von: Zhu, Zirui, et al.
Veröffentlicht: (2025)
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
von: Kwon, Minchan, et al.
Veröffentlicht: (2026)
von: Kwon, Minchan, et al.
Veröffentlicht: (2026)
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
von: Song, Baiyang, et al.
Veröffentlicht: (2026)
von: Song, Baiyang, et al.
Veröffentlicht: (2026)
CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editing
von: Li, Yan, et al.
Veröffentlicht: (2025)
von: Li, Yan, et al.
Veröffentlicht: (2025)
Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction
von: Jia, Mingda, et al.
Veröffentlicht: (2025)
von: Jia, Mingda, et al.
Veröffentlicht: (2025)
TOTNet: Occlusion-Aware Temporal Tracking for Robust Ball Detection in Sports Videos
von: Xu, Hao, et al.
Veröffentlicht: (2025)
von: Xu, Hao, et al.
Veröffentlicht: (2025)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
GaussianEditor: Editing 3D Gaussians Delicately with Text Instructions
von: Wang, Junjie, et al.
Veröffentlicht: (2023)
von: Wang, Junjie, et al.
Veröffentlicht: (2023)
Occlusion-Aware SORT: Observing Occlusion for Robust Multi-Object Tracking
von: Li, Chunjiang, et al.
Veröffentlicht: (2026)
von: Li, Chunjiang, et al.
Veröffentlicht: (2026)
Global Occlusion-Aware Transformer for Robust Stereo Matching
von: Liu, Zihua, et al.
Veröffentlicht: (2023)
von: Liu, Zihua, et al.
Veröffentlicht: (2023)
Masked Autoencoders are Robust Data Augmentors
von: Xu, Haohang, et al.
Veröffentlicht: (2022)
von: Xu, Haohang, et al.
Veröffentlicht: (2022)
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
von: Xu, Chuanzhi, et al.
Veröffentlicht: (2026)
von: Xu, Chuanzhi, et al.
Veröffentlicht: (2026)
Betrayed by Attention: A Simple yet Effective Approach for Self-supervised Video Object Segmentation
von: Ding, Shuangrui, et al.
Veröffentlicht: (2023)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2023)
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
von: Lin, Hangyu, et al.
Veröffentlicht: (2026)
von: Lin, Hangyu, et al.
Veröffentlicht: (2026)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
Agentic Keyframe Search for Video Question Answering
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
Adaptive Keyframe Sampling for Long Video Understanding
von: Tang, Xi, et al.
Veröffentlicht: (2025)
von: Tang, Xi, et al.
Veröffentlicht: (2025)
SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing
von: Liang, Sen, et al.
Veröffentlicht: (2026)
von: Liang, Sen, et al.
Veröffentlicht: (2026)
GenEraser: Generalizable Video Object Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver
von: Chen, Yuqing, et al.
Veröffentlicht: (2026)
von: Chen, Yuqing, et al.
Veröffentlicht: (2026)
Large Model based Sequential Keyframe Extraction for Video Summarization
von: Tan, Kailong, et al.
Veröffentlicht: (2024)
von: Tan, Kailong, et al.
Veröffentlicht: (2024)
OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance
von: Wang, Cong, et al.
Veröffentlicht: (2026)
von: Wang, Cong, et al.
Veröffentlicht: (2026)
The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection
von: He, Qingdong, et al.
Veröffentlicht: (2025)
von: He, Qingdong, et al.
Veröffentlicht: (2025)
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
von: Zhang, Xinyao, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyao, et al.
Veröffentlicht: (2026)
AlignZeg: Mitigating Objective Misalignment for Zero-shot Semantic Segmentation
von: Ge, Jiannan, et al.
Veröffentlicht: (2024)
von: Ge, Jiannan, et al.
Veröffentlicht: (2024)
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
von: Zhang, Xian, et al.
Veröffentlicht: (2025)
von: Zhang, Xian, et al.
Veröffentlicht: (2025)
From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding
von: Lin, Shih-Yao, et al.
Veröffentlicht: (2025)
von: Lin, Shih-Yao, et al.
Veröffentlicht: (2025)
Runge-Kutta Approximation and Decoupled Attention for Rectified Flow Inversion and Semantic Editing
von: Chen, Weiming, et al.
Veröffentlicht: (2025)
von: Chen, Weiming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance
von: Xu, Haohang, et al.
Veröffentlicht: (2026) -
Reasoning to Align: Implicit Reasoning in Diffusion Transformers for Video Editing
von: Li, Yan, et al.
Veröffentlicht: (2026) -
LoVoRA: Text-guided and Mask-free Video Object Removal and Addition with Learnable Object-aware Localization
von: Xiao, Zhihan, et al.
Veröffentlicht: (2025) -
AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection
von: Zhang, Shuheng, et al.
Veröffentlicht: (2025) -
FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions
von: Zhao, Peisen, et al.
Veröffentlicht: (2026)