FluencyVE: Marrying Temporal-Aware Mamba with Bypass Attention for Video Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Cai, Mingshu, Li, Yixuan, Yoshie, Osamu, Ieiri, Yuya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KGEdit: Ambiguity-Aware Knowledge Graphs for Training-Free Precise Video Generation and Editing
by: Cai, Mingshu, et al.
Published: (2026)
by: Cai, Mingshu, et al.
Published: (2026)
Multi-Attribute guided Thermal Face Image Translation based on Latent Diffusion Model
by: Cai, Mingshu, et al.
Published: (2025)
by: Cai, Mingshu, et al.
Published: (2025)
Where, What, Why: Toward Explainable 3D-GS Watermarking
by: Cai, Mingshu, et al.
Published: (2026)
by: Cai, Mingshu, et al.
Published: (2026)
DM$^3$Net: Dual-Camera Super-Resolution via Domain Modulation and Multi-scale Matching
by: Guan, Cong, et al.
Published: (2025)
by: Guan, Cong, et al.
Published: (2025)
Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation
by: He, Minggui, et al.
Published: (2026)
by: He, Minggui, et al.
Published: (2026)
DreamVE: Unified Instruction-based Image and Video Editing
by: Xia, Bin, et al.
Published: (2025)
by: Xia, Bin, et al.
Published: (2025)
TK-Mamba: Marrying KAN With Mamba for Text-Driven 3D Medical Image Segmentation
by: Yang, Haoyu, et al.
Published: (2025)
by: Yang, Haoyu, et al.
Published: (2025)
CLIP-driven rain perception: Adaptive deraining with pattern-aware network routing and mask-guided cross-attention
by: Guan, Cong, et al.
Published: (2025)
by: Guan, Cong, et al.
Published: (2025)
Mamba-YOLO-World: Marrying YOLO-World with Mamba for Open-Vocabulary Detection
by: Wang, Haoxuan, et al.
Published: (2024)
by: Wang, Haoxuan, et al.
Published: (2024)
Temporally Consistent Object Editing in Videos using Extended Attention
by: Zamani, AmirHossein, et al.
Published: (2024)
by: Zamani, AmirHossein, et al.
Published: (2024)
LFSamba: Marry SAM with Mamba for Light Field Salient Object Detection
by: Liu, Zhengyi, et al.
Published: (2024)
by: Liu, Zhengyi, et al.
Published: (2024)
P-Mamba: Marrying Perona Malik Diffusion with Mamba for Efficient Pediatric Echocardiographic Left Ventricular Segmentation
by: Ye, Zi, et al.
Published: (2024)
by: Ye, Zi, et al.
Published: (2024)
Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion
by: He, Yangfan, et al.
Published: (2025)
by: He, Yangfan, et al.
Published: (2025)
VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment
by: Sun, Shangkun, et al.
Published: (2024)
by: Sun, Shangkun, et al.
Published: (2024)
FiVE: A Fine-grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow Models
by: Li, Minghan, et al.
Published: (2025)
by: Li, Minghan, et al.
Published: (2025)
Matten: Video Generation with Mamba-Attention
by: Gao, Yu, et al.
Published: (2024)
by: Gao, Yu, et al.
Published: (2024)
FlexEdit: Marrying Free-Shape Masks to VLLM for Flexible Image Editing
by: Yuan, Tianshuo, et al.
Published: (2024)
by: Yuan, Tianshuo, et al.
Published: (2024)
TV-LiVE: Training-Free, Text-Guided Video Editing via Layer Informed Vitality Exploitation
by: Kim, Min-Jung, et al.
Published: (2025)
by: Kim, Min-Jung, et al.
Published: (2025)
OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
by: He, Haoyang, et al.
Published: (2025)
by: He, Haoyang, et al.
Published: (2025)
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
by: An, Joungbin, et al.
Published: (2025)
by: An, Joungbin, et al.
Published: (2025)
Relit-LiVE: Relight Video by Jointly Learning Environment Video
by: Xiao, Weiqing, et al.
Published: (2026)
by: Xiao, Weiqing, et al.
Published: (2026)
Ada-VE: Training-Free Consistent Video Editing Using Adaptive Motion Prior
by: Mahmud, Tanvir, et al.
Published: (2024)
by: Mahmud, Tanvir, et al.
Published: (2024)
FlowBypass: Rectified Flow Trajectory Bypass for Training-Free Image Editing
by: Han, Menglin, et al.
Published: (2026)
by: Han, Menglin, et al.
Published: (2026)
EPIC Fields: Marrying 3D Geometry and Video Understanding
by: Tschernezki, Vadim, et al.
Published: (2023)
by: Tschernezki, Vadim, et al.
Published: (2023)
Gamba: Marry Gaussian Splatting with Mamba for single view 3D reconstruction
by: Shen, Qiuhong, et al.
Published: (2024)
by: Shen, Qiuhong, et al.
Published: (2024)
FADE: Frequency-Aware Diffusion Model Factorization for Video Editing
by: Zhu, Yixuan, et al.
Published: (2025)
by: Zhu, Yixuan, et al.
Published: (2025)
MiVE: Multiscale Vision-language features for reference-guided video Editing
by: Wang, Tong, et al.
Published: (2026)
by: Wang, Tong, et al.
Published: (2026)
PMQ-VE: Progressive Multi-Frame Quantization for Video Enhancement
by: Feng, ZhanFeng, et al.
Published: (2025)
by: Feng, ZhanFeng, et al.
Published: (2025)
VideoCoF: Unified Video Editing with Temporal Reasoner
by: Yang, Xiangpeng, et al.
Published: (2025)
by: Yang, Xiangpeng, et al.
Published: (2025)
BFA: Best-Feature-Aware Fusion for Multi-View Fine-grained Manipulation
by: Lan, Zihan, et al.
Published: (2025)
by: Lan, Zihan, et al.
Published: (2025)
VideoMamba: Spatio-Temporal Selective State Space Model
by: Park, Jinyoung, et al.
Published: (2024)
by: Park, Jinyoung, et al.
Published: (2024)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
by: Cai, Jianfeng, et al.
Published: (2025)
by: Cai, Jianfeng, et al.
Published: (2025)
STNMamba: Mamba-based Spatial-Temporal Normality Learning for Video Anomaly Detection
by: Li, Zhangxun, et al.
Published: (2024)
by: Li, Zhangxun, et al.
Published: (2024)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
by: Zhang, David Junhao, et al.
Published: (2023)
by: Zhang, David Junhao, et al.
Published: (2023)
Predictive Temporal Attention on Event-based Video Stream for Energy-efficient Situation Awareness
by: Bu, Yiming, et al.
Published: (2024)
by: Bu, Yiming, et al.
Published: (2024)
SMTrack: State-Aware Mamba for Efficient Temporal Modeling in Visual Tracking
by: Ma, Yinchao, et al.
Published: (2026)
by: Ma, Yinchao, et al.
Published: (2026)
DiVE: DiT-based Video Generation with Enhanced Control
by: Jiang, Junpeng, et al.
Published: (2024)
by: Jiang, Junpeng, et al.
Published: (2024)
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Qffusion: Controllable Portrait Video Editing via Quadrant-Grid Attention Learning
by: Li, Maomao, et al.
Published: (2025)
by: Li, Maomao, et al.
Published: (2025)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
by: Cai, Lingling, et al.
Published: (2024)
by: Cai, Lingling, et al.
Published: (2024)
Similar Items
-
KGEdit: Ambiguity-Aware Knowledge Graphs for Training-Free Precise Video Generation and Editing
by: Cai, Mingshu, et al.
Published: (2026) -
Multi-Attribute guided Thermal Face Image Translation based on Latent Diffusion Model
by: Cai, Mingshu, et al.
Published: (2025) -
Where, What, Why: Toward Explainable 3D-GS Watermarking
by: Cai, Mingshu, et al.
Published: (2026) -
DM$^3$Net: Dual-Camera Super-Resolution via Domain Modulation and Multi-scale Matching
by: Guan, Cong, et al.
Published: (2025) -
Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation
by: He, Minggui, et al.
Published: (2026)