Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Ruyang, Sun, Shangkun, Tang, Haoran, Li, Ge, Gao, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
von: Sun, Shangkun, et al.
Veröffentlicht: (2024)
von: Sun, Shangkun, et al.
Veröffentlicht: (2024)
Content-Rich AIGC Video Quality Assessment via Intricate Text Alignment and Motion-Aware Consistency
von: Sun, Shangkun, et al.
Veröffentlicht: (2025)
von: Sun, Shangkun, et al.
Veröffentlicht: (2025)
Video Spatial Reasoning with Object-Centric 3D Rollout
von: Tang, Haoran, et al.
Veröffentlicht: (2025)
von: Tang, Haoran, et al.
Veröffentlicht: (2025)
ST-LLM: Large Language Models Are Effective Temporal Learners
von: Liu, Ruyang, et al.
Veröffentlicht: (2024)
von: Liu, Ruyang, et al.
Veröffentlicht: (2024)
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2026)
von: Cao, Meng, et al.
Veröffentlicht: (2026)
LumosFlow: Motion-Guided Long Video Generation
von: Chen, Jiahao, et al.
Veröffentlicht: (2025)
von: Chen, Jiahao, et al.
Veröffentlicht: (2025)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
von: Liu, Ruyang, et al.
Veröffentlicht: (2023)
von: Liu, Ruyang, et al.
Veröffentlicht: (2023)
Exploring AIGC Video Quality: A Focus on Visual Harmony, Video-Text Consistency and Domain Distribution Gap
von: Qu, Bowen, et al.
Veröffentlicht: (2024)
von: Qu, Bowen, et al.
Veröffentlicht: (2024)
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
von: Nguyen, Hong, et al.
Veröffentlicht: (2025)
von: Nguyen, Hong, et al.
Veröffentlicht: (2025)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
SplatFlow: Learning Multi-frame Optical Flow via Splatting
von: Wang, Bo, et al.
Veröffentlicht: (2023)
von: Wang, Bo, et al.
Veröffentlicht: (2023)
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
MUSE: Mamba is Efficient Multi-scale Learner for Text-video Retrieval
von: Tang, Haoran, et al.
Veröffentlicht: (2024)
von: Tang, Haoran, et al.
Veröffentlicht: (2024)
FlowMotion: Training-Free Flow Guidance for Video Motion Transfer
von: Wang, Zhen, et al.
Veröffentlicht: (2026)
von: Wang, Zhen, et al.
Veröffentlicht: (2026)
Learning Long-form Video Prior via Generative Pre-Training
von: Xie, Jinheng, et al.
Veröffentlicht: (2024)
von: Xie, Jinheng, et al.
Veröffentlicht: (2024)
VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment
von: Sun, Shangkun, et al.
Veröffentlicht: (2024)
von: Sun, Shangkun, et al.
Veröffentlicht: (2024)
Training-Free Video Editing via Optical Flow-Enhanced Score Distillation
von: Zhu, Lianghan, et al.
Veröffentlicht: (2024)
von: Zhu, Lianghan, et al.
Veröffentlicht: (2024)
DeepSeek-OCR 2: Visual Causal Flow
von: Wei, Haoran, et al.
Veröffentlicht: (2026)
von: Wei, Haoran, et al.
Veröffentlicht: (2026)
Video Token Merging for Long-form Video Understanding
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
Motion Semantics Guided Normalizing Flow for Privacy-Preserving Video Anomaly Detection
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
Aligning Latent Spaces with Flow Priors
von: Li, Yizhuo, et al.
Veröffentlicht: (2025)
von: Li, Yizhuo, et al.
Veröffentlicht: (2025)
MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation
von: Lei, Guojun, et al.
Veröffentlicht: (2025)
von: Lei, Guojun, et al.
Veröffentlicht: (2025)
Unsupervised 4D Cardiac Motion Tracking with Spatiotemporal Optical Flow Networks
von: Teng, Long, et al.
Veröffentlicht: (2024)
von: Teng, Long, et al.
Veröffentlicht: (2024)
Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
OpFlowTalker: Realistic and Natural Talking Face Generation via Optical Flow Guidance
von: Ge, Shuheng, et al.
Veröffentlicht: (2024)
von: Ge, Shuheng, et al.
Veröffentlicht: (2024)
Spacewalk-18: A Benchmark for Multimodal and Long-form Procedural Video Understanding in Novel Domains
von: Tang, Zitian, et al.
Veröffentlicht: (2023)
von: Tang, Zitian, et al.
Veröffentlicht: (2023)
VideoGen-Eval: Agent-based System for Video Generation Evaluation
von: Yang, Yuhang, et al.
Veröffentlicht: (2025)
von: Yang, Yuhang, et al.
Veröffentlicht: (2025)
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
von: Shen, Fei, et al.
Veröffentlicht: (2025)
von: Shen, Fei, et al.
Veröffentlicht: (2025)
iMOVE: Instance-Motion-Aware Video Understanding
von: Li, Jiaze, et al.
Veröffentlicht: (2025)
von: Li, Jiaze, et al.
Veröffentlicht: (2025)
FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis
von: Liang, Feng, et al.
Veröffentlicht: (2023)
von: Liang, Feng, et al.
Veröffentlicht: (2023)
ScaleFlow++: Robust and Accurate Estimation of 3D Motion from Video
von: Ling, Han, et al.
Veröffentlicht: (2024)
von: Ling, Han, et al.
Veröffentlicht: (2024)
FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion Bases
von: Poggi, Matteo, et al.
Veröffentlicht: (2025)
von: Poggi, Matteo, et al.
Veröffentlicht: (2025)
CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding
von: Patel, Shrenik, et al.
Veröffentlicht: (2025)
von: Patel, Shrenik, et al.
Veröffentlicht: (2025)
ScaleFlow++: Robust and Accurate Estimation of 3D Motion from Video
von: Ling, Han, et al.
Veröffentlicht: (2024)
von: Ling, Han, et al.
Veröffentlicht: (2024)
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation
von: Shaulov, Ariel, et al.
Veröffentlicht: (2025)
von: Shaulov, Ariel, et al.
Veröffentlicht: (2025)
AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection
von: Zhang, Shuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Shuheng, et al.
Veröffentlicht: (2025)
Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
HMAFlow: Learning More Accurate Optical Flow via Hierarchical Motion Field Alignment
von: Ma, Dianbo, et al.
Veröffentlicht: (2024)
von: Ma, Dianbo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
von: Sun, Shangkun, et al.
Veröffentlicht: (2024) -
Content-Rich AIGC Video Quality Assessment via Intricate Text Alignment and Motion-Aware Consistency
von: Sun, Shangkun, et al.
Veröffentlicht: (2025) -
Video Spatial Reasoning with Object-Centric 3D Rollout
von: Tang, Haoran, et al.
Veröffentlicht: (2025) -
ST-LLM: Large Language Models Are Effective Temporal Learners
von: Liu, Ruyang, et al.
Veröffentlicht: (2024) -
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2026)