MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bhowmik, Aritra, Korzhenkov, Denis, Snoek, Cees G. M., Habibian, Amirhossein, Ghafoorian, Mohsen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
by: Ghafoorian, Mohsen, et al.
Published: (2025)
by: Ghafoorian, Mohsen, et al.
Published: (2025)
PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
by: Korzhenkov, Denis, et al.
Published: (2026)
by: Korzhenkov, Denis, et al.
Published: (2026)
ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers
by: Ghafoorian, Mohsen, et al.
Published: (2026)
by: Ghafoorian, Mohsen, et al.
Published: (2026)
Mobile Video Diffusion
by: Yahia, Haitam Ben, et al.
Published: (2024)
by: Yahia, Haitam Ben, et al.
Published: (2024)
Neodragon: Mobile Video Generation using Diffusion Transformer
by: Karnewar, Animesh, et al.
Published: (2025)
by: Karnewar, Animesh, et al.
Published: (2025)
LocoMotion: Learning Motion-Focused Video-Language Representations
by: Doughty, Hazel, et al.
Published: (2024)
by: Doughty, Hazel, et al.
Published: (2024)
Union-over-Intersections: Object Detection beyond Winner-Takes-All
by: Bhowmik, Aritra, et al.
Published: (2023)
by: Bhowmik, Aritra, et al.
Published: (2023)
Object-Centric Diffusion for Efficient Video Editing
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
MoViE: Mobile Diffusion for Video Editing
by: Karjauv, Adil, et al.
Published: (2024)
by: Karjauv, Adil, et al.
Published: (2024)
FastCAD: Real-Time CAD Retrieval and Alignment from Scans and Videos
by: Langer, Florian, et al.
Published: (2024)
by: Langer, Florian, et al.
Published: (2024)
TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning
by: Bhowmik, Aritra, et al.
Published: (2024)
by: Bhowmik, Aritra, et al.
Published: (2024)
MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
by: Salehi, Mohammadreza, et al.
Published: (2025)
by: Salehi, Mohammadreza, et al.
Published: (2025)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
by: Liu, Huabin, et al.
Published: (2025)
by: Liu, Huabin, et al.
Published: (2025)
Enhancing Novel View Synthesis via Geometry Grounded Set Diffusion
by: Zanjani, Farhad G., et al.
Published: (2026)
by: Zanjani, Farhad G., et al.
Published: (2026)
Low-Resource Vision Challenges for Foundation Models
by: Zhang, Yunhua, et al.
Published: (2024)
by: Zhang, Yunhua, et al.
Published: (2024)
SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning
by: Thoker, Fida Mohammad, et al.
Published: (2025)
by: Thoker, Fida Mohammad, et al.
Published: (2025)
Controllable 3D Placement of Objects with Scene-Aware Diffusion Models
by: Omran, Mohamed, et al.
Published: (2025)
by: Omran, Mohamed, et al.
Published: (2025)
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
by: Chen, Yanlong, et al.
Published: (2026)
by: Chen, Yanlong, et al.
Published: (2026)
SIGMA: Sinkhorn-Guided Masked Video Modeling
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Lost in Time: A New Temporal Benchmark for VideoLLMs
by: Cores, Daniel, et al.
Published: (2024)
by: Cores, Daniel, et al.
Published: (2024)
Beyond Coarse-Grained Matching in Video-Text Retrieval
by: Chen, Aozhu, et al.
Published: (2024)
by: Chen, Aozhu, et al.
Published: (2024)
Multi-Scale Local Speculative Decoding for Image Generation
by: Peruzzo, Elia, et al.
Published: (2026)
by: Peruzzo, Elia, et al.
Published: (2026)
Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
by: Wang, Ziqi, et al.
Published: (2025)
by: Wang, Ziqi, et al.
Published: (2025)
IPO: Interpretable Prompt Optimization for Vision-Language Models
by: Du, Yingjun, et al.
Published: (2024)
by: Du, Yingjun, et al.
Published: (2024)
Scene-Aware Location Modeling for Data Augmentation in Automotive Object Detection
by: Petersen, Jens, et al.
Published: (2025)
by: Petersen, Jens, et al.
Published: (2025)
OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference Optimization
by: Zhang, Jiacheng, et al.
Published: (2024)
by: Zhang, Jiacheng, et al.
Published: (2024)
Dual Guidance Semi-Supervised Action Detection
by: Singh, Ankit, et al.
Published: (2025)
by: Singh, Ankit, et al.
Published: (2025)
SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail
by: Du, Yingjun, et al.
Published: (2023)
by: Du, Yingjun, et al.
Published: (2023)
SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation
by: Nguyen, Duy-Kien, et al.
Published: (2023)
by: Nguyen, Duy-Kien, et al.
Published: (2023)
Clockwork Diffusion: Efficient Generation With Model-Step Distillation
by: Habibian, Amirhossein, et al.
Published: (2023)
by: Habibian, Amirhossein, et al.
Published: (2023)
Elastic ViTs from Pretrained Models without Retraining
by: Simoncini, Walter, et al.
Published: (2025)
by: Simoncini, Walter, et al.
Published: (2025)
Segment Any 3D-Part in a Scene from a Sentence
by: Wu, Hongyu, et al.
Published: (2025)
by: Wu, Hongyu, et al.
Published: (2025)
PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs
by: Dorkenwald, Michael, et al.
Published: (2024)
by: Dorkenwald, Michael, et al.
Published: (2024)
RefAlign: Representation Alignment for Reference-to-Video Generation
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
Prompt Diffusion Robustifies Any-Modality Prompt Learning
by: Du, Yingjun, et al.
Published: (2024)
by: Du, Yingjun, et al.
Published: (2024)
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
by: Sun, Wenfang, et al.
Published: (2026)
by: Sun, Wenfang, et al.
Published: (2026)
Training-Free Semantic Segmentation via LLM-Supervision
by: Sun, Wenfang, et al.
Published: (2024)
by: Sun, Wenfang, et al.
Published: (2024)
QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain
by: Sun, Wenfang, et al.
Published: (2024)
by: Sun, Wenfang, et al.
Published: (2024)
Redefining Normal: A Novel Object-Level Approach for Multi-Object Novelty Detection
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
MoVideo: Motion-Aware Video Generation with Diffusion Models
by: Liang, Jingyun, et al.
Published: (2023)
by: Liang, Jingyun, et al.
Published: (2023)
Similar Items
-
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
by: Ghafoorian, Mohsen, et al.
Published: (2025) -
PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
by: Korzhenkov, Denis, et al.
Published: (2026) -
ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers
by: Ghafoorian, Mohsen, et al.
Published: (2026) -
Mobile Video Diffusion
by: Yahia, Haitam Ben, et al.
Published: (2024) -
Neodragon: Mobile Video Generation using Diffusion Transformer
by: Karnewar, Animesh, et al.
Published: (2025)