Don't Judge by the Look: Towards Motion Coherent Video Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yitian, Bai, Yue, Wang, Huan, Wang, Yizhou, Fu, Yun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accessing Vision Foundation Models via ImageNet-1K
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Don't Play Favorites: Minority Guidance for Diffusion Models
by: Um, Soobin, et al.
Published: (2023)
by: Um, Soobin, et al.
Published: (2023)
REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder
by: Zhang, Yitian, et al.
Published: (2025)
by: Zhang, Yitian, et al.
Published: (2025)
MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching
by: Wu, Yen-Siang, et al.
Published: (2025)
by: Wu, Yen-Siang, et al.
Published: (2025)
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization
by: Choudhury, Rohan, et al.
Published: (2024)
by: Choudhury, Rohan, et al.
Published: (2024)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
by: Yu, Zhuoran, et al.
Published: (2023)
by: Yu, Zhuoran, et al.
Published: (2023)
Don't Get Me Wrong: How to Apply Deep Visual Interpretations to Time Series
by: Loeffler, Christoffer, et al.
Published: (2022)
by: Loeffler, Christoffer, et al.
Published: (2022)
Motion meets Attention: Video Motion Prompts
by: Chen, Qixiang, et al.
Published: (2024)
by: Chen, Qixiang, et al.
Published: (2024)
Distorted or Fabricated? A Survey on Hallucination in Video LLMs
by: Huang, Yiyang, et al.
Published: (2026)
by: Huang, Yiyang, et al.
Published: (2026)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
by: Parmar, Mihir, et al.
Published: (2022)
by: Parmar, Mihir, et al.
Published: (2022)
Don't Collapse Your Features: Why CenterLoss Hurts OOD Detection and Multi-Scale Mahalanobis Wins
by: Ray, Rahul D
Published: (2026)
by: Ray, Rahul D
Published: (2026)
A Closer Look at Multimodal Representation Collapse
by: Chaudhuri, Abhra, et al.
Published: (2025)
by: Chaudhuri, Abhra, et al.
Published: (2025)
Towards Understanding Camera Motions in Any Video
by: Lin, Zhiqiu, et al.
Published: (2025)
by: Lin, Zhiqiu, et al.
Published: (2025)
When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
Towards Principled Representation Learning from Videos for Reinforcement Learning
by: Misra, Dipendra, et al.
Published: (2024)
by: Misra, Dipendra, et al.
Published: (2024)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
by: Wang, Zhouxia, et al.
Published: (2023)
by: Wang, Zhouxia, et al.
Published: (2023)
Shadows Don't Lie and Lines Can't Bend! Generative Models don't know Projective Geometry...for now
by: Sarkar, Ayush, et al.
Published: (2023)
by: Sarkar, Ayush, et al.
Published: (2023)
ReactDance: Hierarchical Representation for High-Fidelity and Coherent Long-Form Reactive Dance Generation
by: Lin, Jingzhong, et al.
Published: (2025)
by: Lin, Jingzhong, et al.
Published: (2025)
Look Through Masks: Towards Masked Face Recognition with De-Occlusion Distillation
by: Li, Chenyu, et al.
Published: (2024)
by: Li, Chenyu, et al.
Published: (2024)
Text-to-3D Gaussian Splatting with Physics-Grounded Motion Generation
by: Wang, Wenqing, et al.
Published: (2024)
by: Wang, Wenqing, et al.
Published: (2024)
MTS-DMAE: Dual-Masked Autoencoder for Unsupervised Multivariate Time Series Representation Learning
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
by: Huang, Yiyang, et al.
Published: (2025)
by: Huang, Yiyang, et al.
Published: (2025)
ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
by: Ai, Jiaxin, et al.
Published: (2025)
by: Ai, Jiaxin, et al.
Published: (2025)
Spectral Motion Alignment for Video Motion Transfer using Diffusion Models
by: Park, Geon Yeong, et al.
Published: (2024)
by: Park, Geon Yeong, et al.
Published: (2024)
Video Motion Transfer with Diffusion Transformers
by: Pondaven, Alexander, et al.
Published: (2024)
by: Pondaven, Alexander, et al.
Published: (2024)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
by: Li, Quanhao, et al.
Published: (2025)
by: Li, Quanhao, et al.
Published: (2025)
BrainDINO: A Brain MRI Foundation Model for Generalizable Clinical Representation Learning
by: Wu, Yizhou, et al.
Published: (2026)
by: Wu, Yizhou, et al.
Published: (2026)
Feature Representation Transferring to Lightweight Models via Perception Coherence
by: Nguyen, Hai-Vy, et al.
Published: (2025)
by: Nguyen, Hai-Vy, et al.
Published: (2025)
Towards Precise Scaling Laws for Video Diffusion Transformers
by: Yin, Yuanyang, et al.
Published: (2024)
by: Yin, Yuanyang, et al.
Published: (2024)
Enhanced Event-Based Video Reconstruction with Motion Compensation
by: Liu, Siying, et al.
Published: (2024)
by: Liu, Siying, et al.
Published: (2024)
VideoPhy: Evaluating Physical Commonsense for Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
From Image to Video: An Empirical Study of Diffusion Representations
by: Vélez, Pedro, et al.
Published: (2025)
by: Vélez, Pedro, et al.
Published: (2025)
MotionV2V: Editing Motion in a Video
by: Burgert, Ryan, et al.
Published: (2025)
by: Burgert, Ryan, et al.
Published: (2025)
MotionCraft: Physics-based Zero-Shot Video Generation
by: Aira, Luca Savant, et al.
Published: (2024)
by: Aira, Luca Savant, et al.
Published: (2024)
Stag-1: Towards Realistic 4D Driving Simulation with Video Generation Model
by: Wang, Lening, et al.
Published: (2024)
by: Wang, Lening, et al.
Published: (2024)
VQ-Style: Disentangling Style and Content in Motion with Residual Quantized Representations
by: Zargarbashi, Fatemeh, et al.
Published: (2026)
by: Zargarbashi, Fatemeh, et al.
Published: (2026)
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
by: Lee, Hosu, et al.
Published: (2024)
by: Lee, Hosu, et al.
Published: (2024)
Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning
by: Kim, Hyeonjin, et al.
Published: (2026)
by: Kim, Hyeonjin, et al.
Published: (2026)
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
by: Li, Quanhao, et al.
Published: (2026)
by: Li, Quanhao, et al.
Published: (2026)
X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning
by: Wang, Yunzhe, et al.
Published: (2025)
by: Wang, Yunzhe, et al.
Published: (2025)
Similar Items
-
Accessing Vision Foundation Models via ImageNet-1K
by: Zhang, Yitian, et al.
Published: (2024) -
Don't Play Favorites: Minority Guidance for Diffusion Models
by: Um, Soobin, et al.
Published: (2023) -
REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder
by: Zhang, Yitian, et al.
Published: (2025) -
MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching
by: Wu, Yen-Siang, et al.
Published: (2025) -
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization
by: Choudhury, Rohan, et al.
Published: (2024)