MotiF: Making Text Count in Image Animation with Motion Focal Loss
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Shijie, Azadi, Samaneh, Girdhar, Rohit, Rambhatla, Saketh, Sun, Chen, Yin, Xi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
by: Girdhar, Rohit, et al.
Published: (2023)
by: Girdhar, Rohit, et al.
Published: (2023)
Diffusion Autoencoders are Scalable Image Tokenizers
by: Chen, Yinbo, et al.
Published: (2025)
by: Chen, Yinbo, et al.
Published: (2025)
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
Generating Multi-Image Synthetic Data for Text-to-Image Customization
by: Kumari, Nupur, et al.
Published: (2025)
by: Kumari, Nupur, et al.
Published: (2025)
SelfEval: Leveraging the discriminative nature of generative models for evaluation
by: Rambhatla, Sai Saketh, et al.
Published: (2023)
by: Rambhatla, Sai Saketh, et al.
Published: (2023)
MotiMotion: Motion-Controlled Video Generation with Visual Reasoning
by: Hsin-Ying, Lee, et al.
Published: (2026)
by: Hsin-Ying, Lee, et al.
Published: (2026)
MotiMem: Motion-Aware Approximate Memory for Energy-Efficient Neural Perception in Autonomous Vehicles
by: Que, Haohua, et al.
Published: (2026)
by: Que, Haohua, et al.
Published: (2026)
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2025)
by: Kumar, Pulkit, et al.
Published: (2025)
Trajectory-aligned Space-time Tokens for Few-shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2024)
by: Kumar, Pulkit, et al.
Published: (2024)
Animate-X: Universal Character Image Animation with Enhanced Motion Representation
by: Tan, Shuai, et al.
Published: (2024)
by: Tan, Shuai, et al.
Published: (2024)
LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Text-Guided Image Editing
by: Soni, Achint, et al.
Published: (2025)
by: Soni, Achint, et al.
Published: (2025)
SketchAnimator: Animate Sketch via Motion Customization of Text-to-Video Diffusion Models
by: Yang, Ruolin, et al.
Published: (2025)
by: Yang, Ruolin, et al.
Published: (2025)
FocalCount: Towards Class-Count Imbalance in Class-Agnostic Counting
by: Zhu, Huilin, et al.
Published: (2025)
by: Zhu, Huilin, et al.
Published: (2025)
UVIS: Unsupervised Video Instance Segmentation
by: Huang, Shuaiyi, et al.
Published: (2024)
by: Huang, Shuaiyi, et al.
Published: (2024)
MVPortrait: Text-Guided Motion and Emotion Control for Multi-view Vivid Portrait Animation
by: Lin, Yukang, et al.
Published: (2025)
by: Lin, Yukang, et al.
Published: (2025)
Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
by: Ma, Xin, et al.
Published: (2025)
by: Ma, Xin, et al.
Published: (2025)
Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models
by: Ma, Xin, et al.
Published: (2024)
by: Ma, Xin, et al.
Published: (2024)
Animate-X++: Universal Character Image Animation with Dynamic Backgrounds
by: Tan, Shuai, et al.
Published: (2025)
by: Tan, Shuai, et al.
Published: (2025)
Human detectors are surprisingly powerful reward models
by: Ashutosh, Kumar, et al.
Published: (2026)
by: Ashutosh, Kumar, et al.
Published: (2026)
Make-It-Vivid: Dressing Your Animatable Biped Cartoon Characters from Text
by: Tang, Junshu, et al.
Published: (2024)
by: Tang, Junshu, et al.
Published: (2024)
Revisiting Reweighted Risk for Calibration: AURC, Focal, and Inverse Focal Loss
by: Zhou, Han, et al.
Published: (2025)
by: Zhou, Han, et al.
Published: (2025)
Federated Vision Transformer with Adaptive Focal Loss for Medical Image Classification
by: Zhao, Xinyuan, et al.
Published: (2026)
by: Zhao, Xinyuan, et al.
Published: (2026)
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
by: Zeng, Guanning, et al.
Published: (2025)
by: Zeng, Guanning, et al.
Published: (2025)
Continuous Piecewise-Affine Based Motion Model for Image Animation
by: Wang, Hexiang, et al.
Published: (2024)
by: Wang, Hexiang, et al.
Published: (2024)
Animate Your Motion: Turning Still Images into Dynamic Videos
by: Li, Mingxiao, et al.
Published: (2024)
by: Li, Mingxiao, et al.
Published: (2024)
Motion-Aware Animatable Gaussian Avatars Deblurring
by: Niu, Muyao, et al.
Published: (2024)
by: Niu, Muyao, et al.
Published: (2024)
Make It Count: Text-to-Image Generation with an Accurate Number of Objects
by: Binyamin, Lital, et al.
Published: (2024)
by: Binyamin, Lital, et al.
Published: (2024)
FocalClick-XL: Towards Unified and High-quality Interactive Segmentation
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
IM-Animation: An Implicit Motion Representation for Identity-decoupled Character Animation
by: Xu, Zhufeng, et al.
Published: (2026)
by: Xu, Zhufeng, et al.
Published: (2026)
HyperMotionX: The Dataset and Benchmark with DiT-Based Pose-Guided Human Image Animation of Complex Motions
by: Xu, Shuolin, et al.
Published: (2025)
by: Xu, Shuolin, et al.
Published: (2025)
Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
by: Hu, Li, et al.
Published: (2023)
by: Hu, Li, et al.
Published: (2023)
X-Dyna: Expressive Dynamic Human Image Animation
by: Chang, Di, et al.
Published: (2025)
by: Chang, Di, et al.
Published: (2025)
Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation
by: Chen, Yingjie, et al.
Published: (2025)
by: Chen, Yingjie, et al.
Published: (2025)
Generating Illustrated Instructions
by: Menon, Sachit, et al.
Published: (2023)
by: Menon, Sachit, et al.
Published: (2023)
GaussianMotion: End-to-End Learning of Animatable Gaussian Avatars with Pose Guidance from Text
by: Shim, Gyumin, et al.
Published: (2025)
by: Shim, Gyumin, et al.
Published: (2025)
AnyI2V: Animating Any Conditional Image with Motion Control
by: Li, Ziye, et al.
Published: (2025)
by: Li, Ziye, et al.
Published: (2025)
MoVer: Motion Verification for Motion Graphics Animations
by: Ma, Jiaju, et al.
Published: (2025)
by: Ma, Jiaju, et al.
Published: (2025)
CountDiffusion: Text-to-Image Synthesis with Training-Free Counting-Guidance Diffusion
by: Li, Yanyu, et al.
Published: (2025)
by: Li, Yanyu, et al.
Published: (2025)
MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image Animation
by: Hu, Xirui, et al.
Published: (2026)
by: Hu, Xirui, et al.
Published: (2026)
Similar Items
-
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
by: Girdhar, Rohit, et al.
Published: (2023) -
Diffusion Autoencoders are Scalable Image Tokenizers
by: Chen, Yinbo, et al.
Published: (2025) -
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024) -
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
by: Lai, Bolin, et al.
Published: (2025) -
Generating Multi-Image Synthetic Data for Text-to-Image Customization
by: Kumari, Nupur, et al.
Published: (2025)