AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yiheng, Li, Zhuo, Hou, Ruibing, Chen, Yingjie, Chang, Hong, Liu, Hao, Shan, Shiguang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild
von: Chen, Baiyu, et al.
Veröffentlicht: (2026)
von: Chen, Baiyu, et al.
Veröffentlicht: (2026)
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
von: Li, Yiheng, et al.
Veröffentlicht: (2024)
von: Li, Yiheng, et al.
Veröffentlicht: (2024)
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
von: Li, Yinqi, et al.
Veröffentlicht: (2025)
von: Li, Yinqi, et al.
Veröffentlicht: (2025)
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
von: Hou, Ruibing, et al.
Veröffentlicht: (2025)
von: Hou, Ruibing, et al.
Veröffentlicht: (2025)
M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
von: Luo, Mingshuang, et al.
Veröffentlicht: (2024)
von: Luo, Mingshuang, et al.
Veröffentlicht: (2024)
AnyI2V: Animating Any Conditional Image with Motion Control
von: Li, Ziye, et al.
Veröffentlicht: (2025)
von: Li, Ziye, et al.
Veröffentlicht: (2025)
EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation
von: Hou, Ruibing, et al.
Veröffentlicht: (2026)
von: Hou, Ruibing, et al.
Veröffentlicht: (2026)
RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios
von: Huang, Jie, et al.
Veröffentlicht: (2024)
von: Huang, Jie, et al.
Veröffentlicht: (2024)
Component-Based Out-of-Distribution Detection
von: Liu, Wenrui, et al.
Veröffentlicht: (2026)
von: Liu, Wenrui, et al.
Veröffentlicht: (2026)
Morph: A Motion-free Physics Optimization Framework for Human Motion Generation
von: Li, Zhuo, et al.
Veröffentlicht: (2024)
von: Li, Zhuo, et al.
Veröffentlicht: (2024)
un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
von: Li, Yinqi, et al.
Veröffentlicht: (2025)
von: Li, Yinqi, et al.
Veröffentlicht: (2025)
UMFC: Unsupervised Multi-Domain Feature Calibration for Vision-Language Models
von: Liang, Jiachen, et al.
Veröffentlicht: (2024)
von: Liang, Jiachen, et al.
Veröffentlicht: (2024)
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
Generalized Semi-Supervised Learning via Self-Supervised Feature Adaptation
von: Liang, Jiachen, et al.
Veröffentlicht: (2024)
von: Liang, Jiachen, et al.
Veröffentlicht: (2024)
Motion Anything: Any to Motion Generation
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
Revisiting Logit Distributions for Reliable Out-of-Distribution Detection
von: Liang, Jiachen, et al.
Veröffentlicht: (2025)
von: Liang, Jiachen, et al.
Veröffentlicht: (2025)
FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
von: Jiang, Yue, et al.
Veröffentlicht: (2025)
von: Jiang, Yue, et al.
Veröffentlicht: (2025)
Any2Any: Unified Arbitrary Modality Translation for Remote Sensing
von: Chen, Haoyang, et al.
Veröffentlicht: (2026)
von: Chen, Haoyang, et al.
Veröffentlicht: (2026)
MDReID: Modality-Decoupled Learning for Any-to-Any Multi-Modal Object Re-Identification
von: Feng, Yingying, et al.
Veröffentlicht: (2025)
von: Feng, Yingying, et al.
Veröffentlicht: (2025)
SAMTok: Representing Any Mask with Two Words
von: Zhou, Yikang, et al.
Veröffentlicht: (2026)
von: Zhou, Yikang, et al.
Veröffentlicht: (2026)
Clothes-Changing Person Re-Identification with Feasibility-Aware Intermediary Matching
von: Zhao, Jiahe, et al.
Veröffentlicht: (2024)
von: Zhao, Jiahe, et al.
Veröffentlicht: (2024)
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
von: Li, Keliang, et al.
Veröffentlicht: (2024)
von: Li, Keliang, et al.
Veröffentlicht: (2024)
Depth Anything at Any Condition
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
ReMoMask: Retrieval-Augmented Masked Motion Generation
von: Li, Zhengdao, et al.
Veröffentlicht: (2025)
von: Li, Zhengdao, et al.
Veröffentlicht: (2025)
OmniControl: Control Any Joint at Any Time for Human Motion Generation
von: Xie, Yiming, et al.
Veröffentlicht: (2023)
von: Xie, Yiming, et al.
Veröffentlicht: (2023)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks
von: Luo, Mingshuang, et al.
Veröffentlicht: (2026)
von: Luo, Mingshuang, et al.
Veröffentlicht: (2026)
AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D Diffusion
von: Li, Hongjie, et al.
Veröffentlicht: (2026)
von: Li, Hongjie, et al.
Veröffentlicht: (2026)
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
von: Li, Shufan, et al.
Veröffentlicht: (2024)
von: Li, Shufan, et al.
Veröffentlicht: (2024)
Segment Any Motion in Videos
von: Huang, Nan, et al.
Veröffentlicht: (2025)
von: Huang, Nan, et al.
Veröffentlicht: (2025)
Segment Any Change
von: Zheng, Zhuo, et al.
Veröffentlicht: (2024)
von: Zheng, Zhuo, et al.
Veröffentlicht: (2024)
AnyAD: Unified Any-Modality Anomaly Detection in Incomplete Multi-Sequence MRI
von: Wu, Changwei, et al.
Veröffentlicht: (2025)
von: Wu, Changwei, et al.
Veröffentlicht: (2025)
AnyTSR: Any-Scale Thermal Super-Resolution for UAV
von: Li, Mengyuan, et al.
Veröffentlicht: (2025)
von: Li, Mengyuan, et al.
Veröffentlicht: (2025)
MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised Learning
von: Hong, Jingshan, et al.
Veröffentlicht: (2025)
von: Hong, Jingshan, et al.
Veröffentlicht: (2025)
AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities
von: Astruc, Guillaume, et al.
Veröffentlicht: (2024)
von: Astruc, Guillaume, et al.
Veröffentlicht: (2024)
InstructSAM: Segment Any Instance with Any Instructions
von: Yuan, Yuqian, et al.
Veröffentlicht: (2026)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2026)
OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
von: Chen, Jinshu, et al.
Veröffentlicht: (2025)
von: Chen, Jinshu, et al.
Veröffentlicht: (2025)
MoSa: Motion Generation with Scalable Autoregressive Modeling
von: Liu, Mengyuan, et al.
Veröffentlicht: (2025)
von: Liu, Mengyuan, et al.
Veröffentlicht: (2025)
AnySR: Realizing Image Super-Resolution as Any-Scale, Any-Resource
von: Zhan, Wengyi, et al.
Veröffentlicht: (2024)
von: Zhan, Wengyi, et al.
Veröffentlicht: (2024)
AnyFit: Controllable Virtual Try-on for Any Combination of Attire Across Any Scenario
von: Li, Yuhan, et al.
Veröffentlicht: (2024)
von: Li, Yuhan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild
von: Chen, Baiyu, et al.
Veröffentlicht: (2026) -
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
von: Li, Yiheng, et al.
Veröffentlicht: (2024) -
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
von: Li, Yinqi, et al.
Veröffentlicht: (2025) -
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
von: Hou, Ruibing, et al.
Veröffentlicht: (2025) -
M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
von: Luo, Mingshuang, et al.
Veröffentlicht: (2024)