Learning Human Skill Generators at Key-Step Levels
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Yilu, Zhu, Chenhui, Wang, Shuai, Wang, Hanlin, Wang, Jing, Zhang, Zhaoxiang, Wang, Limin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025)
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025)
Open-Event Procedure Planning in Instructional Videos
von: Wu, Yilu, et al.
Veröffentlicht: (2024)
von: Wu, Yilu, et al.
Veröffentlicht: (2024)
PDPP: Projected Diffusion for Procedure Planning in Instructional Videos
von: Wang, Hanlin, et al.
Veröffentlicht: (2023)
von: Wang, Hanlin, et al.
Veröffentlicht: (2023)
PixNerd: Pixel Neural Field Diffusion
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
CycleHOI: Improving Human-Object Interaction Detection with Cycle Consistency of Detection and Generation
von: Wang, Yisen, et al.
Veröffentlicht: (2024)
von: Wang, Yisen, et al.
Veröffentlicht: (2024)
Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
Contextual AD Narration with Interleaved Multimodal Sequence
von: Wang, Hanlin, et al.
Veröffentlicht: (2024)
von: Wang, Hanlin, et al.
Veröffentlicht: (2024)
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
Dual DETRs for Multi-Label Temporal Action Detection
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
Linearized Coupling Flow with Shortcut Constraints for One-Step Face Restoration
von: Sun, Xiaohui, et al.
Veröffentlicht: (2026)
von: Sun, Xiaohui, et al.
Veröffentlicht: (2026)
Open Vocabulary 3D Scene Understanding via Geometry Guided Self-Distillation
von: Wang, Pengfei, et al.
Veröffentlicht: (2024)
von: Wang, Pengfei, et al.
Veröffentlicht: (2024)
FreeVS: Generative View Synthesis on Free Driving Trajectory
von: Wang, Qitai, et al.
Veröffentlicht: (2024)
von: Wang, Qitai, et al.
Veröffentlicht: (2024)
OOD-HOI: Text-Driven 3D Whole-Body Human-Object Interactions Generation Beyond Training Domains
von: Zhang, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2024)
HLG: Comprehensive 3D Room Construction via Hierarchical Layout Generation
von: Wang, Xiping, et al.
Veröffentlicht: (2025)
von: Wang, Xiping, et al.
Veröffentlicht: (2025)
Concept Unlearning by Modeling Key Steps of Diffusion Process
von: Zhang, Chaoshuo, et al.
Veröffentlicht: (2025)
von: Zhang, Chaoshuo, et al.
Veröffentlicht: (2025)
Weakly Supervised 3D Object Detection with Multi-Stage Generalization
von: He, Jiawei, et al.
Veröffentlicht: (2023)
von: He, Jiawei, et al.
Veröffentlicht: (2023)
LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis
von: Wang, Hanlin, et al.
Veröffentlicht: (2024)
von: Wang, Hanlin, et al.
Veröffentlicht: (2024)
Arbitrary Generative Video Interpolation
von: Zhang, Guozhen, et al.
Veröffentlicht: (2025)
von: Zhang, Guozhen, et al.
Veröffentlicht: (2025)
RoomCraft: Controllable and Complete 3D Indoor Scene Generation
von: Zhou, Mengqi, et al.
Veröffentlicht: (2025)
von: Zhou, Mengqi, et al.
Veröffentlicht: (2025)
Learning Multi-dimensional Human Preference for Text-to-Image Generation
von: Zhang, Sixian, et al.
Veröffentlicht: (2024)
von: Zhang, Sixian, et al.
Veröffentlicht: (2024)
FlowDCN: Exploring DCN-like Architectures for Fast Image Generation with Arbitrary Resolution
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers
von: Chen, Yuntao, et al.
Veröffentlicht: (2024)
von: Chen, Yuntao, et al.
Veröffentlicht: (2024)
A Curriculum-style Self-training Approach for Source-Free Semantic Segmentation
von: Wang, Yuxi, et al.
Veröffentlicht: (2021)
von: Wang, Yuxi, et al.
Veröffentlicht: (2021)
DMM: Building a Versatile Image Generation Model via Distillation-Based Model Merging
von: Song, Tianhui, et al.
Veröffentlicht: (2025)
von: Song, Tianhui, et al.
Veröffentlicht: (2025)
DDT: Decoupled Diffusion Transformer
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
ZeroI2V: Zero-Cost Adaptation of Pre-trained Transformers from Image to Video
von: Li, Xinhao, et al.
Veröffentlicht: (2023)
von: Li, Xinhao, et al.
Veröffentlicht: (2023)
OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning
von: Gong, Yuan, et al.
Veröffentlicht: (2025)
von: Gong, Yuan, et al.
Veröffentlicht: (2025)
Towards Human-Like Machine Comprehension: Few-Shot Relational Learning in Visually-Rich Documents
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Motion-Aware Generative Frame Interpolation
von: Zhang, Guozhen, et al.
Veröffentlicht: (2025)
von: Zhang, Guozhen, et al.
Veröffentlicht: (2025)
Recovering 3D Human Mesh from Monocular Images: A Survey
von: Tian, Yating, et al.
Veröffentlicht: (2022)
von: Tian, Yating, et al.
Veröffentlicht: (2022)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
LoD-Loc v2: Aerial Visual Localization over Low Level-of-Detail City Models using Explicit Silhouette Alignment
von: Zhu, Juelin, et al.
Veröffentlicht: (2025)
von: Zhu, Juelin, et al.
Veröffentlicht: (2025)
MeMOTR: Long-Term Memory-Augmented Transformer for Multi-Object Tracking
von: Gao, Ruopeng, et al.
Veröffentlicht: (2023)
von: Gao, Ruopeng, et al.
Veröffentlicht: (2023)
Denoising Diffusion Step-aware Models
von: Yang, Shuai, et al.
Veröffentlicht: (2023)
von: Yang, Shuai, et al.
Veröffentlicht: (2023)
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
PPJudge: Towards Human-Aligned Assessment of Artistic Painting Process
von: Jiang, Shiqi, et al.
Veröffentlicht: (2025)
von: Jiang, Shiqi, et al.
Veröffentlicht: (2025)
Using Unreliable Pseudo-Labels for Label-Efficient Semantic Segmentation
von: Wang, Haochen, et al.
Veröffentlicht: (2023)
von: Wang, Haochen, et al.
Veröffentlicht: (2023)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
von: Wang, Chenting, et al.
Veröffentlicht: (2025)
von: Wang, Chenting, et al.
Veröffentlicht: (2025)
TIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion Generation
von: Wang, Yabiao, et al.
Veröffentlicht: (2024)
von: Wang, Yabiao, et al.
Veröffentlicht: (2024)
Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks
von: Yang, Min, et al.
Veröffentlicht: (2024)
von: Yang, Min, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025) -
Open-Event Procedure Planning in Instructional Videos
von: Wu, Yilu, et al.
Veröffentlicht: (2024) -
PDPP: Projected Diffusion for Procedure Planning in Instructional Videos
von: Wang, Hanlin, et al.
Veröffentlicht: (2023) -
PixNerd: Pixel Neural Field Diffusion
von: Wang, Shuai, et al.
Veröffentlicht: (2025) -
CycleHOI: Improving Human-Object Interaction Detection with Cycle Consistency of Detection and Generation
von: Wang, Yisen, et al.
Veröffentlicht: (2024)