Saved in:
| Main Authors: | Zeng, Ailing, Yang, Yuhang, Chen, Weidong, Liu, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.05227 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoGen-Eval: Agent-based System for Video Generation Evaluation
by: Yang, Yuhang, et al.
Published: (2025)
by: Yang, Yuhang, et al.
Published: (2025)
Med-SORA: Symptom to Organ Reasoning in Abdomen CT Images
by: Na, You-Kyoung, et al.
Published: (2025)
by: Na, You-Kyoung, et al.
Published: (2025)
A Preliminary Exploration Towards General Image Restoration
by: Kong, Xiangtao, et al.
Published: (2024)
by: Kong, Xiangtao, et al.
Published: (2024)
Virgo: A Preliminary Exploration on Reproducing o1-like MLLM
by: Du, Yifan, et al.
Published: (2025)
by: Du, Yifan, et al.
Published: (2025)
VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
SORA: Free Second-Order Attacks in Fast Adversarial Training
by: Teymourian, Mazdak, et al.
Published: (2026)
by: Teymourian, Mazdak, et al.
Published: (2026)
HERO: Human Reaction Generation from Videos
by: Yu, Chengjun, et al.
Published: (2025)
by: Yu, Chengjun, et al.
Published: (2025)
TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
by: Chen, Jiaben, et al.
Published: (2025)
by: Chen, Jiaben, et al.
Published: (2025)
X-Pose: Detecting Any Keypoints
by: Yang, Jie, et al.
Published: (2023)
by: Yang, Jie, et al.
Published: (2023)
The Quest for Generalizable Motion Generation: Data, Model, and Evaluation
by: Lin, Jing, et al.
Published: (2025)
by: Lin, Jing, et al.
Published: (2025)
MotionLLM: Understanding Human Behaviors from Human Motions and Videos
by: Chen, Ling-Hao, et al.
Published: (2024)
by: Chen, Ling-Hao, et al.
Published: (2024)
UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation
by: Guo, Qin, et al.
Published: (2025)
by: Guo, Qin, et al.
Published: (2025)
Open-World Human-Object Interaction Detection via Multi-modal Prompts
by: Yang, Jie, et al.
Published: (2024)
by: Yang, Jie, et al.
Published: (2024)
Generative Photographic Control for Scene-Consistent Video Cinematic Editing
by: Sun, Huiqiang, et al.
Published: (2025)
by: Sun, Huiqiang, et al.
Published: (2025)
MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls
by: Bian, Yuxuan, et al.
Published: (2024)
by: Bian, Yuxuan, et al.
Published: (2024)
Dual-path Collaborative Generation Network for Emotional Video Captioning
by: Ye, Cheng, et al.
Published: (2024)
by: Ye, Cheng, et al.
Published: (2024)
EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
by: Liu, Yaofang, et al.
Published: (2023)
by: Liu, Yaofang, et al.
Published: (2023)
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
by: Hu, Siyuan, et al.
Published: (2024)
by: Hu, Siyuan, et al.
Published: (2024)
Gloria: Consistent Character Video Generation via Content Anchors
by: Yang, Yuhang, et al.
Published: (2026)
by: Yang, Yuhang, et al.
Published: (2026)
RGC-VQA: An Exploration Database for Robotic-Generated Video Quality Assessment
by: Jin, Jianing, et al.
Published: (2025)
by: Jin, Jianing, et al.
Published: (2025)
GPAvatar: Generalizable and Precise Head Avatar from Image(s)
by: Chu, Xuangeng, et al.
Published: (2024)
by: Chu, Xuangeng, et al.
Published: (2024)
Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations
by: Mao, Shunqi, et al.
Published: (2025)
by: Mao, Shunqi, et al.
Published: (2025)
Information Bottleneck Approach to Spatial Attention Learning
by: Lai, Qiuxia, et al.
Published: (2021)
by: Lai, Qiuxia, et al.
Published: (2021)
A Preliminary Study on GPT-Image Generation Model for Image Restoration
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
by: Ju, Xuan, et al.
Published: (2024)
by: Ju, Xuan, et al.
Published: (2024)
Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases
by: Liu, Xinpeng, et al.
Published: (2023)
by: Liu, Xinpeng, et al.
Published: (2023)
DreamJourney: Perpetual View Generation with Video Diffusion Models
by: Pan, Bo, et al.
Published: (2025)
by: Pan, Bo, et al.
Published: (2025)
TubeMLLM: A Foundation Model for Topology Knowledge Exploration in Vessel-like Anatomy
by: Liu, Yaoyu, et al.
Published: (2026)
by: Liu, Yaoyu, et al.
Published: (2026)
ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
by: Bu, Jiazi, et al.
Published: (2024)
by: Bu, Jiazi, et al.
Published: (2024)
Lie Flow: Video Dynamic Fields Modeling and Predicting with Lie Algebra as Geometric Physics Principle
by: Qiao, Weidong, et al.
Published: (2026)
by: Qiao, Weidong, et al.
Published: (2026)
EEA: Exploration-Exploitation Agent for Long Video Understanding
by: Yang, Te, et al.
Published: (2025)
by: Yang, Te, et al.
Published: (2025)
Do Current Video LLMs Have Strong OCR Abilities? A Preliminary Study
by: Fei, Yulin, et al.
Published: (2024)
by: Fei, Yulin, et al.
Published: (2024)
DPoser: Diffusion Model as Robust 3D Human Pose Prior
by: Lu, Junzhe, et al.
Published: (2023)
by: Lu, Junzhe, et al.
Published: (2023)
Manifold-Aware Exploration for Reinforcement Learning in Video Generation
by: Zheng, Mingzhe, et al.
Published: (2026)
by: Zheng, Mingzhe, et al.
Published: (2026)
OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation
by: Xu, Guowei, et al.
Published: (2025)
by: Xu, Guowei, et al.
Published: (2025)
ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models
by: Fang, Zixun, et al.
Published: (2025)
by: Fang, Zixun, et al.
Published: (2025)
VideoRoPE: What Makes for Good Video Rotary Position Embedding?
by: Wei, Xilin, et al.
Published: (2025)
by: Wei, Xilin, et al.
Published: (2025)
Towards Long Video Understanding via Fine-detailed Video Story Generation
by: You, Zeng, et al.
Published: (2024)
by: You, Zeng, et al.
Published: (2024)
Opening the Black Box: Preliminary Insights into Affective Modeling in Multimodal Foundation Models
by: Zhang, Zhen, et al.
Published: (2026)
by: Zhang, Zhen, et al.
Published: (2026)
ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
by: Chen, Lin, et al.
Published: (2024)
by: Chen, Lin, et al.
Published: (2024)
Similar Items
-
VideoGen-Eval: Agent-based System for Video Generation Evaluation
by: Yang, Yuhang, et al.
Published: (2025) -
Med-SORA: Symptom to Organ Reasoning in Abdomen CT Images
by: Na, You-Kyoung, et al.
Published: (2025) -
A Preliminary Exploration Towards General Image Restoration
by: Kong, Xiangtao, et al.
Published: (2024) -
Virgo: A Preliminary Exploration on Reproducing o1-like MLLM
by: Du, Yifan, et al.
Published: (2025) -
VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer
by: Liu, Xinyu, et al.
Published: (2025)