Saved in:
| Main Authors: | Yang, Yuhang, Fan, Ke, Sun, Shangkun, Li, Hongxiang, Zeng, Ailing, Han, FeiLin, Zhai, Wei, Liu, Wei, Cao, Yang, Zha, Zheng-Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.23452 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HERO: Human Reaction Generation from Videos
by: Yu, Chengjun, et al.
Published: (2025)
by: Yu, Chengjun, et al.
Published: (2025)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
by: Xi, Haocheng, et al.
Published: (2025)
by: Xi, Haocheng, et al.
Published: (2025)
DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
by: Xi, Haocheng, et al.
Published: (2026)
by: Xi, Haocheng, et al.
Published: (2026)
The Dawn of Video Generation: Preliminary Explorations with SORA-like Models
by: Zeng, Ailing, et al.
Published: (2024)
by: Zeng, Ailing, et al.
Published: (2024)
Gloria: Consistent Character Video Generation via Content Anchors
by: Yang, Yuhang, et al.
Published: (2026)
by: Yang, Yuhang, et al.
Published: (2026)
TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
by: Han, Guangyi, et al.
Published: (2025)
by: Han, Guangyi, et al.
Published: (2025)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
by: Zheng, Mingzhe, et al.
Published: (2025)
by: Zheng, Mingzhe, et al.
Published: (2025)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
by: Zheng, Mingzhe, et al.
Published: (2024)
by: Zheng, Mingzhe, et al.
Published: (2024)
Improved Video VAE for Latent Video Diffusion Model
by: Wu, Pingyu, et al.
Published: (2024)
by: Wu, Pingyu, et al.
Published: (2024)
LEMON: Learning 3D Human-Object Interaction Relation from 2D Images
by: Yang, Yuhang, et al.
Published: (2023)
by: Yang, Yuhang, et al.
Published: (2023)
End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
by: Shao, Yawen, et al.
Published: (2024)
by: Shao, Yawen, et al.
Published: (2024)
EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views
by: Yang, Yuhang, et al.
Published: (2024)
by: Yang, Yuhang, et al.
Published: (2024)
ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models
by: Fang, Zixun, et al.
Published: (2025)
by: Fang, Zixun, et al.
Published: (2025)
GRACE: Estimating Geometry-level 3D Human-Scene Contact from 2D Images
by: Wang, Chengfeng, et al.
Published: (2025)
by: Wang, Chengfeng, et al.
Published: (2025)
Exploring AIGC Video Quality: A Focus on Visual Harmony, Video-Text Consistency and Domain Distribution Gap
by: Qu, Bowen, et al.
Published: (2024)
by: Qu, Bowen, et al.
Published: (2024)
Grounding 3D Scene Affordance From Egocentric Interactions
by: Liu, Cuiyu, et al.
Published: (2024)
by: Liu, Cuiyu, et al.
Published: (2024)
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
by: Liu, Ruyang, et al.
Published: (2025)
by: Liu, Ruyang, et al.
Published: (2025)
VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment
by: Sun, Shangkun, et al.
Published: (2024)
by: Sun, Shangkun, et al.
Published: (2024)
RAIN: Real-time Animation of Infinite Video Stream
by: Shu, Zhilei, et al.
Published: (2024)
by: Shu, Zhilei, et al.
Published: (2024)
VanGogh: A Unified Multimodal Diffusion-based Framework for Video Colorization
by: Fang, Zixun, et al.
Published: (2025)
by: Fang, Zixun, et al.
Published: (2025)
MATE: Motion-Augmented Temporal Consistency for Event-based Point Tracking
by: Han, Han, et al.
Published: (2024)
by: Han, Han, et al.
Published: (2024)
Content-Rich AIGC Video Quality Assessment via Intricate Text Alignment and Motion-Aware Consistency
by: Sun, Shangkun, et al.
Published: (2025)
by: Sun, Shangkun, et al.
Published: (2025)
Event Stream Filtering via Probability Flux Estimation
by: Chen, Jinze, et al.
Published: (2025)
by: Chen, Jinze, et al.
Published: (2025)
Visual-Geometric Collaborative Guidance for Affordance Learning
by: Luo, Hongchen, et al.
Published: (2024)
by: Luo, Hongchen, et al.
Published: (2024)
Leverage Task Context for Object Affordance Ranking
by: Huang, Haojie, et al.
Published: (2024)
by: Huang, Haojie, et al.
Published: (2024)
ViViD: Video Virtual Try-on using Diffusion Models
by: Fang, Zixun, et al.
Published: (2024)
by: Fang, Zixun, et al.
Published: (2024)
UniFinEval: Towards Unified Evaluation of Financial Multimodal Models across Text, Images and Videos
by: Yang, Zhi, et al.
Published: (2026)
by: Yang, Zhi, et al.
Published: (2026)
Unbiased Gradient Estimation for Event Binning via Functional Backpropagation
by: Chen, Jinze, et al.
Published: (2026)
by: Chen, Jinze, et al.
Published: (2026)
Intention-driven Ego-to-Exo Video Generation
by: Luo, Hongchen, et al.
Published: (2024)
by: Luo, Hongchen, et al.
Published: (2024)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
by: Sun, Shangkun, et al.
Published: (2024)
by: Sun, Shangkun, et al.
Published: (2024)
EMoTive: Event-guided Trajectory Modeling for 3D Motion Estimation
by: Wan, Zengyu, et al.
Published: (2025)
by: Wan, Zengyu, et al.
Published: (2025)
EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
by: Liu, Yaofang, et al.
Published: (2023)
by: Liu, Yaofang, et al.
Published: (2023)
Event-based Visual Deformation Measurement
by: Wu, Yuliang, et al.
Published: (2026)
by: Wu, Yuliang, et al.
Published: (2026)
Event-based Asynchronous HDR Imaging by Temporal Incident Light Modulation
by: Wu, Yuliang, et al.
Published: (2024)
by: Wu, Yuliang, et al.
Published: (2024)
MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling
by: Yang, Jian, et al.
Published: (2024)
by: Yang, Jian, et al.
Published: (2024)
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
by: Peng, Tianhao, et al.
Published: (2025)
by: Peng, Tianhao, et al.
Published: (2025)
Get In Video: Add Anything You Want to the Video
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
Similar Items
-
HERO: Human Reaction Generation from Videos
by: Yu, Chengjun, et al.
Published: (2025) -
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
by: Xi, Haocheng, et al.
Published: (2025) -
DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
by: Chen, Junyu, et al.
Published: (2025) -
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
by: Yang, Shuo, et al.
Published: (2025) -
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
by: Xi, Haocheng, et al.
Published: (2026)