VC-Agent: An Interactive Agent for Customized Video Dataset Collection
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yidan, Xu, Mutian, Hao, Yiming, Zhou, Kun, Chang, Jiahao, Liu, Xiaoqiang, Wan, Pengfei, Fu, Hongbo, Han, Xiaoguang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation
di: Zhao, Hongxiang, et al.
Pubblicazione: (2025)
di: Zhao, Hongxiang, et al.
Pubblicazione: (2025)
Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth Diffusion
di: Xu, Mutian, et al.
Pubblicazione: (2025)
di: Xu, Mutian, et al.
Pubblicazione: (2025)
LoFA: Learning to Predict Personalized Priors for Fast Adaptation of Visual Generative Models
di: Hao, Yiming, et al.
Pubblicazione: (2025)
di: Hao, Yiming, et al.
Pubblicazione: (2025)
StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration
di: Hu, Panwen, et al.
Pubblicazione: (2024)
di: Hu, Panwen, et al.
Pubblicazione: (2024)
ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos
di: Chen, Yuantao, et al.
Pubblicazione: (2026)
di: Chen, Yuantao, et al.
Pubblicazione: (2026)
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
di: Hu, Jiahao, et al.
Pubblicazione: (2024)
di: Hu, Jiahao, et al.
Pubblicazione: (2024)
SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation
di: Xu, Mutian, et al.
Pubblicazione: (2023)
di: Xu, Mutian, et al.
Pubblicazione: (2023)
3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
di: Fang, Zhixue, et al.
Pubblicazione: (2026)
di: Fang, Zhixue, et al.
Pubblicazione: (2026)
Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation
di: Xu, Mutian, et al.
Pubblicazione: (2026)
di: Xu, Mutian, et al.
Pubblicazione: (2026)
Motion Inversion for Video Customization
di: Wang, Luozhou, et al.
Pubblicazione: (2024)
di: Wang, Luozhou, et al.
Pubblicazione: (2024)
GaussReg: Fast 3D Registration with Gaussian Splatting
di: Chang, Jiahao, et al.
Pubblicazione: (2024)
di: Chang, Jiahao, et al.
Pubblicazione: (2024)
A Survey of Interactive Generative Video
di: Yu, Jiwen, et al.
Pubblicazione: (2025)
di: Yu, Jiwen, et al.
Pubblicazione: (2025)
SketchVideo: Sketch-based Video Generation and Editing
di: Liu, Feng-Lin, et al.
Pubblicazione: (2025)
di: Liu, Feng-Lin, et al.
Pubblicazione: (2025)
CustomSketching: Sketch Concept Extraction for Sketch-based Image Synthesis and Editing
di: Xiao, Chufeng, et al.
Pubblicazione: (2024)
di: Xiao, Chufeng, et al.
Pubblicazione: (2024)
Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
di: Li, Jiahua, et al.
Pubblicazione: (2025)
di: Li, Jiahua, et al.
Pubblicazione: (2025)
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
di: Huang, Yuzhou, et al.
Pubblicazione: (2025)
di: Huang, Yuzhou, et al.
Pubblicazione: (2025)
Sketch2Human: Deep Human Generation with Disentangled Geometry and Appearance Control
di: Qu, Linzi, et al.
Pubblicazione: (2024)
di: Qu, Linzi, et al.
Pubblicazione: (2024)
UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation
di: Xu, Yiyan, et al.
Pubblicazione: (2026)
di: Xu, Yiyan, et al.
Pubblicazione: (2026)
VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
di: Zhi, Zhuo, et al.
Pubblicazione: (2025)
di: Zhi, Zhuo, et al.
Pubblicazione: (2025)
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
di: Yang, Shiyuan, et al.
Pubblicazione: (2024)
di: Yang, Shiyuan, et al.
Pubblicazione: (2024)
Agent Attention: On the Integration of Softmax and Linear Attention
di: Han, Dongchen, et al.
Pubblicazione: (2023)
di: Han, Dongchen, et al.
Pubblicazione: (2023)
VC-LLM: Automated Advertisement Video Creation from Raw Footage using Multi-modal LLMs
di: Qian, Dongjun, et al.
Pubblicazione: (2025)
di: Qian, Dongjun, et al.
Pubblicazione: (2025)
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
di: Chen, Kaijin, et al.
Pubblicazione: (2026)
di: Chen, Kaijin, et al.
Pubblicazione: (2026)
SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation
di: Xu, Xuancheng, et al.
Pubblicazione: (2025)
di: Xu, Xuancheng, et al.
Pubblicazione: (2025)
ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation
di: Chang, Jiahao, et al.
Pubblicazione: (2025)
di: Chang, Jiahao, et al.
Pubblicazione: (2025)
ReSeDis: A Dataset for Referring-based Object Search across Large-Scale Image Collections
di: Huang, Ziling, et al.
Pubblicazione: (2025)
di: Huang, Ziling, et al.
Pubblicazione: (2025)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
StructLayoutFormer:Conditional Structured Layout Generation via Structure Serialization and Disentanglement
di: Hu, Xin, et al.
Pubblicazione: (2025)
di: Hu, Xin, et al.
Pubblicazione: (2025)
MVHumanNet++: A Large-scale Dataset of Multi-view Daily Dressing Human Captures with Richer Annotations for 3D Human Digitization
di: Li, Chenghong, et al.
Pubblicazione: (2025)
di: Li, Chenghong, et al.
Pubblicazione: (2025)
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
di: Yin, Zhiyu, et al.
Pubblicazione: (2026)
di: Yin, Zhiyu, et al.
Pubblicazione: (2026)
Hi3DGen: High-fidelity 3D Geometry Generation from Images via Normal Bridging
di: Ye, Chongjie, et al.
Pubblicazione: (2025)
di: Ye, Chongjie, et al.
Pubblicazione: (2025)
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
di: Chen, Ming, et al.
Pubblicazione: (2025)
di: Chen, Ming, et al.
Pubblicazione: (2025)
MonoHair: High-Fidelity Hair Modeling from a Monocular Video
di: Wu, Keyu, et al.
Pubblicazione: (2024)
di: Wu, Keyu, et al.
Pubblicazione: (2024)
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
di: Zhou, Yiyang, et al.
Pubblicazione: (2025)
di: Zhou, Yiyang, et al.
Pubblicazione: (2025)
FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization
di: Song, Quanjian, et al.
Pubblicazione: (2026)
di: Song, Quanjian, et al.
Pubblicazione: (2026)
MV-Performer: Taming Video Diffusion Model for Faithful and Synchronized Multi-view Performer Synthesis
di: Zhi, Yihao, et al.
Pubblicazione: (2025)
di: Zhi, Yihao, et al.
Pubblicazione: (2025)
Lighting-grounded Video Generation with Renderer-based Agent Reasoning
di: Cai, Ziqi, et al.
Pubblicazione: (2026)
di: Cai, Ziqi, et al.
Pubblicazione: (2026)
GauStudio: A Modular Framework for 3D Gaussian Splatting and Beyond
di: Ye, Chongjie, et al.
Pubblicazione: (2024)
di: Ye, Chongjie, et al.
Pubblicazione: (2024)
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models
di: Huang, Chi-Pin, et al.
Pubblicazione: (2025)
di: Huang, Chi-Pin, et al.
Pubblicazione: (2025)
GameFactory: Creating New Games with Generative Interactive Videos
di: Yu, Jiwen, et al.
Pubblicazione: (2025)
di: Yu, Jiwen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation
di: Zhao, Hongxiang, et al.
Pubblicazione: (2025) -
Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth Diffusion
di: Xu, Mutian, et al.
Pubblicazione: (2025) -
LoFA: Learning to Predict Personalized Priors for Fast Adaptation of Visual Generative Models
di: Hao, Yiming, et al.
Pubblicazione: (2025) -
StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration
di: Hu, Panwen, et al.
Pubblicazione: (2024) -
ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos
di: Chen, Yuantao, et al.
Pubblicazione: (2026)