Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Yuheng, Gao, Xiangbo, Chen, Tianhao, Chen, Xinghao, Yin, Qing, Tu, Zhengzhong, Lee, Dongman |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing
por: Lionar, Stefan, et al.
Publicado: (2025)
por: Lionar, Stefan, et al.
Publicado: (2025)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
por: Cheng, Shihao, et al.
Publicado: (2026)
por: Cheng, Shihao, et al.
Publicado: (2026)
Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
por: He, Liu, et al.
Publicado: (2024)
por: He, Liu, et al.
Publicado: (2024)
Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive Smoothing
por: Qi, Leyi, et al.
Publicado: (2026)
por: Qi, Leyi, et al.
Publicado: (2026)
Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation
por: Lin, Jiantao, et al.
Publicado: (2025)
por: Lin, Jiantao, et al.
Publicado: (2025)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
por: Guan, Jiazhi, et al.
Publicado: (2025)
por: Guan, Jiazhi, et al.
Publicado: (2025)
VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
por: Cha, SeungJu, et al.
Publicado: (2025)
por: Cha, SeungJu, et al.
Publicado: (2025)
Casual3DHDR: Deblurring High Dynamic Range 3D Gaussian Splatting from Casually Captured Videos
por: Gong, Shucheng, et al.
Publicado: (2025)
por: Gong, Shucheng, et al.
Publicado: (2025)
InteractDiffusion: Interaction Control in Text-to-Image Diffusion Models
por: Hoe, Jiun Tian, et al.
Publicado: (2023)
por: Hoe, Jiun Tian, et al.
Publicado: (2023)
InteractEdit: Zero-Shot Editing of Human-Object Interactions in Images
por: Hoe, Jiun Tian, et al.
Publicado: (2025)
por: Hoe, Jiun Tian, et al.
Publicado: (2025)
Unveiling Deep Shadows: A Survey and Benchmark on Image and Video Shadow Detection, Removal, and Generation in the Deep Learning Era
por: Hu, Xiaowei, et al.
Publicado: (2024)
por: Hu, Xiaowei, et al.
Publicado: (2024)
STAR: Skeleton-aware Text-based 4D Avatar Generation with In-Network Motion Retargeting
por: Chai, Zenghao, et al.
Publicado: (2024)
por: Chai, Zenghao, et al.
Publicado: (2024)
Representing Long Volumetric Video with Temporal Gaussian Hierarchy
por: Xu, Zhen, et al.
Publicado: (2024)
por: Xu, Zhen, et al.
Publicado: (2024)
DreamCinema: Cinematic Transfer with Free Camera and 3D Character
por: Chen, Weiliang, et al.
Publicado: (2024)
por: Chen, Weiliang, et al.
Publicado: (2024)
FairyGen: Storied Cartoon Video from a Single Child-Drawn Character
por: Zheng, Jiayi, et al.
Publicado: (2025)
por: Zheng, Jiayi, et al.
Publicado: (2025)
Neural Network-Based Tracking and 3D Reconstruction of Baseball Pitch Trajectories from Single-View 2D Video
por: Hsieh, Jhen
Publicado: (2024)
por: Hsieh, Jhen
Publicado: (2024)
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
por: Girdhar, Rohit, et al.
Publicado: (2023)
por: Girdhar, Rohit, et al.
Publicado: (2023)
HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
por: Dong, Wenqi, et al.
Publicado: (2025)
por: Dong, Wenqi, et al.
Publicado: (2025)
ImagenHub: Standardizing the evaluation of conditional image generation models
por: Ku, Max, et al.
Publicado: (2023)
por: Ku, Max, et al.
Publicado: (2023)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
por: Razlighi, AmirHossein Naghi, et al.
Publicado: (2026)
por: Razlighi, AmirHossein Naghi, et al.
Publicado: (2026)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
por: Guan, Jiazhi, et al.
Publicado: (2024)
por: Guan, Jiazhi, et al.
Publicado: (2024)
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
por: Zhao, Junchuan, et al.
Publicado: (2026)
por: Zhao, Junchuan, et al.
Publicado: (2026)
EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
por: Flynn, John, et al.
Publicado: (2026)
por: Flynn, John, et al.
Publicado: (2026)
MusicScore: A Dataset for Music Score Modeling and Generation
por: Lin, Yuheng, et al.
Publicado: (2024)
por: Lin, Yuheng, et al.
Publicado: (2024)
Laplacian Analysis Meets Dynamics Modelling: Gaussian Splatting for 4D Reconstruction
por: Zhou, Yifan, et al.
Publicado: (2025)
por: Zhou, Yifan, et al.
Publicado: (2025)
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
por: Huang, Yiheng, et al.
Publicado: (2025)
por: Huang, Yiheng, et al.
Publicado: (2025)
SVGS: Enhancing Gaussian Splatting Using Primitives with Spatially Varying Colors
por: Xu, Rui, et al.
Publicado: (2024)
por: Xu, Rui, et al.
Publicado: (2024)
Break-for-Make: Modular Low-Rank Adaptations for Composable Content-Style Customization
por: Xu, Yu, et al.
Publicado: (2024)
por: Xu, Yu, et al.
Publicado: (2024)
SemanticGarment: Semantic-Controlled Generation and Editing of 3D Gaussian Garments
por: Wang, Ruiyan, et al.
Publicado: (2025)
por: Wang, Ruiyan, et al.
Publicado: (2025)
MesonGS++: Post-training Compression of 3D Gaussian Splatting with Hyperparameter Searching
por: Xie, Shuzhao, et al.
Publicado: (2026)
por: Xie, Shuzhao, et al.
Publicado: (2026)
Generating Digital Models Using Text-to-3D and Image-to-3D Prompts: Critical Case Study
por: Ziatdinov, Rushan, et al.
Publicado: (2025)
por: Ziatdinov, Rushan, et al.
Publicado: (2025)
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
por: Xu, Chuanzhi, et al.
Publicado: (2026)
por: Xu, Chuanzhi, et al.
Publicado: (2026)
MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation
por: Gupta, Prerit, et al.
Publicado: (2025)
por: Gupta, Prerit, et al.
Publicado: (2025)
DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions
por: Zhang, Hengyuan, et al.
Publicado: (2025)
por: Zhang, Hengyuan, et al.
Publicado: (2025)
ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion
por: Wang, Xuanchen, et al.
Publicado: (2025)
por: Wang, Xuanchen, et al.
Publicado: (2025)
Perceive-Sample-Compress: Towards Real-Time 3D Gaussian Splatting
por: Wang, Zijian, et al.
Publicado: (2025)
por: Wang, Zijian, et al.
Publicado: (2025)
Splatography: Sparse multi-view dynamic Gaussian Splatting for filmmaking challenges
por: Azzarelli, Adrian, et al.
Publicado: (2025)
por: Azzarelli, Adrian, et al.
Publicado: (2025)
altiro3D: Scene representation from single image and novel view synthesis
por: Canessa, E., et al.
Publicado: (2023)
por: Canessa, E., et al.
Publicado: (2023)
Exploring Palette based Color Guidance in Diffusion Models
por: Qiu, Qianru, et al.
Publicado: (2025)
por: Qiu, Qianru, et al.
Publicado: (2025)
Real-Time Position-Aware View Synthesis from Single-View Input
por: Gond, Manu, et al.
Publicado: (2024)
por: Gond, Manu, et al.
Publicado: (2024)
Ejemplares similares
-
TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing
por: Lionar, Stefan, et al.
Publicado: (2025) -
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
por: Cheng, Shihao, et al.
Publicado: (2026) -
Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
por: He, Liu, et al.
Publicado: (2024) -
Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive Smoothing
por: Qi, Leyi, et al.
Publicado: (2026) -
Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation
por: Lin, Jiantao, et al.
Publicado: (2025)