Scaling Image and Video Generation via Test-Time Evolutionary Search
Fuente:
arXiv
Guardado en:
| Autores principales: | He, Haoran, Liang, Jiajun, Wang, Xintao, Wan, Pengfei, Zhang, Di, Gai, Kun, Pan, Ling |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
por: Huang, Yuzhou, et al.
Publicado: (2025)
por: Huang, Yuzhou, et al.
Publicado: (2025)
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
por: Cheng, Junhao, et al.
Publicado: (2026)
por: Cheng, Junhao, et al.
Publicado: (2026)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
por: Wu, Shengqiong, et al.
Publicado: (2025)
por: Wu, Shengqiong, et al.
Publicado: (2025)
GARDO: Reinforcing Diffusion Models without Reward Hacking
por: He, Haoran, et al.
Publicado: (2025)
por: He, Haoran, et al.
Publicado: (2025)
UniVideo: Unified Understanding, Generation, and Editing for Videos
por: Wei, Cong, et al.
Publicado: (2025)
por: Wei, Cong, et al.
Publicado: (2025)
Pre-Trained Video Generative Models as World Simulators
por: He, Haoran, et al.
Publicado: (2025)
por: He, Haoran, et al.
Publicado: (2025)
Improving Video Generation with Human Feedback
por: Liu, Jie, et al.
Publicado: (2025)
por: Liu, Jie, et al.
Publicado: (2025)
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
por: Ju, Xuan, et al.
Publicado: (2025)
por: Ju, Xuan, et al.
Publicado: (2025)
CamCloneMaster: Enabling Reference-based Camera Control for Video Generation
por: Luo, Yawen, et al.
Publicado: (2025)
por: Luo, Yawen, et al.
Publicado: (2025)
A Survey of Interactive Generative Video
por: Yu, Jiwen, et al.
Publicado: (2025)
por: Yu, Jiwen, et al.
Publicado: (2025)
Flow-GRPO: Training Flow Matching Models via Online RL
por: Liu, Jie, et al.
Publicado: (2025)
por: Liu, Jie, et al.
Publicado: (2025)
UNIC: Unified In-Context Video Editing
por: Ye, Zixuan, et al.
Publicado: (2025)
por: Ye, Zixuan, et al.
Publicado: (2025)
RelightMaster: Precise Video Relighting with Multi-plane Light Images
por: Bian, Weikang, et al.
Publicado: (2025)
por: Bian, Weikang, et al.
Publicado: (2025)
UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation
por: Xu, Yiyan, et al.
Publicado: (2026)
por: Xu, Yiyan, et al.
Publicado: (2026)
CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation
por: Wang, Qinghe, et al.
Publicado: (2025)
por: Wang, Qinghe, et al.
Publicado: (2025)
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
por: He, Xuanhua, et al.
Publicado: (2025)
por: He, Xuanhua, et al.
Publicado: (2025)
Video-T1: Test-Time Scaling for Video Generation
por: Liu, Fangfu, et al.
Publicado: (2025)
por: Liu, Fangfu, et al.
Publicado: (2025)
VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
por: Cai, Minghong, et al.
Publicado: (2025)
por: Cai, Minghong, et al.
Publicado: (2025)
TimeSearch: Hierarchical Video Search with Spotlight and Reflection for Human-like Long Video Understanding
por: Pan, Junwen, et al.
Publicado: (2025)
por: Pan, Junwen, et al.
Publicado: (2025)
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
por: Niu, Muyao, et al.
Publicado: (2024)
por: Niu, Muyao, et al.
Publicado: (2024)
Imbalance in Balance: Online Concept Balancing in Generation Models
por: Shi, Yukai, et al.
Publicado: (2025)
por: Shi, Yukai, et al.
Publicado: (2025)
MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
por: Wang, Qinghe, et al.
Publicado: (2025)
por: Wang, Qinghe, et al.
Publicado: (2025)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
por: Wu, Shengqiong, et al.
Publicado: (2025)
por: Wu, Shengqiong, et al.
Publicado: (2025)
StyleMaster: Stylize Your Video with Artistic Generation and Translation
por: Ye, Zixuan, et al.
Publicado: (2024)
por: Ye, Zixuan, et al.
Publicado: (2024)
GameFactory: Creating New Games with Generative Interactive Videos
por: Yu, Jiwen, et al.
Publicado: (2025)
por: Yu, Jiwen, et al.
Publicado: (2025)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
por: Wang, Qixun, et al.
Publicado: (2025)
por: Wang, Qixun, et al.
Publicado: (2025)
TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
por: Pan, Junwen, et al.
Publicado: (2025)
por: Pan, Junwen, et al.
Publicado: (2025)
SemanticGen: Video Generation in Semantic Space
por: Bai, Jianhong, et al.
Publicado: (2025)
por: Bai, Jianhong, et al.
Publicado: (2025)
Towards Precise Scaling Laws for Video Diffusion Transformers
por: Yin, Yuanyang, et al.
Publicado: (2024)
por: Yin, Yuanyang, et al.
Publicado: (2024)
Retrieval, Refinement, and Ranking for Text-to-Video Generation via Prompt Optimization and Test-Time Scaling
por: Rahman, Zillur, et al.
Publicado: (2026)
por: Rahman, Zillur, et al.
Publicado: (2026)
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
por: Hou, Liang, et al.
Publicado: (2025)
por: Hou, Liang, et al.
Publicado: (2025)
VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
por: Wang, Qunzhong, et al.
Publicado: (2025)
por: Wang, Qunzhong, et al.
Publicado: (2025)
DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers
por: Shi, Minglei, et al.
Publicado: (2025)
por: Shi, Minglei, et al.
Publicado: (2025)
Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
por: Park, Joonhyung, et al.
Publicado: (2025)
por: Park, Joonhyung, et al.
Publicado: (2025)
SketchVideo: Sketch-based Video Generation and Editing
por: Liu, Feng-Lin, et al.
Publicado: (2025)
por: Liu, Feng-Lin, et al.
Publicado: (2025)
3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
por: Fang, Zhixue, et al.
Publicado: (2026)
por: Fang, Zhixue, et al.
Publicado: (2026)
Position: Interactive Generative Video as Next-Generation Game Engine
por: Yu, Jiwen, et al.
Publicado: (2025)
por: Yu, Jiwen, et al.
Publicado: (2025)
VINO: A Unified Visual Generator with Interleaved OmniModal Context
por: Chen, Junyi, et al.
Publicado: (2026)
por: Chen, Junyi, et al.
Publicado: (2026)
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
por: Yang, Zhenhao, et al.
Publicado: (2026)
por: Yang, Zhenhao, et al.
Publicado: (2026)
VideoTetris: Towards Compositional Text-to-Video Generation
por: Tian, Ye, et al.
Publicado: (2024)
por: Tian, Ye, et al.
Publicado: (2024)
Ejemplares similares
-
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
por: Huang, Yuzhou, et al.
Publicado: (2025) -
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
por: Cheng, Junhao, et al.
Publicado: (2026) -
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
por: Wu, Shengqiong, et al.
Publicado: (2025) -
GARDO: Reinforcing Diffusion Models without Reward Hacking
por: He, Haoran, et al.
Publicado: (2025) -
UniVideo: Unified Understanding, Generation, and Editing for Videos
por: Wei, Cong, et al.
Publicado: (2025)