ROICtrl: Boosting Instance Control for Visual Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Gu, Yuchao, Zhou, Yipin, Ye, Yunfan, Nie, Yixin, Yu, Licheng, Ma, Pingchuan, Lin, Kevin Qinghong, Shou, Mike Zheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
por: Lin, Kevin Qinghong, et al.
Publicado: (2025)
por: Lin, Kevin Qinghong, et al.
Publicado: (2025)
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
por: Hu, Siyuan, et al.
Publicado: (2025)
por: Hu, Siyuan, et al.
Publicado: (2025)
Paper2Video: Automatic Video Generation from Scientific Papers
por: Zhu, Zeyu, et al.
Publicado: (2025)
por: Zhu, Zeyu, et al.
Publicado: (2025)
Code2Video: A Code-centric Paradigm for Educational Video Generation
por: Chen, Yanzhe, et al.
Publicado: (2025)
por: Chen, Yanzhe, et al.
Publicado: (2025)
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
por: Gu, Yuchao, et al.
Publicado: (2025)
por: Gu, Yuchao, et al.
Publicado: (2025)
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
por: Liu, Ye, et al.
Publicado: (2025)
por: Liu, Ye, et al.
Publicado: (2025)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
por: Ouyang, Mingyu, et al.
Publicado: (2026)
por: Ouyang, Mingyu, et al.
Publicado: (2026)
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
por: Wang, Jiaqi, et al.
Publicado: (2025)
por: Wang, Jiaqi, et al.
Publicado: (2025)
Personalized Vision via Visual In-Context Learning
por: Jiang, Yuxin, et al.
Publicado: (2025)
por: Jiang, Yuxin, et al.
Publicado: (2025)
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
por: Xie, Jinheng, et al.
Publicado: (2024)
por: Xie, Jinheng, et al.
Publicado: (2024)
Edit Transfer: Learning Image Editing via Vision In-Context Relations
por: Chen, Lan, et al.
Publicado: (2025)
por: Chen, Lan, et al.
Publicado: (2025)
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
por: Song, Yiren, et al.
Publicado: (2026)
por: Song, Yiren, et al.
Publicado: (2026)
Bootstrapping SparseFormers from Vision Foundation Models
por: Gao, Ziteng, et al.
Publicado: (2023)
por: Gao, Ziteng, et al.
Publicado: (2023)
P-Flow: Prompting Visual Effects Generation
por: Zhao, Rui, et al.
Publicado: (2026)
por: Zhao, Rui, et al.
Publicado: (2026)
Olaf-World: Orienting Latent Actions for Video World Modeling
por: Jiang, Yuxin, et al.
Publicado: (2026)
por: Jiang, Yuxin, et al.
Publicado: (2026)
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
por: Ouyang, Mingyu, et al.
Publicado: (2026)
por: Ouyang, Mingyu, et al.
Publicado: (2026)
Computer-Use Agents as Judges for Generative User Interface
por: Lin, Kevin Qinghong, et al.
Publicado: (2025)
por: Lin, Kevin Qinghong, et al.
Publicado: (2025)
FramePrompt: In-context Controllable Animation with Zero Structural Changes
por: Fang, Guian, et al.
Publicado: (2025)
por: Fang, Guian, et al.
Publicado: (2025)
Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model
por: Mao, Qi, et al.
Publicado: (2025)
por: Mao, Qi, et al.
Publicado: (2025)
COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
por: Wang, Alex Jinpeng, et al.
Publicado: (2024)
por: Wang, Alex Jinpeng, et al.
Publicado: (2024)
GEB+: A Benchmark for Generic Event Boundary Captioning, Grounding and Retrieval
por: Wang, Yuxuan, et al.
Publicado: (2022)
por: Wang, Yuxuan, et al.
Publicado: (2022)
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
por: Wu, Weijia, et al.
Publicado: (2024)
por: Wu, Weijia, et al.
Publicado: (2024)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
por: Lin, Kevin Qinghong, et al.
Publicado: (2024)
por: Lin, Kevin Qinghong, et al.
Publicado: (2024)
Reinforcement Learning for Large Model: A Survey
por: Wu, Weijia, et al.
Publicado: (2025)
por: Wu, Weijia, et al.
Publicado: (2025)
DiffSim: Taming Diffusion Models for Evaluating Visual Similarity
por: Song, Yiren, et al.
Publicado: (2024)
por: Song, Yiren, et al.
Publicado: (2024)
Learning Long-form Video Prior via Generative Pre-Training
por: Xie, Jinheng, et al.
Publicado: (2024)
por: Xie, Jinheng, et al.
Publicado: (2024)
Factorized Visual Tokenization and Generation
por: Bai, Zechen, et al.
Publicado: (2024)
por: Bai, Zechen, et al.
Publicado: (2024)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
por: Lin, Kevin Qinghong, et al.
Publicado: (2024)
por: Lin, Kevin Qinghong, et al.
Publicado: (2024)
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
por: Gu, Yuchao, et al.
Publicado: (2026)
por: Gu, Yuchao, et al.
Publicado: (2026)
Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator
por: Zhao, Henry Hengyuan, et al.
Publicado: (2023)
por: Zhao, Henry Hengyuan, et al.
Publicado: (2023)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
por: Zhang, David Junhao, et al.
Publicado: (2023)
por: Zhang, David Junhao, et al.
Publicado: (2023)
GUI Action Narrator: Where and When Did That Action Take Place?
por: Wu, Qinchen, et al.
Publicado: (2024)
por: Wu, Qinchen, et al.
Publicado: (2024)
DragAnything: Motion Control for Anything using Entity Representation
por: Wu, Weijia, et al.
Publicado: (2024)
por: Wu, Weijia, et al.
Publicado: (2024)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
por: Ilaslan, Muhammet Furkan, et al.
Publicado: (2024)
por: Ilaslan, Muhammet Furkan, et al.
Publicado: (2024)
Learning Video Context as Interleaved Multimodal Sequences
por: Lin, Kevin Qinghong, et al.
Publicado: (2024)
por: Lin, Kevin Qinghong, et al.
Publicado: (2024)
Automated Movie Generation via Multi-Agent CoT Planning
por: Wu, Weijia, et al.
Publicado: (2025)
por: Wu, Weijia, et al.
Publicado: (2025)
DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles
por: Zhao, Rui, et al.
Publicado: (2025)
por: Zhao, Rui, et al.
Publicado: (2025)
MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
por: Song, Yiren, et al.
Publicado: (2025)
por: Song, Yiren, et al.
Publicado: (2025)
TPDiff: Temporal Pyramid Video Diffusion Model
por: Ran, Lingmin, et al.
Publicado: (2025)
por: Ran, Lingmin, et al.
Publicado: (2025)
D-AR: Diffusion via Autoregressive Models
por: Gao, Ziteng, et al.
Publicado: (2025)
por: Gao, Ziteng, et al.
Publicado: (2025)
Ejemplares similares
-
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
por: Lin, Kevin Qinghong, et al.
Publicado: (2025) -
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
por: Hu, Siyuan, et al.
Publicado: (2025) -
Paper2Video: Automatic Video Generation from Scientific Papers
por: Zhu, Zeyu, et al.
Publicado: (2025) -
Code2Video: A Code-centric Paradigm for Educational Video Generation
por: Chen, Yanzhe, et al.
Publicado: (2025) -
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
por: Gu, Yuchao, et al.
Publicado: (2025)