VisionCreator-R1: A Reflection-Enhanced Native Visual-Generation Agentic Model
Fuente:
arXiv
Salvato in:
| Autori principali: | Lai, Jinxiang, Zhao, Wenzhe, Lu, Zexin, Zhang, Hualei, Yang, Qinyu, Quan, Rongwei, Li, Zhimin, Shao, Shuai, Guo, Song, Lu, Qinglin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VisionCreator: A Native Visual-Generation Agentic Model with Understanding, Thinking, Planning and Creation
di: Lai, Jinxiang, et al.
Pubblicazione: (2026)
di: Lai, Jinxiang, et al.
Pubblicazione: (2026)
EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect Creation
di: Yang, Shiyuan, et al.
Pubblicazione: (2026)
di: Yang, Shiyuan, et al.
Pubblicazione: (2026)
What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity
di: Li, Haoxi, et al.
Pubblicazione: (2026)
di: Li, Haoxi, et al.
Pubblicazione: (2026)
Spider: Any-to-Many Multimodal LLM
di: Lai, Jinxiang, et al.
Pubblicazione: (2024)
di: Lai, Jinxiang, et al.
Pubblicazione: (2024)
Stealing Creator's Workflow: A Creator-Inspired Agentic Framework with Iterative Feedback Loop for Improved Scientific Short-form Generation
di: Park, Jong Inn, et al.
Pubblicazione: (2025)
di: Park, Jong Inn, et al.
Pubblicazione: (2025)
VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer
di: Liu, Xinyu, et al.
Pubblicazione: (2025)
di: Liu, Xinyu, et al.
Pubblicazione: (2025)
Large Language Models Reflect the Ideology of their Creators
di: Buyl, Maarten, et al.
Pubblicazione: (2024)
di: Buyl, Maarten, et al.
Pubblicazione: (2024)
HLLM-Creator: Hierarchical LLM-based Personalized Creative Generation
di: Chen, Junyi, et al.
Pubblicazione: (2025)
di: Chen, Junyi, et al.
Pubblicazione: (2025)
Juan Soriano. Creator of Visual Parables
VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping
di: Dong, Haotian, et al.
Pubblicazione: (2025)
di: Dong, Haotian, et al.
Pubblicazione: (2025)
Aperiodic Pupil Fluctuations at Rest Predict Orienting of Visual Attention
di: Rongwei Wang, et al.
Pubblicazione: (2025)
di: Rongwei Wang, et al.
Pubblicazione: (2025)
Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
di: Weng, Fenghua, et al.
Pubblicazione: (2025)
di: Weng, Fenghua, et al.
Pubblicazione: (2025)
Betti numbers of normal edge rings (II)
di: Wang, Zexin, et al.
Pubblicazione: (2025)
di: Wang, Zexin, et al.
Pubblicazione: (2025)
The edge rings of compact graphs
di: Wang, Zexin, et al.
Pubblicazione: (2023)
di: Wang, Zexin, et al.
Pubblicazione: (2023)
The resolutions of generalized co-letterplace ideals and their powers
di: Lu, Dancheng, et al.
Pubblicazione: (2023)
di: Lu, Dancheng, et al.
Pubblicazione: (2023)
Betti numbers of normal edge rings (\bf{I})
di: Wang, Zexin, et al.
Pubblicazione: (2024)
di: Wang, Zexin, et al.
Pubblicazione: (2024)
Betti Numbers of Edge Ideals of Weighted Oriented Crown Graphs
di: Wang, Zexin, et al.
Pubblicazione: (2025)
di: Wang, Zexin, et al.
Pubblicazione: (2025)
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models
di: Niu, Junbo, et al.
Pubblicazione: (2025)
di: Niu, Junbo, et al.
Pubblicazione: (2025)
DeepPresenter: Environment-Grounded Reflection for Agentic Presentation Generation
di: Zheng, Hao, et al.
Pubblicazione: (2026)
di: Zheng, Hao, et al.
Pubblicazione: (2026)
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
di: Wang, Yukun, et al.
Pubblicazione: (2026)
di: Wang, Yukun, et al.
Pubblicazione: (2026)
ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs
di: Zhang, Ben, et al.
Pubblicazione: (2025)
di: Zhang, Ben, et al.
Pubblicazione: (2025)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
di: Zhou, Yuqi, et al.
Pubblicazione: (2025)
di: Zhou, Yuqi, et al.
Pubblicazione: (2025)
GPTDrawer: Enhancing Visual Synthesis through ChatGPT
di: Li, Kun, et al.
Pubblicazione: (2024)
di: Li, Kun, et al.
Pubblicazione: (2024)
Relative Reality
di: Yang, Rongwei
Pubblicazione: (2025)
di: Yang, Rongwei
Pubblicazione: (2025)
SHAP-AAD: DeepSHAP-Guided Channel Reduction for EEG Auditory Attention Detection
di: Salmi, Rayan, et al.
Pubblicazione: (2025)
di: Salmi, Rayan, et al.
Pubblicazione: (2025)
The Informal Labor of Content Creators: Situating Xiaohongshu's Key Opinion Consumers in Relationships to Marketers, Consumer Brands, and the Platform
di: Yi, Huiran, et al.
Pubblicazione: (2024)
di: Yi, Huiran, et al.
Pubblicazione: (2024)
Measuring the Spin of the Galactic Center Supermassive Black Hole with Two Pulsars
di: Hu, Zexin, et al.
Pubblicazione: (2024)
di: Hu, Zexin, et al.
Pubblicazione: (2024)
Fundamental Physics with Pulsars around Sagittarius A$^\star$
di: Shao, Lijing, et al.
Pubblicazione: (2025)
di: Shao, Lijing, et al.
Pubblicazione: (2025)
Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
di: Li, Jiaqi, et al.
Pubblicazione: (2025)
di: Li, Jiaqi, et al.
Pubblicazione: (2025)
Spectral properties of Toeplitz operators with harmonic function symbols on the Bergman space
di: Cui, Puyu, et al.
Pubblicazione: (2025)
di: Cui, Puyu, et al.
Pubblicazione: (2025)
Interference inhibition of multimodal information in digital interfaces and its rule of cognitive processing
di: Junkai Shao, et al.
Pubblicazione: (2024)
di: Junkai Shao, et al.
Pubblicazione: (2024)
FedRIR: Rethinking Information Representation in Federated Learning
di: Huang, Yongqiang, et al.
Pubblicazione: (2025)
di: Huang, Yongqiang, et al.
Pubblicazione: (2025)
NativeTok: Native Visual Tokenization for Improved Image Generation
di: Wu, Bin, et al.
Pubblicazione: (2026)
di: Wu, Bin, et al.
Pubblicazione: (2026)
From Performers to Creators: Understanding Retired Women's Perceptions of Technology-Enhanced Dance Performance
di: Zheng, Danlin, et al.
Pubblicazione: (2026)
di: Zheng, Danlin, et al.
Pubblicazione: (2026)
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
di: Gu, Bohai, et al.
Pubblicazione: (2026)
di: Gu, Bohai, et al.
Pubblicazione: (2026)
Object Referring-Guided Scanpath Prediction with Perception-Enhanced Vision-Language Models
di: Quan, Rong, et al.
Pubblicazione: (2026)
di: Quan, Rong, et al.
Pubblicazione: (2026)
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning
di: Wang, Jianing, et al.
Pubblicazione: (2026)
di: Wang, Jianing, et al.
Pubblicazione: (2026)
SongCreator: Lyrics-based Universal Song Generation
di: Lei, Shun, et al.
Pubblicazione: (2024)
di: Lei, Shun, et al.
Pubblicazione: (2024)
BAMF-SLAM: Bundle Adjusted Multi-Fisheye Visual-Inertial SLAM Using Recurrent Field Transforms
di: Zhang, Wei, et al.
Pubblicazione: (2023)
di: Zhang, Wei, et al.
Pubblicazione: (2023)
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
di: Liu, Xianyang, et al.
Pubblicazione: (2025)
di: Liu, Xianyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VisionCreator: A Native Visual-Generation Agentic Model with Understanding, Thinking, Planning and Creation
di: Lai, Jinxiang, et al.
Pubblicazione: (2026) -
EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect Creation
di: Yang, Shiyuan, et al.
Pubblicazione: (2026) -
What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity
di: Li, Haoxi, et al.
Pubblicazione: (2026) -
Spider: Any-to-Many Multimodal LLM
di: Lai, Jinxiang, et al.
Pubblicazione: (2024) -
Stealing Creator's Workflow: A Creator-Inspired Agentic Framework with Iterative Feedback Loop for Improved Scientific Short-form Generation
di: Park, Jong Inn, et al.
Pubblicazione: (2025)