Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Keming, Yang, Zuhao, Zhang, Kaichen, Wang, Shizun, Zhu, Haowei, Leng, Sicong, Yang, Zhongyu, Wang, Qijie, Wang, Sudong, Wang, Ziting, Wang, Zili, Zhang, Hui, Wang, Haonan, Zhou, Hang, Pu, Yifan, Li, Xingxuan, Zhan, Fangneng, Li, Bo, Bing, Lidong, Song, Yuxin, Liu, Ziwei, Chen, Wenhu, Wang, Jingdong, Wang, Xinchao, Qi, Xiaojuan, Lu, Shijian, Wang, Bin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
by: Yang, Zuhao, et al.
Published: (2026)
by: Yang, Zuhao, et al.
Published: (2026)
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
by: Zhang, Kaichen, et al.
Published: (2025)
by: Zhang, Kaichen, et al.
Published: (2025)
WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors
by: Wu, Keming, et al.
Published: (2026)
by: Wu, Keming, et al.
Published: (2026)
PE3R: Perception-Efficient 3D Reconstruction
by: Hu, Jie, et al.
Published: (2025)
by: Hu, Jie, et al.
Published: (2025)
C4D: 4D Made from 3D through Dual Correspondences
by: Wang, Shizun, et al.
Published: (2025)
by: Wang, Shizun, et al.
Published: (2025)
MindBridge: A Cross-Subject Brain Decoding Framework
by: Wang, Shizun, et al.
Published: (2024)
by: Wang, Shizun, et al.
Published: (2024)
GFlow: Recovering 4D World from Monocular Video
by: Wang, Shizun, et al.
Published: (2024)
by: Wang, Shizun, et al.
Published: (2024)
Test3R: Learning to Reconstruct 3D at Test Time
by: Yuan, Yuheng, et al.
Published: (2025)
by: Yuan, Yuheng, et al.
Published: (2025)
Make Geometry Matter for Spatial Reasoning
by: Zhang, Shihua, et al.
Published: (2026)
by: Zhang, Shihua, et al.
Published: (2026)
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
by: Wang, Sudong, et al.
Published: (2026)
by: Wang, Sudong, et al.
Published: (2026)
Hash3D: Training-free Acceleration for 3D Generation
by: Yang, Xingyi, et al.
Published: (2024)
by: Yang, Xingyi, et al.
Published: (2024)
Language Model as Visual Explainer
by: Yang, Xingyi, et al.
Published: (2024)
by: Yang, Xingyi, et al.
Published: (2024)
Neural Metamorphosis
by: Yang, Xingyi, et al.
Published: (2024)
by: Yang, Xingyi, et al.
Published: (2024)
Compositional Video Generation as Flow Equalization
by: Yang, Xingyi, et al.
Published: (2024)
by: Yang, Xingyi, et al.
Published: (2024)
Kolmogorov-Arnold Transformer
by: Yang, Xingyi, et al.
Published: (2024)
by: Yang, Xingyi, et al.
Published: (2024)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
Accelerating Diffusion-based Super-Resolution with Dynamic Time-Spatial Sampling
by: Qin, Rui, et al.
Published: (2025)
by: Qin, Rui, et al.
Published: (2025)
Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
by: Zhou, Kaichen, et al.
Published: (2026)
by: Zhou, Kaichen, et al.
Published: (2026)
CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification
by: Wang, Qijie, et al.
Published: (2024)
by: Wang, Qijie, et al.
Published: (2024)
Distributional Robustness Bounds Generalization Errors
by: Wang, Shixiong, et al.
Published: (2022)
by: Wang, Shixiong, et al.
Published: (2022)
UniDet-D: A Unified Dynamic Spectral Attention Model for Object Detection under Adverse Weathers
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
PAGE-4D: VGGT-4D Perception via Disentangled Pose and Geometry Estimation
by: Zhou, Kaichen, et al.
Published: (2025)
by: Zhou, Kaichen, et al.
Published: (2025)
Guiding Visual Autoregressive Models through Spectrum Weakening
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
FusionTransNet for Smart Urban Mobility: Spatiotemporal Traffic Forecasting Through Multimodal Network Integration
by: Wang, Binwu, et al.
Published: (2024)
by: Wang, Binwu, et al.
Published: (2024)
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
by: Song, Yuxin, et al.
Published: (2025)
by: Song, Yuxin, et al.
Published: (2025)
DiffV2IR: Visible-to-Infrared Diffusion Model via Vision-Language Understanding
by: Ran, Lingyan, et al.
Published: (2025)
by: Ran, Lingyan, et al.
Published: (2025)
Efficient Gaussian Splatting for Monocular Dynamic Scene Rendering via Sparse Time-Variant Attribute Modeling
by: Kong, Hanyang, et al.
Published: (2025)
by: Kong, Hanyang, et al.
Published: (2025)
Don't Forget the Nonlinearity: Unlocking Activation Functions in Efficient Fine-Tuning
by: Yin, Bo, et al.
Published: (2025)
by: Yin, Bo, et al.
Published: (2025)
Unsegment Anything by Simulating Deformation
by: Lu, Jiahao, et al.
Published: (2024)
by: Lu, Jiahao, et al.
Published: (2024)
Relation Rectification in Diffusion Model
by: Wu, Yinwei, et al.
Published: (2024)
by: Wu, Yinwei, et al.
Published: (2024)
FlashSplat: 2D to 3D Gaussian Splatting Segmentation Solved Optimally
by: Shen, Qiuhong, et al.
Published: (2024)
by: Shen, Qiuhong, et al.
Published: (2024)
Manipulating Photogalvanic Effects in Two-Dimensional Multiferroic Breathing Kagome Materials
by: Wang, Haonan, et al.
Published: (2024)
by: Wang, Haonan, et al.
Published: (2024)
Focus on Neighbors and Know the Whole: Towards Consistent Dense Multiview Text-to-Image Generator for 3D Creation
by: Li, Bonan, et al.
Published: (2024)
by: Li, Bonan, et al.
Published: (2024)
ViMU: Benchmarking Video Metaphorical Understanding
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
MambaOut: Do We Really Need Mamba for Vision?
by: Yu, Weihao, et al.
Published: (2024)
by: Yu, Weihao, et al.
Published: (2024)
Neural Lineage
by: Yu, Runpeng, et al.
Published: (2024)
by: Yu, Runpeng, et al.
Published: (2024)
Sponge Tool Attack: Stealthy Denial-of-Efficiency against Tool-Augmented Agentic Reasoning
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
Ungeneralizable Examples
by: Ye, Jingwen, et al.
Published: (2024)
by: Ye, Jingwen, et al.
Published: (2024)
Flipped Classroom: Aligning Teacher Attention with Student in Generalized Category Discovery
by: Lin, Haonan, et al.
Published: (2024)
by: Lin, Haonan, et al.
Published: (2024)
Similar Items
-
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
by: Yang, Zuhao, et al.
Published: (2026) -
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
by: Yang, Zuhao, et al.
Published: (2025) -
OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
by: Zhang, Kaichen, et al.
Published: (2025) -
WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors
by: Wu, Keming, et al.
Published: (2026) -
PE3R: Perception-Efficient 3D Reconstruction
by: Hu, Jie, et al.
Published: (2025)