Guardado en:
| Autores principales: | Wang, Yikai, Wang, Zhouxia, Wu, Zhonghua, Tao, Qingyi, Liao, Kang, Loy, Chen Change |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2508.12811 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Denoising as Adaptation: Noise-Space Domain Adaptation for Image Restoration
por: Liao, Kang, et al.
Publicado: (2024)
por: Liao, Kang, et al.
Publicado: (2024)
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
por: Xie, Jiahao, et al.
Publicado: (2023)
por: Xie, Jiahao, et al.
Publicado: (2023)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
por: Wang, Zhouxia, et al.
Publicado: (2023)
por: Wang, Zhouxia, et al.
Publicado: (2023)
ObjCtrl-2.5D: Training-free Object Control with Camera Poses
por: Wang, Zhouxia, et al.
Publicado: (2024)
por: Wang, Zhouxia, et al.
Publicado: (2024)
MOWA: Multiple-in-One Image Warping Model
por: Liao, Kang, et al.
Publicado: (2024)
por: Liao, Kang, et al.
Publicado: (2024)
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
por: Liao, Kang, et al.
Publicado: (2025)
por: Liao, Kang, et al.
Publicado: (2025)
SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer
por: Gong, Zerui, et al.
Publicado: (2025)
por: Gong, Zerui, et al.
Publicado: (2025)
MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention
por: Wang, Yuhan, et al.
Publicado: (2025)
por: Wang, Yuhan, et al.
Publicado: (2025)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
por: Wu, Size, et al.
Publicado: (2025)
por: Wu, Size, et al.
Publicado: (2025)
VLANeXt: Recipes for Building Strong VLA Models
por: Wu, Xiao-Ming, et al.
Publicado: (2026)
por: Wu, Xiao-Ming, et al.
Publicado: (2026)
Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models
por: Miao, Yanting, et al.
Publicado: (2026)
por: Miao, Yanting, et al.
Publicado: (2026)
Unified Lexical Representation for Interpretable Visual-Language Alignment
por: Li, Yifan, et al.
Publicado: (2024)
por: Li, Yifan, et al.
Publicado: (2024)
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
por: Wu, Size, et al.
Publicado: (2025)
por: Wu, Size, et al.
Publicado: (2025)
Meta-CoT: Enhancing Granularity and Generalization in Image Editing
por: Zhang, Shiyi, et al.
Publicado: (2026)
por: Zhang, Shiyi, et al.
Publicado: (2026)
Small Scale Data-Free Knowledge Distillation
por: Liu, He, et al.
Publicado: (2024)
por: Liu, He, et al.
Publicado: (2024)
SpecPL: Disentangling Spectral Granularity for Prompt Learning
por: Zhou, Jingtao, et al.
Publicado: (2026)
por: Zhou, Jingtao, et al.
Publicado: (2026)
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
por: Yu, Junwei, et al.
Publicado: (2025)
por: Yu, Junwei, et al.
Publicado: (2025)
Contextual Object Detection with Multimodal Large Language Models
por: Zang, Yuhang, et al.
Publicado: (2023)
por: Zang, Yuhang, et al.
Publicado: (2023)
High-Resolution Image Synthesis via Next-Token Prediction
por: Chen, Dengsheng, et al.
Publicado: (2024)
por: Chen, Dengsheng, et al.
Publicado: (2024)
Visual Generation Without Guidance
por: Chen, Huayu, et al.
Publicado: (2025)
por: Chen, Huayu, et al.
Publicado: (2025)
Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization
por: Gallici, Matteo, et al.
Publicado: (2025)
por: Gallici, Matteo, et al.
Publicado: (2025)
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
por: Tang, Haotian, et al.
Publicado: (2024)
por: Tang, Haotian, et al.
Publicado: (2024)
PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
por: Meng, Ziqiao, et al.
Publicado: (2025)
por: Meng, Ziqiao, et al.
Publicado: (2025)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
por: Li, Liupeng, et al.
Publicado: (2026)
por: Li, Liupeng, et al.
Publicado: (2026)
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
por: Li, Yuming, et al.
Publicado: (2025)
por: Li, Yuming, et al.
Publicado: (2025)
Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
por: Wang, Xiao, et al.
Publicado: (2025)
por: Wang, Xiao, et al.
Publicado: (2025)
AEGPO: Adaptive Entropy-Guided Policy Optimization for Diffusion Models
por: Li, Yuming, et al.
Publicado: (2026)
por: Li, Yuming, et al.
Publicado: (2026)
What Happens Next? Anticipating Future Motion by Generating Point Trajectories
por: Boduljak, Gabrijel, et al.
Publicado: (2025)
por: Boduljak, Gabrijel, et al.
Publicado: (2025)
RepAct: The Re-parameterizable Adaptive Activation Function
por: Wu, Xian, et al.
Publicado: (2024)
por: Wu, Xian, et al.
Publicado: (2024)
AutoGeo: Automating Geometric Image Dataset Creation for Enhanced Geometry Understanding
por: Huang, Zihan, et al.
Publicado: (2024)
por: Huang, Zihan, et al.
Publicado: (2024)
Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification
por: Tan, Zhaorui, et al.
Publicado: (2024)
por: Tan, Zhaorui, et al.
Publicado: (2024)
SpectralAR: Spectral Autoregressive Visual Generation
por: Huang, Yuanhui, et al.
Publicado: (2025)
por: Huang, Yuanhui, et al.
Publicado: (2025)
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
por: Wang, Yulin, et al.
Publicado: (2024)
por: Wang, Yulin, et al.
Publicado: (2024)
MatAnyone: Stable Video Matting with Consistent Memory Propagation
por: Yang, Peiqing, et al.
Publicado: (2025)
por: Yang, Peiqing, et al.
Publicado: (2025)
OmniPrism: Learning Disentangled Visual Concept for Image Generation
por: Li, Yangyang, et al.
Publicado: (2024)
por: Li, Yangyang, et al.
Publicado: (2024)
Jodi: Unification of Visual Generation and Understanding via Joint Modeling
por: Xu, Yifeng, et al.
Publicado: (2025)
por: Xu, Yifeng, et al.
Publicado: (2025)
IMTS is Worth Time $\times$ Channel Patches: Visual Masked Autoencoders for Irregular Multivariate Time Series Prediction
por: Hu, Zhangyi, et al.
Publicado: (2025)
por: Hu, Zhangyi, et al.
Publicado: (2025)
Generalizing Consistency Policy to Visual RL with Prioritized Proximal Experience Regularization
por: Li, Haoran, et al.
Publicado: (2024)
por: Li, Haoran, et al.
Publicado: (2024)
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
por: You, Zuyao, et al.
Publicado: (2025)
por: You, Zuyao, et al.
Publicado: (2025)
Pre-Training Meta-Rule Selection Policy for Visual Generative Abductive Learning
por: Jin, Yu, et al.
Publicado: (2025)
por: Jin, Yu, et al.
Publicado: (2025)
Ejemplares similares
-
Denoising as Adaptation: Noise-Space Domain Adaptation for Image Restoration
por: Liao, Kang, et al.
Publicado: (2024) -
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
por: Xie, Jiahao, et al.
Publicado: (2023) -
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
por: Wang, Zhouxia, et al.
Publicado: (2023) -
ObjCtrl-2.5D: Training-free Object Control with Camera Poses
por: Wang, Zhouxia, et al.
Publicado: (2024) -
MOWA: Multiple-in-One Image Warping Model
por: Liao, Kang, et al.
Publicado: (2024)