Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Size, Zhang, Wenwei, Xu, Lumin, Jin, Sheng, Wu, Zhonghua, Tao, Qingyi, Liu, Wentao, Li, Wei, Loy, Chen Change |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
di: Wu, Size, et al.
Pubblicazione: (2025)
di: Wu, Size, et al.
Pubblicazione: (2025)
F-LMM: Grounding Frozen Large Multimodal Models
di: Wu, Size, et al.
Pubblicazione: (2024)
di: Wu, Size, et al.
Pubblicazione: (2024)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
di: Wu, Size, et al.
Pubblicazione: (2023)
di: Wu, Size, et al.
Pubblicazione: (2023)
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
di: Liao, Kang, et al.
Pubblicazione: (2025)
di: Liao, Kang, et al.
Pubblicazione: (2025)
Next Visual Granularity Generation
di: Wang, Yikai, et al.
Pubblicazione: (2025)
di: Wang, Yikai, et al.
Pubblicazione: (2025)
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
di: Xu, Shilin, et al.
Pubblicazione: (2023)
di: Xu, Shilin, et al.
Pubblicazione: (2023)
SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer
di: Gong, Zerui, et al.
Pubblicazione: (2025)
di: Gong, Zerui, et al.
Pubblicazione: (2025)
Controllable Human-centric Keyframe Interpolation with Generative Prior
di: Guo, Zujin, et al.
Pubblicazione: (2025)
di: Guo, Zujin, et al.
Pubblicazione: (2025)
MOWA: Multiple-in-One Image Warping Model
di: Liao, Kang, et al.
Pubblicazione: (2024)
di: Liao, Kang, et al.
Pubblicazione: (2024)
OMG-Seg: Is One Model Good Enough For All Segmentation?
di: Li, Xiangtai, et al.
Pubblicazione: (2024)
di: Li, Xiangtai, et al.
Pubblicazione: (2024)
MatAnyone: Stable Video Matting with Consistent Memory Propagation
di: Yang, Peiqing, et al.
Pubblicazione: (2025)
di: Yang, Peiqing, et al.
Pubblicazione: (2025)
UniFS: Universal Few-shot Instance Perception with Point Representations
di: Jin, Sheng, et al.
Pubblicazione: (2024)
di: Jin, Sheng, et al.
Pubblicazione: (2024)
AITTI: Learning Adaptive Inclusive Token for Text-to-Image Generation
di: Hou, Xinyu, et al.
Pubblicazione: (2024)
di: Hou, Xinyu, et al.
Pubblicazione: (2024)
Enhanced Generative Structure Prior for Chinese Text Image Super-resolution
di: Li, Xiaoming, et al.
Pubblicazione: (2025)
di: Li, Xiaoming, et al.
Pubblicazione: (2025)
Generalizable Implicit Motion Modeling for Video Frame Interpolation
di: Guo, Zujin, et al.
Pubblicazione: (2024)
di: Guo, Zujin, et al.
Pubblicazione: (2024)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
di: Jiao, Yang, et al.
Pubblicazione: (2025)
di: Jiao, Yang, et al.
Pubblicazione: (2025)
Transformer-Based Visual Segmentation: A Survey
di: Li, Xiangtai, et al.
Pubblicazione: (2023)
di: Li, Xiangtai, et al.
Pubblicazione: (2023)
MVIP-NeRF: Multi-view 3D Inpainting on NeRF Scenes via Diffusion Prior
di: Chen, Honghua, et al.
Pubblicazione: (2024)
di: Chen, Honghua, et al.
Pubblicazione: (2024)
TCFormer: Visual Recognition via Token Clustering Transformer
di: Zeng, Wang, et al.
Pubblicazione: (2024)
di: Zeng, Wang, et al.
Pubblicazione: (2024)
DifFace: Blind Face Restoration with Diffused Error Contraction
di: Yue, Zongsheng, et al.
Pubblicazione: (2022)
di: Yue, Zongsheng, et al.
Pubblicazione: (2022)
Contextual Object Detection with Multimodal Large Language Models
di: Zang, Yuhang, et al.
Pubblicazione: (2023)
di: Zang, Yuhang, et al.
Pubblicazione: (2023)
Kalman-Inspired Feature Propagation for Video Face Super-Resolution
di: Feng, Ruicheng, et al.
Pubblicazione: (2024)
di: Feng, Ruicheng, et al.
Pubblicazione: (2024)
Generative Photographic Control for Scene-Consistent Video Cinematic Editing
di: Sun, Huiqiang, et al.
Pubblicazione: (2025)
di: Sun, Huiqiang, et al.
Pubblicazione: (2025)
Control Color: Multimodal Diffusion-based Interactive Image Colorization
di: Liang, Zhexin, et al.
Pubblicazione: (2024)
di: Liang, Zhexin, et al.
Pubblicazione: (2024)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
di: Zhang, Tao, et al.
Pubblicazione: (2024)
di: Zhang, Tao, et al.
Pubblicazione: (2024)
Knowledge Visualization: A Benchmark and Method for Knowledge-Intensive Text-to-Image Generation
di: Zhao, Ran, et al.
Pubblicazione: (2026)
di: Zhao, Ran, et al.
Pubblicazione: (2026)
GKGNet: Group K-Nearest Neighbor based Graph Convolutional Network for Multi-Label Image Recognition
di: Yao, Ruijie, et al.
Pubblicazione: (2023)
di: Yao, Ruijie, et al.
Pubblicazione: (2023)
Learning 3D Garment Animation from Trajectories of A Piece of Cloth
di: Shao, Yidi, et al.
Pubblicazione: (2025)
di: Shao, Yidi, et al.
Pubblicazione: (2025)
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
di: Yang, Jie, et al.
Pubblicazione: (2025)
di: Yang, Jie, et al.
Pubblicazione: (2025)
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
di: Ma, Yiyang, et al.
Pubblicazione: (2024)
di: Ma, Yiyang, et al.
Pubblicazione: (2024)
KptLLM: Unveiling the Power of Large Language Model for Keypoint Comprehension
di: Yang, Jie, et al.
Pubblicazione: (2024)
di: Yang, Jie, et al.
Pubblicazione: (2024)
HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
di: Qiu, Xuerui, et al.
Pubblicazione: (2026)
di: Qiu, Xuerui, et al.
Pubblicazione: (2026)
Sigma: Semantically Informative Pre-training for Skeleton-based Sign Language Understanding
di: Pu, Muxin, et al.
Pubblicazione: (2025)
di: Pu, Muxin, et al.
Pubblicazione: (2025)
ObjCtrl-2.5D: Training-free Object Control with Camera Poses
di: Wang, Zhouxia, et al.
Pubblicazione: (2024)
di: Wang, Zhouxia, et al.
Pubblicazione: (2024)
Denoising as Adaptation: Noise-Space Domain Adaptation for Image Restoration
di: Liao, Kang, et al.
Pubblicazione: (2024)
di: Liao, Kang, et al.
Pubblicazione: (2024)
Omegance: A Single Parameter for Various Granularities in Diffusion-Based Synthesis
di: Hou, Xinyu, et al.
Pubblicazione: (2024)
di: Hou, Xinyu, et al.
Pubblicazione: (2024)
Trans-Adapter: A Plug-and-Play Framework for Transparent Image Inpainting
di: Dai, Yuekun, et al.
Pubblicazione: (2025)
di: Dai, Yuekun, et al.
Pubblicazione: (2025)
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation
di: Yang, Shuai, et al.
Pubblicazione: (2024)
di: Yang, Shuai, et al.
Pubblicazione: (2024)
EdgeSAM: Prompt-In-the-Loop Distillation for SAM
di: Zhou, Chong, et al.
Pubblicazione: (2023)
di: Zhou, Chong, et al.
Pubblicazione: (2023)
MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention
di: Wang, Yuhan, et al.
Pubblicazione: (2025)
di: Wang, Yuhan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
di: Wu, Size, et al.
Pubblicazione: (2025) -
F-LMM: Grounding Frozen Large Multimodal Models
di: Wu, Size, et al.
Pubblicazione: (2024) -
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
di: Wu, Size, et al.
Pubblicazione: (2023) -
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
di: Liao, Kang, et al.
Pubblicazione: (2025) -
Next Visual Granularity Generation
di: Wang, Yikai, et al.
Pubblicazione: (2025)