HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Xuerui, Cui, Yutao, Zhang, Guozhen, Li, Junzhe, Hu, JiaKui, Zhang, Xiao, Li, Yang, Liu, Songtao, Yang, Miles, Shi, Yu, Zhong, Zhao, Bo, Liefeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
by: Li, Junzhe, et al.
Published: (2025)
by: Li, Junzhe, et al.
Published: (2025)
Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations
by: Zhao, Yaqi, et al.
Published: (2026)
by: Zhao, Yaqi, et al.
Published: (2026)
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
by: Liu, Xiangyue, et al.
Published: (2026)
by: Liu, Xiangyue, et al.
Published: (2026)
High-Performance Temporal Reversible Spiking Neural Networks with $O(L)$ Training Memory and $O(1)$ Inference Cost
by: Hu, JiaKui, et al.
Published: (2024)
by: Hu, JiaKui, et al.
Published: (2024)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
by: Jiao, Yang, et al.
Published: (2025)
by: Jiao, Yang, et al.
Published: (2025)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
by: Shan, Sizhe, et al.
Published: (2025)
by: Shan, Sizhe, et al.
Published: (2025)
Bridging Degradation Discrimination and Generation for Universal Image Restoration
by: Hu, JiaKui, et al.
Published: (2026)
by: Hu, JiaKui, et al.
Published: (2026)
Universal Image Restoration Pre-training via Degradation Classification
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model
by: Li, Junzhe, et al.
Published: (2025)
by: Li, Junzhe, et al.
Published: (2025)
Auto-Regressively Generating Multi-View Consistent Images
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
Towards Fine-grained Interactive Segmentation in Images and Videos
by: Yao, Yuan, et al.
Published: (2025)
by: Yao, Yuan, et al.
Published: (2025)
MoSAM: Motion-Guided Segment Anything Model with Spatial-Temporal Memory Selection
by: Yang, Qiushi, et al.
Published: (2025)
by: Yang, Qiushi, et al.
Published: (2025)
When Graph Traversal Meets Structured Preferences: Unified Framework and Complexity Results
by: Rong, Guozhen, et al.
Published: (2026)
by: Rong, Guozhen, et al.
Published: (2026)
Enhancing Image Restoration Transformer via Adaptive Translation Equivariance
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
Universal Image Restoration Pre-training via Masked Degradation Classification
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
by: Zhang, Xin, et al.
Published: (2023)
by: Zhang, Xin, et al.
Published: (2023)
Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context
by: Hu, JiaKui, et al.
Published: (2026)
by: Hu, JiaKui, et al.
Published: (2026)
UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization
by: He, Junjie, et al.
Published: (2024)
by: He, Junjie, et al.
Published: (2024)
HYDRA: Model Factorization Framework for Black-Box LLM Personalization
by: Zhuang, Yuchen, et al.
Published: (2024)
by: Zhuang, Yuchen, et al.
Published: (2024)
Skin Tokens: A Learned Compact Representation for Unified Autoregressive Rigging
by: Zhang, Jia-peng, et al.
Published: (2026)
by: Zhang, Jia-peng, et al.
Published: (2026)
Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation
by: Chen, Yingjie, et al.
Published: (2025)
by: Chen, Yingjie, et al.
Published: (2025)
Precise: SDE-Consistent Stochastic Sampling for RL Post-Training of Flow-Matching Models
by: Zou, Jade, et al.
Published: (2026)
by: Zou, Jade, et al.
Published: (2026)
4DPC$^2$hat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping
by: Zhang, Xindan, et al.
Published: (2026)
by: Zhang, Xindan, et al.
Published: (2026)
GaussianIP: Identity-Preserving Realistic 3D Human Generation via Human-Centric Diffusion Prior
by: Tang, Zichen, et al.
Published: (2025)
by: Tang, Zichen, et al.
Published: (2025)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Some Mizohata-Takeuchi-type estimate for exponential sums
by: Yang, Xuerui
Published: (2025)
by: Yang, Xuerui
Published: (2025)
AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation
by: He, Junjie, et al.
Published: (2025)
by: He, Junjie, et al.
Published: (2025)
Multi-modal Relation Distillation for Unified 3D Representation Learning
by: Wang, Huiqun, et al.
Published: (2024)
by: Wang, Huiqun, et al.
Published: (2024)
Sheet as Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding
by: Lei, Yiming, et al.
Published: (2026)
by: Lei, Yiming, et al.
Published: (2026)
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2025)
by: Fan, Lijie, et al.
Published: (2025)
Soft-Prompting with Graph-of-Thought for Multi-modal Representation Learning
by: Yang, Juncheng, et al.
Published: (2024)
by: Yang, Juncheng, et al.
Published: (2024)
StableDrag: Stable Dragging for Point-based Image Editing
by: Cui, Yutao, et al.
Published: (2024)
by: Cui, Yutao, et al.
Published: (2024)
Motion-Aware Generative Frame Interpolation
by: Zhang, Guozhen, et al.
Published: (2025)
by: Zhang, Guozhen, et al.
Published: (2025)
VFIMamba: Video Frame Interpolation with State Space Models
by: Zhang, Guozhen, et al.
Published: (2024)
by: Zhang, Guozhen, et al.
Published: (2024)
ToolTok: Tool Tokenization for Efficient and Generalizable GUI Agents
by: Wang, Xiaoce, et al.
Published: (2026)
by: Wang, Xiaoce, et al.
Published: (2026)
UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
by: An, Ruichuan, et al.
Published: (2025)
by: An, Ruichuan, et al.
Published: (2025)
Understanding Representation Learnability of Nonlinear Self-Supervised Learning
by: Yang, Ruofeng, et al.
Published: (2024)
by: Yang, Ruofeng, et al.
Published: (2024)
EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing
by: Zhang, Wei, et al.
Published: (2024)
by: Zhang, Wei, et al.
Published: (2024)
Similar Items
-
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
by: Li, Junzhe, et al.
Published: (2025) -
Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
by: Hu, JiaKui, et al.
Published: (2025) -
UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations
by: Zhao, Yaqi, et al.
Published: (2026) -
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
by: Liu, Xiangyue, et al.
Published: (2026) -
High-Performance Temporal Reversible Spiking Neural Networks with $O(L)$ Training Memory and $O(1)$ Inference Cost
by: Hu, JiaKui, et al.
Published: (2024)