Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Lyu, Yuanhuiyi, Wong, Chi Kit, Liao, Chenfei, Jiang, Lutao, Zheng, Xu, Lu, Zexin, Zhang, Linfeng, Hu, Xuming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next
by: Pan, Ye, et al.
Published: (2026)
by: Pan, Ye, et al.
Published: (2026)
RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive Learning
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision
by: Lyu, Yuanhuiyi, et al.
Published: (2026)
by: Lyu, Yuanhuiyi, et al.
Published: (2026)
BrightDreamer: Generic 3D Gaussian Generative Framework for Fast Text-to-3D Synthesis
by: Jiang, Lutao, et al.
Published: (2024)
by: Jiang, Lutao, et al.
Published: (2024)
MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
OmniSAM: Omnidirectional Segment Anything Model for UDA in Panoramic Semantic Segmentation
by: Zhong, Ding, et al.
Published: (2025)
by: Zhong, Ding, et al.
Published: (2025)
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
by: Liao, Chenfei, et al.
Published: (2025)
by: Liao, Chenfei, et al.
Published: (2025)
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
by: Li, Jungang, et al.
Published: (2024)
by: Li, Jungang, et al.
Published: (2024)
MemorySAM: Memorize Modalities and Semantics with Segment Anything Model 2 for Multi-modal Semantic Segmentation
by: Liao, Chenfei, et al.
Published: (2025)
by: Liao, Chenfei, et al.
Published: (2025)
EventBind: Learning a Unified Representation to Bind Them All for Event-based Open-world Understanding
by: Zhou, Jiazhou, et al.
Published: (2023)
by: Zhou, Jiazhou, et al.
Published: (2023)
PANORAMA: The Rise of Omnidirectional Vision in the Embodied AI Era
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object Detection
by: Zhou, Jiazhou, et al.
Published: (2025)
by: Zhou, Jiazhou, et al.
Published: (2025)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
Image Anything: Towards Reasoning-coherent and Training-free Multi-modal Image Generation
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
by: Zou, Xin, et al.
Published: (2025)
by: Zou, Xin, et al.
Published: (2025)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
by: Jiang, Jingjing, et al.
Published: (2025)
by: Jiang, Jingjing, et al.
Published: (2025)
Unified Personalized Understanding, Generating and Editing
by: Zhong, Yu, et al.
Published: (2026)
by: Zhong, Yu, et al.
Published: (2026)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
by: Qu, Liao, et al.
Published: (2024)
by: Qu, Liao, et al.
Published: (2024)
CompoNeRF: Text-guided Multi-object Compositional NeRF with Editable 3D Scene Layout
by: Bai, Haotian, et al.
Published: (2023)
by: Bai, Haotian, et al.
Published: (2023)
UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
AI for Service: Proactive Assistance with AI Glasses
by: Wen, Zichen, et al.
Published: (2025)
by: Wen, Zichen, et al.
Published: (2025)
MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation
by: Wang, Hongpeng, et al.
Published: (2026)
by: Wang, Hongpeng, et al.
Published: (2026)
Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
A General Framework to Boost 3D GS Initialization for Text-to-3D Generation by Lexical Richness
by: Jiang, Lutao, et al.
Published: (2024)
by: Jiang, Lutao, et al.
Published: (2024)
Omni-Weather: A Unified Multimodal Model for Weather Radar Understanding and Generation
by: Zhou, Zhiwang, et al.
Published: (2025)
by: Zhou, Zhiwang, et al.
Published: (2025)
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
by: Zou, Jialv, et al.
Published: (2025)
by: Zou, Jialv, et al.
Published: (2025)
Learning Modality-agnostic Representation for Semantic Segmentation from Any Modalities
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
by: Lu, Yanzuo, et al.
Published: (2025)
by: Lu, Yanzuo, et al.
Published: (2025)
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
by: Xu, Chenkai, et al.
Published: (2025)
by: Xu, Chenkai, et al.
Published: (2025)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models
by: Pan, Jiadong, et al.
Published: (2026)
by: Pan, Jiadong, et al.
Published: (2026)
Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding
by: Jiang, Yibo, et al.
Published: (2026)
by: Jiang, Yibo, et al.
Published: (2026)
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
by: Tian, Rui, et al.
Published: (2025)
by: Tian, Rui, et al.
Published: (2025)
EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
by: He, Xin, et al.
Published: (2025)
by: He, Xin, et al.
Published: (2025)
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
by: Zhao, Shanshan, et al.
Published: (2025)
by: Zhao, Shanshan, et al.
Published: (2025)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
DiMeR: Disentangled Mesh Reconstruction Model
by: Jiang, Lutao, et al.
Published: (2025)
by: Jiang, Lutao, et al.
Published: (2025)
Similar Items
-
EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next
by: Pan, Ye, et al.
Published: (2026) -
RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive Learning
by: Lyu, Yuanhuiyi, et al.
Published: (2025) -
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
by: Zheng, Xu, et al.
Published: (2025) -
StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision
by: Lyu, Yuanhuiyi, et al.
Published: (2026) -
BrightDreamer: Generic 3D Gaussian Generative Framework for Fast Text-to-3D Synthesis
by: Jiang, Lutao, et al.
Published: (2024)