Enhancing Spatial Understanding in Image Generation via Reward Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Zhenyu, Feng, Chaoran, Deng, Yufan, Wu, Jie, Li, Xiaojie, Wang, Rui, Chen, Yunpeng, Zhou, Daquan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HumanNet: Scaling Human-centric Video Learning to One Million Hours
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
Rethinking Video Generation Model for the Embodied World
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle
von: Tang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Tang, Zhenyu, et al.
Veröffentlicht: (2024)
GS2E: Gaussian Splatting is an Effective Data Generator for Event Stream Generation
von: Li, Yuchen, et al.
Veröffentlicht: (2025)
von: Li, Yuchen, et al.
Veröffentlicht: (2025)
SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
von: Zhou, Sashuai, et al.
Veröffentlicht: (2026)
von: Zhou, Sashuai, et al.
Veröffentlicht: (2026)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
von: Zhou, Yupeng, et al.
Veröffentlicht: (2024)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2024)
Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models
von: Pan, Jiadong, et al.
Veröffentlicht: (2026)
von: Pan, Jiadong, et al.
Veröffentlicht: (2026)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
von: Zhang, Kewei, et al.
Veröffentlicht: (2026)
von: Zhang, Kewei, et al.
Veröffentlicht: (2026)
DiffHarmony: Latent Diffusion Model Meets Image Harmonization
von: Zhou, Pengfei, et al.
Veröffentlicht: (2024)
von: Zhou, Pengfei, et al.
Veröffentlicht: (2024)
Spatial Preference Rewarding for MLLMs Spatial Understanding
von: Qiu, Han, et al.
Veröffentlicht: (2025)
von: Qiu, Han, et al.
Veröffentlicht: (2025)
TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation
von: Zhang, Hongyu, et al.
Veröffentlicht: (2026)
von: Zhang, Hongyu, et al.
Veröffentlicht: (2026)
SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning
von: Long, Yancheng, et al.
Veröffentlicht: (2026)
von: Long, Yancheng, et al.
Veröffentlicht: (2026)
ID-Aligner: Enhancing Identity-Preserving Text-to-Image Generation with Reward Feedback Learning
von: Chen, Weifeng, et al.
Veröffentlicht: (2024)
von: Chen, Weifeng, et al.
Veröffentlicht: (2024)
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
von: Fu, Yiyang, et al.
Veröffentlicht: (2026)
von: Fu, Yiyang, et al.
Veröffentlicht: (2026)
SpatiaLoc: Leveraging Multi-Level Spatial Enhanced Descriptors for Cross-Modal Localization
von: Shang, Tianyi, et al.
Veröffentlicht: (2026)
von: Shang, Tianyi, et al.
Veröffentlicht: (2026)
EvaGaussians: Event Stream Assisted Gaussian Splatting from Blurry Images
von: Yu, Wangbo, et al.
Veröffentlicht: (2024)
von: Yu, Wangbo, et al.
Veröffentlicht: (2024)
OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning
von: Gong, Yuan, et al.
Veröffentlicht: (2025)
von: Gong, Yuan, et al.
Veröffentlicht: (2025)
Video Generation Models Are Good Latent Reward Models
von: Mi, Xiaoyue, et al.
Veröffentlicht: (2025)
von: Mi, Xiaoyue, et al.
Veröffentlicht: (2025)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2024)
von: Zhou, Shijie, et al.
Veröffentlicht: (2024)
Multi-Scale Diffusion: Enhancing Spatial Layout in High-Resolution Panoramic Image Generation
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2024)
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
von: Chen, Jiahua, et al.
Veröffentlicht: (2026)
von: Chen, Jiahua, et al.
Veröffentlicht: (2026)
Unified Reward Model for Multimodal Understanding and Generation
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
Understanding Reward Hacking in Text-to-Image Reinforcement Learning
von: Hong, Yunqi, et al.
Veröffentlicht: (2026)
von: Hong, Yunqi, et al.
Veröffentlicht: (2026)
Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
von: Li, Yueyan, et al.
Veröffentlicht: (2025)
von: Li, Yueyan, et al.
Veröffentlicht: (2025)
UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning
von: Tian, Rui, et al.
Veröffentlicht: (2025)
von: Tian, Rui, et al.
Veröffentlicht: (2025)
Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment
von: Ba, Ying, et al.
Veröffentlicht: (2025)
von: Ba, Ying, et al.
Veröffentlicht: (2025)
TwinDiffusion: Enhancing Coherence and Efficiency in Panoramic Image Generation with Diffusion Models
von: Zhou, Teng, et al.
Veröffentlicht: (2024)
von: Zhou, Teng, et al.
Veröffentlicht: (2024)
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation
von: Chen, Wei, et al.
Veröffentlicht: (2026)
von: Chen, Wei, et al.
Veröffentlicht: (2026)
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models
von: Zhang, Gaoyang, et al.
Veröffentlicht: (2024)
von: Zhang, Gaoyang, et al.
Veröffentlicht: (2024)
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
von: Zhang, Xuying, et al.
Veröffentlicht: (2025)
von: Zhang, Xuying, et al.
Veröffentlicht: (2025)
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
von: Xia, Zhongyu, et al.
Veröffentlicht: (2026)
von: Xia, Zhongyu, et al.
Veröffentlicht: (2026)
The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
von: Mao, Weijia, et al.
Veröffentlicht: (2025)
von: Mao, Weijia, et al.
Veröffentlicht: (2025)
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
von: Liang, Huizhi, et al.
Veröffentlicht: (2026)
von: Liang, Huizhi, et al.
Veröffentlicht: (2026)
CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models
von: Hao, Xiangzhao, et al.
Veröffentlicht: (2026)
von: Hao, Xiangzhao, et al.
Veröffentlicht: (2026)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
von: Xu, Lin, et al.
Veröffentlicht: (2024)
von: Xu, Lin, et al.
Veröffentlicht: (2024)
AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
von: Zhou, Ziyin, et al.
Veröffentlicht: (2025)
von: Zhou, Ziyin, et al.
Veröffentlicht: (2025)
AE-NeRF: Augmenting Event-Based Neural Radiance Fields for Non-ideal Conditions and Larger Scene
von: Feng, Chaoran, et al.
Veröffentlicht: (2025)
von: Feng, Chaoran, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HumanNet: Scaling Human-centric Video Learning to One Million Hours
von: Deng, Yufan, et al.
Veröffentlicht: (2026) -
Rethinking Video Generation Model for the Embodied World
von: Deng, Yufan, et al.
Veröffentlicht: (2026) -
Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle
von: Tang, Zhenyu, et al.
Veröffentlicht: (2024) -
GS2E: Gaussian Splatting is an Effective Data Generator for Event Stream Generation
von: Li, Yuchen, et al.
Veröffentlicht: (2025) -
SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
von: Zhou, Sashuai, et al.
Veröffentlicht: (2026)