SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness
Fuente:
arXiv
Salvato in:
| Autori principali: | Qiu, Haiyi, Pan, Kaihang, Li, Jiacheng, Li, Juncheng, Tang, Siliang, Zhuang, Yueting |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
di: Yu, Qifan, et al.
Pubblicazione: (2024)
di: Yu, Qifan, et al.
Pubblicazione: (2024)
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
di: Qiu, Haiyi, et al.
Pubblicazione: (2024)
di: Qiu, Haiyi, et al.
Pubblicazione: (2024)
WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing
di: Pan, Kaihang, et al.
Pubblicazione: (2025)
di: Pan, Kaihang, et al.
Pubblicazione: (2025)
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
di: Pan, Kaihang, et al.
Pubblicazione: (2025)
di: Pan, Kaihang, et al.
Pubblicazione: (2025)
OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
di: Bu, Wendong, et al.
Pubblicazione: (2025)
di: Bu, Wendong, et al.
Pubblicazione: (2025)
Auto-Encoding Morph-Tokens for Multimodal LLM
di: Pan, Kaihang, et al.
Pubblicazione: (2024)
di: Pan, Kaihang, et al.
Pubblicazione: (2024)
Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
di: Pan, Kaihang, et al.
Pubblicazione: (2025)
di: Pan, Kaihang, et al.
Pubblicazione: (2025)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
di: Li, Juncheng, et al.
Pubblicazione: (2023)
di: Li, Juncheng, et al.
Pubblicazione: (2023)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
di: Qin, Bosheng, et al.
Pubblicazione: (2023)
di: Qin, Bosheng, et al.
Pubblicazione: (2023)
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
di: Ge, Zhiqi, et al.
Pubblicazione: (2024)
di: Ge, Zhiqi, et al.
Pubblicazione: (2024)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
di: Fan, Zhaoyu, et al.
Pubblicazione: (2025)
di: Fan, Zhaoyu, et al.
Pubblicazione: (2025)
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
di: Pan, Kaihang, et al.
Pubblicazione: (2024)
di: Pan, Kaihang, et al.
Pubblicazione: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
di: Chow, Wei, et al.
Pubblicazione: (2024)
di: Chow, Wei, et al.
Pubblicazione: (2024)
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
di: Bu, Wendong, et al.
Pubblicazione: (2025)
di: Bu, Wendong, et al.
Pubblicazione: (2025)
OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning
di: Pan, Kaihang, et al.
Pubblicazione: (2026)
di: Pan, Kaihang, et al.
Pubblicazione: (2026)
Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
di: Pan, Kaihang, et al.
Pubblicazione: (2025)
di: Pan, Kaihang, et al.
Pubblicazione: (2025)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
di: Dai, Guangyu, et al.
Pubblicazione: (2025)
di: Dai, Guangyu, et al.
Pubblicazione: (2025)
LASER: Tuning-Free LLM-Driven Attention Control for Efficient Text-conditioned Image-to-Animation
di: Zheng, Haoyu, et al.
Pubblicazione: (2024)
di: Zheng, Haoyu, et al.
Pubblicazione: (2024)
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
di: Zheng, Haoyu, et al.
Pubblicazione: (2025)
di: Zheng, Haoyu, et al.
Pubblicazione: (2025)
InstructSAM: Segment Any Instance with Any Instructions
di: Yuan, Yuqian, et al.
Pubblicazione: (2026)
di: Yuan, Yuqian, et al.
Pubblicazione: (2026)
T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from Text
di: Yin, Aoxiong, et al.
Pubblicazione: (2024)
di: Yin, Aoxiong, et al.
Pubblicazione: (2024)
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
di: Qian, Long, et al.
Pubblicazione: (2024)
di: Qian, Long, et al.
Pubblicazione: (2024)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
di: Miao, Bingchen, et al.
Pubblicazione: (2024)
di: Miao, Bingchen, et al.
Pubblicazione: (2024)
Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
di: Yu, Qifan, et al.
Pubblicazione: (2024)
di: Yu, Qifan, et al.
Pubblicazione: (2024)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
di: Wang, Wei, et al.
Pubblicazione: (2026)
di: Wang, Wei, et al.
Pubblicazione: (2026)
HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
di: Yu, Qifan, et al.
Pubblicazione: (2023)
di: Yu, Qifan, et al.
Pubblicazione: (2023)
De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
di: Gao, Minghe, et al.
Pubblicazione: (2023)
di: Gao, Minghe, et al.
Pubblicazione: (2023)
Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
di: Gao, Minghe, et al.
Pubblicazione: (2025)
di: Gao, Minghe, et al.
Pubblicazione: (2025)
Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales
di: Gao, Minghe, et al.
Pubblicazione: (2024)
di: Gao, Minghe, et al.
Pubblicazione: (2024)
MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation
di: Zheng, Haoyu, et al.
Pubblicazione: (2024)
di: Zheng, Haoyu, et al.
Pubblicazione: (2024)
CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents
di: Wang, Keyu, et al.
Pubblicazione: (2026)
di: Wang, Keyu, et al.
Pubblicazione: (2026)
Ask Questions with Double Hints: Visual Question Generation with Answer-awareness and Region-reference
di: Shen, Kai, et al.
Pubblicazione: (2024)
di: Shen, Kai, et al.
Pubblicazione: (2024)
SpatialFusion additional data
di: Yates, Josephine
Pubblicazione: (2026)
di: Yates, Josephine
Pubblicazione: (2026)
3D Dynamics-Aware Manipulation: Endowing Manipulation Policies with 3D Foresight
di: He, Yuxin, et al.
Pubblicazione: (2025)
di: He, Yuxin, et al.
Pubblicazione: (2025)
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
di: Lin, Tianwei, et al.
Pubblicazione: (2025)
di: Lin, Tianwei, et al.
Pubblicazione: (2025)
Towards Physically Executable 3D Gaussian for Embodied Navigation
di: Miao, Bingchen, et al.
Pubblicazione: (2025)
di: Miao, Bingchen, et al.
Pubblicazione: (2025)
Unified Multimodal Coherent Field: Synchronous Semantic-Spatial-Vision Fusion for Brain Tumor Segmentation
di: Zhang, Mingda, et al.
Pubblicazione: (2025)
di: Zhang, Mingda, et al.
Pubblicazione: (2025)
UniFS: Unified Multi-Contrast MRI Reconstruction via Frequency-Spatial Fusion
di: Li, Jialin, et al.
Pubblicazione: (2025)
di: Li, Jialin, et al.
Pubblicazione: (2025)
MV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D Generation
di: Li, Baicheng, et al.
Pubblicazione: (2026)
di: Li, Baicheng, et al.
Pubblicazione: (2026)
Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models
di: Pan, Jiadong, et al.
Pubblicazione: (2026)
di: Pan, Jiadong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
di: Yu, Qifan, et al.
Pubblicazione: (2024) -
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
di: Qiu, Haiyi, et al.
Pubblicazione: (2024) -
WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing
di: Pan, Kaihang, et al.
Pubblicazione: (2025) -
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
di: Pan, Kaihang, et al.
Pubblicazione: (2025) -
OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
di: Bu, Wendong, et al.
Pubblicazione: (2025)