Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, JiaKui, Zhao, Shanshan, Chen, Qing-Guo, Qiu, Xuerui, Liu, Jialun, Xu, Zhao, Luo, Weihua, Zhang, Kaifu, Lu, Yanye |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Auto-Regressively Generating Multi-View Consistent Images
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
Bridging Degradation Discrimination and Generation for Universal Image Restoration
by: Hu, JiaKui, et al.
Published: (2026)
by: Hu, JiaKui, et al.
Published: (2026)
Universal Image Restoration Pre-training via Degradation Classification
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
Enhancing Image Restoration Transformer via Adaptive Translation Equivariance
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
Universal Image Restoration Pre-training via Masked Degradation Classification
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
by: Qiu, Xuerui, et al.
Published: (2026)
by: Qiu, Xuerui, et al.
Published: (2026)
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
by: Zhao, Shanshan, et al.
Published: (2025)
by: Zhao, Shanshan, et al.
Published: (2025)
Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context
by: Hu, JiaKui, et al.
Published: (2026)
by: Hu, JiaKui, et al.
Published: (2026)
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
by: Duan, Lunhao, et al.
Published: (2024)
by: Duan, Lunhao, et al.
Published: (2024)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
by: Gao, Sensen, et al.
Published: (2025)
by: Gao, Sensen, et al.
Published: (2025)
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
by: Lu, Shiyin, et al.
Published: (2024)
by: Lu, Shiyin, et al.
Published: (2024)
High-Performance Temporal Reversible Spiking Neural Networks with $O(L)$ Training Memory and $O(1)$ Inference Cost
by: Hu, JiaKui, et al.
Published: (2024)
by: Hu, JiaKui, et al.
Published: (2024)
SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
MCGS: Multiview Consistency Enhancement for Sparse-View 3D Gaussian Radiance Fields
by: Xiao, Yuru, et al.
Published: (2024)
by: Xiao, Yuru, et al.
Published: (2024)
Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs
by: Sun, Hui, et al.
Published: (2025)
by: Sun, Hui, et al.
Published: (2025)
Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMs
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Omni-Video: Democratizing Unified Video Understanding and Generation
by: Tan, Zhiyu, et al.
Published: (2025)
by: Tan, Zhiyu, et al.
Published: (2025)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
by: Cui, Tianyu, et al.
Published: (2025)
by: Cui, Tianyu, et al.
Published: (2025)
EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
by: Tian, Wenjie, et al.
Published: (2026)
by: Tian, Wenjie, et al.
Published: (2026)
Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees
by: Chen, Sijia, et al.
Published: (2024)
by: Chen, Sijia, et al.
Published: (2024)
TimeOmni-VL: Unified Models for Time Series Understanding and Generation
by: Guan, Tong, et al.
Published: (2026)
by: Guan, Tong, et al.
Published: (2026)
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
by: Zhao, Shifang, et al.
Published: (2025)
by: Zhao, Shifang, et al.
Published: (2025)
Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
by: Yang, Wenhao, et al.
Published: (2025)
by: Yang, Wenhao, et al.
Published: (2025)
Joint Tensor and Inter-View Low-Rank Recovery for Incomplete Multiview Clustering
by: Wang, Jianyu, et al.
Published: (2025)
by: Wang, Jianyu, et al.
Published: (2025)
Wings: Learning Multimodal LLMs without Text-only Forgetting
by: Zhang, Yi-Kai, et al.
Published: (2024)
by: Zhang, Yi-Kai, et al.
Published: (2024)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
by: Yan, Qianqi, et al.
Published: (2026)
by: Yan, Qianqi, et al.
Published: (2026)
A Joint Sparse Self-Representation Learning Method for Multiview Clustering
by: Jia, Mengxue, et al.
Published: (2025)
by: Jia, Mengxue, et al.
Published: (2025)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
by: Dai, Yifan, et al.
Published: (2026)
by: Dai, Yifan, et al.
Published: (2026)
A Unified Agentic Framework for Evaluating Conditional Image Generation
by: Wang, Jifang, et al.
Published: (2025)
by: Wang, Jifang, et al.
Published: (2025)
LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech
by: Yang, Fei, et al.
Published: (2026)
by: Yang, Fei, et al.
Published: (2026)
A State-Transition Framework for Efficient LLM Reasoning
by: Zhang, Liang, et al.
Published: (2026)
by: Zhang, Liang, et al.
Published: (2026)
Pragmatist: Multiview Conditional Diffusion Models for High-Fidelity 3D Reconstruction from Unposed Sparse Views
by: Zhang, Songchun, et al.
Published: (2024)
by: Zhang, Songchun, et al.
Published: (2024)
Building Decision Making Models Through Language Model Regime
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
OmniFD: A Unified Model for Versatile Face Forgery Detection
by: Liu, Haotian, et al.
Published: (2025)
by: Liu, Haotian, et al.
Published: (2025)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
by: Cheng, Xize, et al.
Published: (2024)
by: Cheng, Xize, et al.
Published: (2024)
Similar Items
-
Auto-Regressively Generating Multi-View Consistent Images
by: Hu, JiaKui, et al.
Published: (2025) -
Bridging Degradation Discrimination and Generation for Universal Image Restoration
by: Hu, JiaKui, et al.
Published: (2026) -
Universal Image Restoration Pre-training via Degradation Classification
by: Hu, JiaKui, et al.
Published: (2025) -
Enhancing Image Restoration Transformer via Adaptive Translation Equivariance
by: Hu, JiaKui, et al.
Published: (2025) -
Universal Image Restoration Pre-training via Masked Degradation Classification
by: Hu, JiaKui, et al.
Published: (2025)