JoPano: Unified Panorama Generation via Joint Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Wancheng, An, Chen, He, Zhenliang, Kan, Meina, Shan, Shiguang, Wang, Lukun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jodi: Unification of Visual Generation and Understanding via Joint Modeling
by: Xu, Yifeng, et al.
Published: (2025)
by: Xu, Yifeng, et al.
Published: (2025)
OSI: One-step Inversion Excels in Extracting Diffusion Watermarks
by: Chen, Yuwei, et al.
Published: (2026)
by: Chen, Yuwei, et al.
Published: (2026)
PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World
by: Wang, Changpeng, et al.
Published: (2026)
by: Wang, Changpeng, et al.
Published: (2026)
DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware Diffusion
by: Ye, Weicai, et al.
Published: (2024)
by: Ye, Weicai, et al.
Published: (2024)
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
by: Li, Yiheng, et al.
Published: (2024)
by: Li, Yiheng, et al.
Published: (2024)
Trigger without Trace: Towards Stealthy Backdoor Attack on Text-to-Image Diffusion Models
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
CtrLoRA: An Extensible and Efficient Framework for Controllable Image Generation
by: Xu, Yifeng, et al.
Published: (2024)
by: Xu, Yifeng, et al.
Published: (2024)
Neural Gate: Mitigating Privacy Risks in LVLMs via Neuron-Level Gradient Gating
by: Cao, Xiangkui, et al.
Published: (2026)
by: Cao, Xiangkui, et al.
Published: (2026)
CognitionCapturerPro: Towards High-Fidelity Visual Decoding from EEG/MEG via Multi-modal Information and Asymmetric Alignment
by: Zhang, Kaifan, et al.
Published: (2026)
by: Zhang, Kaifan, et al.
Published: (2026)
PanoWorld: Geometry-Consistent Panoramic Video World Modeling
by: Jiang, Le, et al.
Published: (2026)
by: Jiang, Le, et al.
Published: (2026)
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
by: Yan, Bei, et al.
Published: (2024)
by: Yan, Bei, et al.
Published: (2024)
PanoLora: Bridging Perspective and Panoramic Video Generation with LoRA Adaptation
by: Dong, Zeyu, et al.
Published: (2025)
by: Dong, Zeyu, et al.
Published: (2025)
Semantic Generative Tuning for Unified Multimodal Models
by: Yu, Songsong, et al.
Published: (2026)
by: Yu, Songsong, et al.
Published: (2026)
Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
by: Yan, Bei, et al.
Published: (2024)
by: Yan, Bei, et al.
Published: (2024)
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
by: Li, Yiheng, et al.
Published: (2026)
by: Li, Yiheng, et al.
Published: (2026)
PanoTPS-Net: Panoramic Room Layout Estimation via Thin Plate Spline Transformation
by: Ibrahem, Hatem, et al.
Published: (2025)
by: Ibrahem, Hatem, et al.
Published: (2025)
StyleBrush: Style Extraction and Transfer from a Single Image
by: Feng, Wancheng, et al.
Published: (2024)
by: Feng, Wancheng, et al.
Published: (2024)
VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
by: Wang, Sibo, et al.
Published: (2024)
by: Wang, Sibo, et al.
Published: (2024)
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
by: Yang, Junqi, et al.
Published: (2026)
by: Yang, Junqi, et al.
Published: (2026)
Apollo: Unified Multi-Task Audio-Video Joint Generation
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
HPNet: Dynamic Trajectory Forecasting with Historical Prediction Attention
by: Tang, Xiaolong, et al.
Published: (2024)
by: Tang, Xiaolong, et al.
Published: (2024)
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
by: Chen, Haoyu, et al.
Published: (2026)
by: Chen, Haoyu, et al.
Published: (2026)
Pose-Robust Calibration Strategy for Point-of-Gaze Estimation on Mobile Phones
by: Zhao, Yujie, et al.
Published: (2025)
by: Zhao, Yujie, et al.
Published: (2025)
EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
by: Cai, Xinyan, et al.
Published: (2025)
by: Cai, Xinyan, et al.
Published: (2025)
A Survey on Text-Driven 360-Degree Panorama Generation
by: Wang, Hai, et al.
Published: (2025)
by: Wang, Hai, et al.
Published: (2025)
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation
by: Huang, Xiaohu, et al.
Published: (2025)
by: Huang, Xiaohu, et al.
Published: (2025)
LensWalk: Agentic Video Understanding by Planning How You See in Videos
by: Li, Keliang, et al.
Published: (2026)
by: Li, Keliang, et al.
Published: (2026)
Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
by: Kang, Xueyang, et al.
Published: (2025)
by: Kang, Xueyang, et al.
Published: (2025)
CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
by: Cao, Zongsheng, et al.
Published: (2025)
by: Cao, Zongsheng, et al.
Published: (2025)
360PanT: Training-Free Text-Driven 360-Degree Panorama-to-Panorama Translation
by: Wang, Hai, et al.
Published: (2024)
by: Wang, Hai, et al.
Published: (2024)
Unified Editing of Panorama, 3D Scenes, and Videos Through Disentangled Self-Attention Injection
by: Kwon, Gihyun, et al.
Published: (2024)
by: Kwon, Gihyun, et al.
Published: (2024)
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling
by: Wang, Yuran, et al.
Published: (2025)
by: Wang, Yuran, et al.
Published: (2025)
RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models
by: Lin, Yijing, et al.
Published: (2025)
by: Lin, Yijing, et al.
Published: (2025)
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
by: Zhang, Huichao, et al.
Published: (2026)
by: Zhang, Huichao, et al.
Published: (2026)
UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
by: Han, Ruiyan, et al.
Published: (2026)
by: Han, Ruiyan, et al.
Published: (2026)
P2DNav: Panorama-to-Downview Reasoning for Zero-shot Vision-and-Language Navigation
by: Sheng, Kai, et al.
Published: (2026)
by: Sheng, Kai, et al.
Published: (2026)
Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
by: Mao, Jiawei, et al.
Published: (2025)
by: Mao, Jiawei, et al.
Published: (2025)
IAD-Unify: A Region-Grounded Unified Model for Industrial Anomaly Segmentation, Understanding, and Generation
by: Zheng, Haoyu, et al.
Published: (2026)
by: Zheng, Haoyu, et al.
Published: (2026)
OmniAlpha: Aligning Transparency-Aware Generation via Multi-Task Unified Reinforcement Learning
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
PointCG: Self-supervised Point Cloud Learning via Joint Completion and Generation
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
Similar Items
-
Jodi: Unification of Visual Generation and Understanding via Joint Modeling
by: Xu, Yifeng, et al.
Published: (2025) -
OSI: One-step Inversion Excels in Extracting Diffusion Watermarks
by: Chen, Yuwei, et al.
Published: (2026) -
PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World
by: Wang, Changpeng, et al.
Published: (2026) -
DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware Diffusion
by: Ye, Weicai, et al.
Published: (2024) -
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
by: Li, Yiheng, et al.
Published: (2024)