Jodi: Unification of Visual Generation and Understanding via Joint Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yifeng, He, Zhenliang, Kan, Meina, Shan, Shiguang, Chen, Xilin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JoPano: Unified Panorama Generation via Joint Modeling
by: Feng, Wancheng, et al.
Published: (2025)
by: Feng, Wancheng, et al.
Published: (2025)
CtrLoRA: An Extensible and Efficient Framework for Controllable Image Generation
by: Xu, Yifeng, et al.
Published: (2024)
by: Xu, Yifeng, et al.
Published: (2024)
Dual Attention Guided Defense Against Malicious Edits
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Towards Transferable Defense Against Malicious Image Edits
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Semantic Mismatch and Perceptual Degradation: A New Perspective on Image Editing Immunity
by: Dong, Shuai, et al.
Published: (2025)
by: Dong, Shuai, et al.
Published: (2025)
OSI: One-step Inversion Excels in Extracting Diffusion Watermarks
by: Chen, Yuwei, et al.
Published: (2026)
by: Chen, Yuwei, et al.
Published: (2026)
Trigger without Trace: Towards Stealthy Backdoor Attack on Text-to-Image Diffusion Models
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Neural Gate: Mitigating Privacy Risks in LVLMs via Neuron-Level Gradient Gating
by: Cao, Xiangkui, et al.
Published: (2026)
by: Cao, Xiangkui, et al.
Published: (2026)
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
by: Li, Yiheng, et al.
Published: (2024)
by: Li, Yiheng, et al.
Published: (2024)
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
by: Yan, Bei, et al.
Published: (2024)
by: Yan, Bei, et al.
Published: (2024)
Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
by: Yan, Bei, et al.
Published: (2024)
by: Yan, Bei, et al.
Published: (2024)
Generalized Semi-Supervised Learning via Self-Supervised Feature Adaptation
by: Liang, Jiachen, et al.
Published: (2024)
by: Liang, Jiachen, et al.
Published: (2024)
Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual Representation
by: Han, Boyu, et al.
Published: (2026)
by: Han, Boyu, et al.
Published: (2026)
Understanding Visual Concepts Across Models
by: Trabucco, Brandon, et al.
Published: (2024)
by: Trabucco, Brandon, et al.
Published: (2024)
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
by: Yang, Junqi, et al.
Published: (2026)
by: Yang, Junqi, et al.
Published: (2026)
A Geometric Unification of Concept Learning with Concept Cones
by: Rocchi--Henry, Alexandre, et al.
Published: (2025)
by: Rocchi--Henry, Alexandre, et al.
Published: (2025)
Understanding Implosion in Text-to-Image Generative Models
by: Ding, Wenxin, et al.
Published: (2024)
by: Ding, Wenxin, et al.
Published: (2024)
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
by: Cho, Jang Hyun, et al.
Published: (2025)
by: Cho, Jang Hyun, et al.
Published: (2025)
EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds
by: Chen, Lu, et al.
Published: (2025)
by: Chen, Lu, et al.
Published: (2025)
PackDiT: Joint Human Motion and Text Generation via Mutual Prompting
by: Jiang, Zhongyu, et al.
Published: (2025)
by: Jiang, Zhongyu, et al.
Published: (2025)
HPNet: Dynamic Trajectory Forecasting with Historical Prediction Attention
by: Tang, Xiaolong, et al.
Published: (2024)
by: Tang, Xiaolong, et al.
Published: (2024)
Visual Generation Without Guidance
by: Chen, Huayu, et al.
Published: (2025)
by: Chen, Huayu, et al.
Published: (2025)
Composition Vision-Language Understanding via Segment and Depth Anything Model
by: Huo, Mingxiao, et al.
Published: (2024)
by: Huo, Mingxiao, et al.
Published: (2024)
VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
by: Wang, Sibo, et al.
Published: (2024)
by: Wang, Sibo, et al.
Published: (2024)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023)
by: Lai, Zhengfeng, et al.
Published: (2023)
Generative Visual Code Mobile World Models
by: Koh, Woosung, et al.
Published: (2026)
by: Koh, Woosung, et al.
Published: (2026)
Next Visual Granularity Generation
by: Wang, Yikai, et al.
Published: (2025)
by: Wang, Yikai, et al.
Published: (2025)
Learning Separable Hidden Unit Contributions for Speaker-Adaptive Lip-Reading
by: Luo, Songtao, et al.
Published: (2023)
by: Luo, Songtao, et al.
Published: (2023)
RGB-Th-Bench: A Dense benchmark for Visual-Thermal Understanding of Vision Language Models
by: Moshtaghi, Mehdi, et al.
Published: (2025)
by: Moshtaghi, Mehdi, et al.
Published: (2025)
VideoNSA: Native Sparse Attention Scales Video Understanding
by: Song, Enxin, et al.
Published: (2025)
by: Song, Enxin, et al.
Published: (2025)
MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
by: Liu, Ziyu, et al.
Published: (2024)
by: Liu, Ziyu, et al.
Published: (2024)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
by: Rao, Zhefan, et al.
Published: (2024)
by: Rao, Zhefan, et al.
Published: (2024)
Reversible Unfolding Network for Concealed Visual Perception with Generative Refinement
by: He, Chunming, et al.
Published: (2025)
by: He, Chunming, et al.
Published: (2025)
OmniPrism: Learning Disentangled Visual Concept for Image Generation
by: Li, Yangyang, et al.
Published: (2024)
by: Li, Yangyang, et al.
Published: (2024)
VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation
by: Zou, Bocheng, et al.
Published: (2024)
by: Zou, Bocheng, et al.
Published: (2024)
SafeFix: Targeted Model Repair via Controlled Image Generation
by: Xu, Ouyang, et al.
Published: (2025)
by: Xu, Ouyang, et al.
Published: (2025)
SpectralAR: Spectral Autoregressive Visual Generation
by: Huang, Yuanhui, et al.
Published: (2025)
by: Huang, Yuanhui, et al.
Published: (2025)
The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design
by: Liu, Anjie, et al.
Published: (2026)
by: Liu, Anjie, et al.
Published: (2026)
What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
by: Farid, Karim, et al.
Published: (2025)
by: Farid, Karim, et al.
Published: (2025)
Similar Items
-
JoPano: Unified Panorama Generation via Joint Modeling
by: Feng, Wancheng, et al.
Published: (2025) -
CtrLoRA: An Extensible and Efficient Framework for Controllable Image Generation
by: Xu, Yifeng, et al.
Published: (2024) -
Dual Attention Guided Defense Against Malicious Edits
by: Zhang, Jie, et al.
Published: (2025) -
Towards Transferable Defense Against Malicious Image Edits
by: Zhang, Jie, et al.
Published: (2025) -
Semantic Mismatch and Perceptual Degradation: A New Perspective on Image Editing Immunity
by: Dong, Shuai, et al.
Published: (2025)