Product of Experts for Visual Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yunzhi, Murtuza-Lanier, Carson, Li, Zizhang, Du, Yilun, Wu, Jiajun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Scene Language: Representing Scenes with Programs, Words, and Embeddings
by: Zhang, Yunzhi, et al.
Published: (2024)
by: Zhang, Yunzhi, et al.
Published: (2024)
Ctrl-VI: Controllable Video Synthesis via Variational Inference
by: Duan, Haoyi, et al.
Published: (2025)
by: Duan, Haoyi, et al.
Published: (2025)
RealWonder: Real-Time Physical Action-Conditioned Video Generation
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
WonderPlay: Dynamic 3D Scene Generation from a Single Image and Actions
by: Li, Zizhang, et al.
Published: (2025)
by: Li, Zizhang, et al.
Published: (2025)
Coupled Diffusion Sampling for Training-Free Multi-View Image Editing
by: Alzayer, Hadi, et al.
Published: (2025)
by: Alzayer, Hadi, et al.
Published: (2025)
HyKid: An Open MRI Dataset with Expert-Annotated Multi-Structure and Choroid Plexus in Pediatric Hydrocephalus
by: Xu, Yunzhi, et al.
Published: (2025)
by: Xu, Yunzhi, et al.
Published: (2025)
Long-Text-to-Image Generation via Compositional Prompt Decomposition
by: Huang, Jen-Yuan, et al.
Published: (2026)
by: Huang, Jen-Yuan, et al.
Published: (2026)
Diffusion Self-Distillation for Zero-Shot Customized Image Generation
by: Cai, Shengqu, et al.
Published: (2024)
by: Cai, Shengqu, et al.
Published: (2024)
Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models
by: Wang, Runqian, et al.
Published: (2025)
by: Wang, Runqian, et al.
Published: (2025)
GPIC: A Giant Permissive Image Corpus for Visual Generation
by: Chandrasegaran, Keshigeyan, et al.
Published: (2026)
by: Chandrasegaran, Keshigeyan, et al.
Published: (2026)
Compositional Generative Modeling: A Single Model is Not All You Need
by: Du, Yilun, et al.
Published: (2024)
by: Du, Yilun, et al.
Published: (2024)
Adversarial Attacks on VQA-NLE: Exposing and Alleviating Inconsistencies in Visual Question Answering Explanations
by: Yeh, Yahsin, et al.
Published: (2025)
by: Yeh, Yahsin, et al.
Published: (2025)
Language-Informed Visual Concept Learning
by: Lee, Sharon, et al.
Published: (2023)
by: Lee, Sharon, et al.
Published: (2023)
3D Congealing: 3D-Aware Image Alignment in the Wild
by: Zhang, Yunzhi, et al.
Published: (2024)
by: Zhang, Yunzhi, et al.
Published: (2024)
Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning
by: Cheng, Hanbo, et al.
Published: (2026)
by: Cheng, Hanbo, et al.
Published: (2026)
DiffusionAgent: Navigating Expert Models for Agentic Image Generation
by: Qin, Jie, et al.
Published: (2024)
by: Qin, Jie, et al.
Published: (2024)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
by: Zhang, Qizhe, et al.
Published: (2024)
by: Zhang, Qizhe, et al.
Published: (2024)
VIGC: Visual Instruction Generation and Correction
by: Wang, Bin, et al.
Published: (2023)
by: Wang, Bin, et al.
Published: (2023)
PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation
by: Zhan, Jiahao, et al.
Published: (2026)
by: Zhan, Jiahao, et al.
Published: (2026)
Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test Time Adaptation
by: Zhang, Rongyu, et al.
Published: (2024)
by: Zhang, Rongyu, et al.
Published: (2024)
Make the Pertinent Salient: Task-Relevant Reconstruction for Visual Control with Distractions
by: Kim, Kyungmin, et al.
Published: (2024)
by: Kim, Kyungmin, et al.
Published: (2024)
Multi-Sourced Compositional Generalization in Visual Question Answering
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation
by: Zhang, Guoqing, et al.
Published: (2025)
by: Zhang, Guoqing, et al.
Published: (2025)
MaS-VQA: A Mask-and-Select Framework for Knowledge-Based Visual Question Answering
by: Mao, Xianwei, et al.
Published: (2026)
by: Mao, Xianwei, et al.
Published: (2026)
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
AdaptGCD: Multi-Expert Adapter Tuning for Generalized Category Discovery
by: Qu, Yuxun, et al.
Published: (2024)
by: Qu, Yuxun, et al.
Published: (2024)
Evaluating the Capability of Video Question Generation for Expert Knowledge Elicitation
by: Zhang, Huaying, et al.
Published: (2025)
by: Zhang, Huaying, et al.
Published: (2025)
ExpertSim: Fast Particle Detector Simulation Using Mixture-of-Generative-Experts
by: Będkowski, Patryk, et al.
Published: (2025)
by: Będkowski, Patryk, et al.
Published: (2025)
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion
by: Zhang, Junkai, et al.
Published: (2025)
by: Zhang, Junkai, et al.
Published: (2025)
From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries
by: Hsu, Joy, et al.
Published: (2025)
by: Hsu, Joy, et al.
Published: (2025)
Content-Conditioned Generation of Stylized Free hand Sketches
by: Liu, Jiajun, et al.
Published: (2024)
by: Liu, Jiajun, et al.
Published: (2024)
Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
by: Dharmarajan, Karthik, et al.
Published: (2025)
by: Dharmarajan, Karthik, et al.
Published: (2025)
Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
by: Lu, Shuai, et al.
Published: (2026)
by: Lu, Shuai, et al.
Published: (2026)
Grounding Video Models to Actions through Goal Conditioned Exploration
by: Luo, Yunhao, et al.
Published: (2024)
by: Luo, Yunhao, et al.
Published: (2024)
Learning the 3D Fauna of the Web
by: Li, Zizhang, et al.
Published: (2024)
by: Li, Zizhang, et al.
Published: (2024)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
by: Liu, Xiaolin, et al.
Published: (2026)
by: Liu, Xiaolin, et al.
Published: (2026)
Synthesizing Multimodal Geometry Datasets from Scratch and Enabling Visual Alignment via Plotting Code
by: Lin, Haobo, et al.
Published: (2026)
by: Lin, Haobo, et al.
Published: (2026)
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
WonderZoom: Multi-Scale 3D World Generation
by: Cao, Jin, et al.
Published: (2025)
by: Cao, Jin, et al.
Published: (2025)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
by: Xu, Yu, et al.
Published: (2026)
by: Xu, Yu, et al.
Published: (2026)
Similar Items
-
The Scene Language: Representing Scenes with Programs, Words, and Embeddings
by: Zhang, Yunzhi, et al.
Published: (2024) -
Ctrl-VI: Controllable Video Synthesis via Variational Inference
by: Duan, Haoyi, et al.
Published: (2025) -
RealWonder: Real-Time Physical Action-Conditioned Video Generation
by: Liu, Wei, et al.
Published: (2026) -
WonderPlay: Dynamic 3D Scene Generation from a Single Image and Actions
by: Li, Zizhang, et al.
Published: (2025) -
Coupled Diffusion Sampling for Training-Free Multi-View Image Editing
by: Alzayer, Hadi, et al.
Published: (2025)