Multimodal ELBO with Diffusion Decoders
Fuente:
arXiv
Saved in:
| Main Authors: | Wesego, Daniel, Rooshenas, Pedram |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Score-Based Multimodal Autoencoder
by: Wesego, Daniel, et al.
Published: (2023)
by: Wesego, Daniel, et al.
Published: (2023)
TARO: Temporal Adversarial Rectification Optimization Using Diffusion Models as Purifiers
by: Wesego, Daniel, et al.
Published: (2026)
by: Wesego, Daniel, et al.
Published: (2026)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
by: Wang, Yuduo, et al.
Published: (2023)
by: Wang, Yuduo, et al.
Published: (2023)
Gradient-free Decoder Inversion in Latent Diffusion Models
by: Hong, Seongmin, et al.
Published: (2024)
by: Hong, Seongmin, et al.
Published: (2024)
ELBO-T2IAlign: A Generic ELBO-Based Method for Calibrating Pixel-level Text-Image Alignment in Diffusion Models
by: Zhou, Qin, et al.
Published: (2025)
by: Zhou, Qin, et al.
Published: (2025)
LatentHDR: Decoupling Exposure from Diffusion via Conditional Latent-to-Latent Mapping for Text/Image-to-Panoramic HDR
by: Fekri, Pedram, et al.
Published: (2026)
by: Fekri, Pedram, et al.
Published: (2026)
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
by: Agarwal, Vatsal, et al.
Published: (2025)
by: Agarwal, Vatsal, et al.
Published: (2025)
ACDC: Autoregressive Coherent Multimodal Generation using Diffusion Correction
by: Chung, Hyungjin, et al.
Published: (2024)
by: Chung, Hyungjin, et al.
Published: (2024)
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
by: Liang, Zhixuan, et al.
Published: (2025)
by: Liang, Zhixuan, et al.
Published: (2025)
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
by: Fan, Xiang, et al.
Published: (2026)
by: Fan, Xiang, et al.
Published: (2026)
LILAC: Long-sequence Incremental Low-latency Arbitrary Motion Stylization via Streaming VAE-Diffusion with Causal Decoding
by: Ren, Peng, et al.
Published: (2025)
by: Ren, Peng, et al.
Published: (2025)
MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models
by: Ganesan, Mugilan, et al.
Published: (2025)
by: Ganesan, Mugilan, et al.
Published: (2025)
NPNet: A Non-Parametric Network with Adaptive Gaussian-Fourier Positional Encoding for 3D Classification and Segmentation
by: Saeid, Mohammad, et al.
Published: (2026)
by: Saeid, Mohammad, et al.
Published: (2026)
Causal Decoding for Hallucination-Resistant Multimodal Large Language Models
by: Tan, Shiwei, et al.
Published: (2026)
by: Tan, Shiwei, et al.
Published: (2026)
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
by: Hong, Ji Woo, et al.
Published: (2026)
by: Hong, Ji Woo, et al.
Published: (2026)
DepMicroDiff: Diffusion-Based Dependency-Aware Multimodal Imputation for Microbiome Data
by: Sadia, Rabeya Tus, et al.
Published: (2025)
by: Sadia, Rabeya Tus, et al.
Published: (2025)
Mask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change Captioning
by: Sun, Dongwei, et al.
Published: (2024)
by: Sun, Dongwei, et al.
Published: (2024)
REED-VAE: RE-Encode Decode Training for Iterative Image Editing with Diffusion Models
by: Almog, Gal, et al.
Published: (2025)
by: Almog, Gal, et al.
Published: (2025)
VGS-Decoding: Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs
by: Kolli, Govinda, et al.
Published: (2026)
by: Kolli, Govinda, et al.
Published: (2026)
Multimodal Latent Language Modeling with Next-Token Diffusion
by: Sun, Yutao, et al.
Published: (2024)
by: Sun, Yutao, et al.
Published: (2024)
Flows and Diffusions on the Neural Manifold
by: Saragih, Daniel, et al.
Published: (2025)
by: Saragih, Daniel, et al.
Published: (2025)
Invariant Representation Guided Multimodal Sentiment Decoding with Sequential Variation Regularization
by: Xu, Guoyang, et al.
Published: (2024)
by: Xu, Guoyang, et al.
Published: (2024)
MDiFF: Exploiting Multimodal Score-based Diffusion Models for New Fashion Product Performance Forecasting
by: Avogaro, Andrea, et al.
Published: (2024)
by: Avogaro, Andrea, et al.
Published: (2024)
Spatial Gated Multi-Layer Perceptron for Land Use and Land Cover Mapping
by: Jamali, Ali, et al.
Published: (2023)
by: Jamali, Ali, et al.
Published: (2023)
Multimodal Latent Diffusion Model for Complex Sewing Pattern Generation
by: Liu, Shengqi, et al.
Published: (2024)
by: Liu, Shengqi, et al.
Published: (2024)
Controlling Multimodal LLMs via Reward-guided Decoding
by: Mañas, Oscar, et al.
Published: (2025)
by: Mañas, Oscar, et al.
Published: (2025)
Score-based Membership Inference on Diffusion Models
by: Rao, Mingxing, et al.
Published: (2025)
by: Rao, Mingxing, et al.
Published: (2025)
Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces
by: Rojas, Kevin, et al.
Published: (2025)
by: Rojas, Kevin, et al.
Published: (2025)
CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text
by: Rani, Anju, et al.
Published: (2025)
by: Rani, Anju, et al.
Published: (2025)
Memory-Efficient Vision Transformers: An Activation-Aware Mixed-Rank Compression Strategy
by: Azizi, Seyedarmin, et al.
Published: (2024)
by: Azizi, Seyedarmin, et al.
Published: (2024)
DMODE: Differential Monocular Object Distance Estimation Module without Class Specific Information
by: Agand, Pedram, et al.
Published: (2022)
by: Agand, Pedram, et al.
Published: (2022)
Latent Diffusion Inversion Requires Understanding the Latent Space
by: Rao, Mingxing, et al.
Published: (2025)
by: Rao, Mingxing, et al.
Published: (2025)
DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders
by: Hong, Susung, et al.
Published: (2025)
by: Hong, Susung, et al.
Published: (2025)
Finding Shared Decodable Concepts and their Negations in the Brain
by: Efird, Cory, et al.
Published: (2024)
by: Efird, Cory, et al.
Published: (2024)
VMDT: Decoding the Trustworthiness of Video Foundation Models
by: Potter, Yujin, et al.
Published: (2025)
by: Potter, Yujin, et al.
Published: (2025)
Visual Diffusion Models are Geometric Solvers
by: Goren, Nir, et al.
Published: (2025)
by: Goren, Nir, et al.
Published: (2025)
Dif4FF: Leveraging Multimodal Diffusion Models and Graph Neural Networks for Accurate New Fashion Product Performance Forecasting
by: Avogaro, Andrea, et al.
Published: (2024)
by: Avogaro, Andrea, et al.
Published: (2024)
DeCLIP: Decoding CLIP representations for deepfake localization
by: Smeu, Stefan, et al.
Published: (2024)
by: Smeu, Stefan, et al.
Published: (2024)
Decoding Federated Learning: The FedNAM+ Conformal Revolution
by: Balija, Sree Bhargavi, et al.
Published: (2025)
by: Balija, Sree Bhargavi, et al.
Published: (2025)
Domain Independent SVM for Transfer Learning in Brain Decoding
by: Zhou, Shuo, et al.
Published: (2019)
by: Zhou, Shuo, et al.
Published: (2019)
Similar Items
-
Score-Based Multimodal Autoencoder
by: Wesego, Daniel, et al.
Published: (2023) -
TARO: Temporal Adversarial Rectification Optimization Using Diffusion Models as Purifiers
by: Wesego, Daniel, et al.
Published: (2026) -
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
by: Wang, Yuduo, et al.
Published: (2023) -
Gradient-free Decoder Inversion in Latent Diffusion Models
by: Hong, Seongmin, et al.
Published: (2024) -
ELBO-T2IAlign: A Generic ELBO-Based Method for Calibrating Pixel-level Text-Image Alignment in Diffusion Models
by: Zhou, Qin, et al.
Published: (2025)