MuLan: Multimodal-LLM Agent for Progressive and Interactive Multi-Object Diffusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Sen, Wang, Ruochen, Hsieh, Cho-Jui, Cheng, Minhao, Zhou, Tianyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding the Impact of Negative Prompts: When and How Do They Take Effect?
von: Ban, Yuanhao, et al.
Veröffentlicht: (2024)
von: Ban, Yuanhao, et al.
Veröffentlicht: (2024)
MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?
von: Li, Xirui, et al.
Veröffentlicht: (2024)
von: Li, Xirui, et al.
Veröffentlicht: (2024)
R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
von: Zhou, Hengguang, et al.
Veröffentlicht: (2025)
von: Zhou, Hengguang, et al.
Veröffentlicht: (2025)
The Crystal Ball Hypothesis in diffusion models: Anticipating object positions from initial noise
von: Ban, Yuanhao, et al.
Veröffentlicht: (2024)
von: Ban, Yuanhao, et al.
Veröffentlicht: (2024)
Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
von: Bai, Andrew, et al.
Veröffentlicht: (2025)
von: Bai, Andrew, et al.
Veröffentlicht: (2025)
On Discrete Prompt Optimization for Diffusion Models
von: Wang, Ruochen, et al.
Veröffentlicht: (2024)
von: Wang, Ruochen, et al.
Veröffentlicht: (2024)
MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost
von: Xing, Sen, et al.
Veröffentlicht: (2024)
von: Xing, Sen, et al.
Veröffentlicht: (2024)
Mitigating Bias in Dataset Distillation
von: Cui, Justin, et al.
Veröffentlicht: (2024)
von: Cui, Justin, et al.
Veröffentlicht: (2024)
QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models
von: Kao, Kuei-Chun, et al.
Veröffentlicht: (2025)
von: Kao, Kuei-Chun, et al.
Veröffentlicht: (2025)
MuSEAgent: A Multimodal Reasoning Agent with Stateful Experiences
von: Wang, Shijian, et al.
Veröffentlicht: (2026)
von: Wang, Shijian, et al.
Veröffentlicht: (2026)
Invisible Backdoor Attacks on Diffusion Models
von: Li, Sen, et al.
Veröffentlicht: (2024)
von: Li, Sen, et al.
Veröffentlicht: (2024)
MuLan: A Study of Fact Mutability in Language Models
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
LanDA: Language-Guided Multi-Source Domain Adaptation
von: Wang, Zhenbin, et al.
Veröffentlicht: (2024)
von: Wang, Zhenbin, et al.
Veröffentlicht: (2024)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
von: Li, Xirui, et al.
Veröffentlicht: (2024)
von: Li, Xirui, et al.
Veröffentlicht: (2024)
Understanding Reward Hacking in Text-to-Image Reinforcement Learning
von: Hong, Yunqi, et al.
Veröffentlicht: (2026)
von: Hong, Yunqi, et al.
Veröffentlicht: (2026)
From Category to Scenery: An End-to-End Framework for Multi-Person Human-Object Interaction Recognition in Videos
von: Qiao, Tanqiu, et al.
Veröffentlicht: (2024)
von: Qiao, Tanqiu, et al.
Veröffentlicht: (2024)
Nodule-Aligned Latent Space Learning with LLM-Driven Multimodal Diffusion for Lung Nodule Progression Prediction
von: Song, James, et al.
Veröffentlicht: (2026)
von: Song, James, et al.
Veröffentlicht: (2026)
MuSLR: Multimodal Symbolic Logical Reasoning
von: Xu, Jundong, et al.
Veröffentlicht: (2025)
von: Xu, Jundong, et al.
Veröffentlicht: (2025)
MuDG: Taming Multi-modal Diffusion with Gaussian Splatting for Urban Scene Reconstruction
von: Zou, Yingshuang, et al.
Veröffentlicht: (2025)
von: Zou, Yingshuang, et al.
Veröffentlicht: (2025)
Geometric Visual Fusion Graph Neural Networks for Multi-Person Human-Object Interaction Recognition in Videos
von: Qiao, Tanqiu, et al.
Veröffentlicht: (2025)
von: Qiao, Tanqiu, et al.
Veröffentlicht: (2025)
PICS: Pairwise Image Compositing with Spatial Interactions
von: Zhou, Hang, et al.
Veröffentlicht: (2026)
von: Zhou, Hang, et al.
Veröffentlicht: (2026)
LanEvil: Benchmarking the Robustness of Lane Detection to Environmental Illusions
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorld
von: Yang, Yijun, et al.
Veröffentlicht: (2023)
von: Yang, Yijun, et al.
Veröffentlicht: (2023)
JIGMARK: A Black-Box Approach for Enhancing Image Watermarks against Diffusion Model Edits
von: Pan, Minzhou, et al.
Veröffentlicht: (2024)
von: Pan, Minzhou, et al.
Veröffentlicht: (2024)
Evolving from Single-modal to Multi-modal Facial Deepfake Detection: Progress and Challenges
von: Liu, Ping, et al.
Veröffentlicht: (2024)
von: Liu, Ping, et al.
Veröffentlicht: (2024)
One-Forcing: Towards Stable One-Step Autoregressive Video Generation
von: Feng, Jiaqi, et al.
Veröffentlicht: (2026)
von: Feng, Jiaqi, et al.
Veröffentlicht: (2026)
Embedding Space Selection for Detecting Memorization and Fingerprinting in Generative Models
von: He, Jack, et al.
Veröffentlicht: (2024)
von: He, Jack, et al.
Veröffentlicht: (2024)
Control Color: Multimodal Diffusion-based Interactive Image Colorization
von: Liang, Zhexin, et al.
Veröffentlicht: (2024)
von: Liang, Zhexin, et al.
Veröffentlicht: (2024)
Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
von: Hong, Yunqi, et al.
Veröffentlicht: (2025)
von: Hong, Yunqi, et al.
Veröffentlicht: (2025)
DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders
von: Hong, Susung, et al.
Veröffentlicht: (2025)
von: Hong, Susung, et al.
Veröffentlicht: (2025)
MuSc-V2: Zero-Shot Multimodal Industrial Anomaly Classification and Segmentation with Mutual Scoring of Unlabeled Samples
von: Li, Xurui, et al.
Veröffentlicht: (2025)
von: Li, Xurui, et al.
Veröffentlicht: (2025)
Multimodal Graph Network Modeling for Human-Object Interaction Detection with PDE Graph Diffusion
von: Ji, Wenxuan, et al.
Veröffentlicht: (2025)
von: Ji, Wenxuan, et al.
Veröffentlicht: (2025)
Adversarial Examples Detection with Bayesian Neural Network
von: Li, Yao, et al.
Veröffentlicht: (2021)
von: Li, Yao, et al.
Veröffentlicht: (2021)
Large Language Models are Interpretable Learners
von: Wang, Ruochen, et al.
Veröffentlicht: (2024)
von: Wang, Ruochen, et al.
Veröffentlicht: (2024)
MuGa-VTON: Multi-Garment Virtual Try-On via Diffusion Transformers with Prompt Customization
von: Deria, Ankan, et al.
Veröffentlicht: (2025)
von: Deria, Ankan, et al.
Veröffentlicht: (2025)
RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba
von: Lu, Andong, et al.
Veröffentlicht: (2024)
von: Lu, Andong, et al.
Veröffentlicht: (2024)
MuRF: Multi-Baseline Radiance Fields
von: Xu, Haofei, et al.
Veröffentlicht: (2023)
von: Xu, Haofei, et al.
Veröffentlicht: (2023)
MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation
von: Zhou, Bohan, et al.
Veröffentlicht: (2025)
von: Zhou, Bohan, et al.
Veröffentlicht: (2025)
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models
von: Wu, Zongyu, et al.
Veröffentlicht: (2025)
von: Wu, Zongyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Understanding the Impact of Negative Prompts: When and How Do They Take Effect?
von: Ban, Yuanhao, et al.
Veröffentlicht: (2024) -
MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?
von: Li, Xirui, et al.
Veröffentlicht: (2024) -
R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
von: Zhou, Hengguang, et al.
Veröffentlicht: (2025) -
The Crystal Ball Hypothesis in diffusion models: Anticipating object positions from initial noise
von: Ban, Yuanhao, et al.
Veröffentlicht: (2024) -
Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
von: Bai, Andrew, et al.
Veröffentlicht: (2025)